Skip to content
TrackPodcasts
scienceMar 13, 202623:47

244 – Scaling Open Textbook Variants with PreTeXt and AI

About this episode

Lily and David continue their discussions on converting open textbooks into PreTeXt. They focus on the “Learning Statistics with …” ecosystem, where an original open book has spawned variants for R, JASP, Jamovi, CogStat, French, and potential new versions such as R-Instat. They explore how PreTeXt could better manage multiple independently maintained variants by identifying what differs, easing updates from a base text, and supporting responsible human ownership.

Get every episode summarized

Each time The IDEMS Podcast publishes, we email you a written briefing from the transcript — the topics, who appeared, and any specific claims, with the ad reads skipped.

Email me new episodes

Free for 3 shows. No card needed.

Hosts & guests

Transcript ready

235 searchable segments. Every word is indexed and playable.

244 – Scaling Open Textbook Variants with PreTeXt and AI

The IDEMS Podcast

0:00
23:47

Full transcript

The IDEMS Podcast244 – Scaling Open Textbook Variants with PreTeXt and AI. Machine-transcribed; use the interactive transcript above to jump the player to any line.

Hello and welcome to the items podcast. I'm Lily Clements, a data scientist and I'm here with David Stone, a founding director of items. Hi David. Hi Lily. What are we discussing today? I thought it could be a good time for a book update on the pretext books. They're moving fast, so while I think our last update was relatively recently, it feels like an appropriate time for another one. It sounds good. So how many have you got to? 20. Well actually, I'm now on my 21st. I was made to stop here, but I came across an email of... And for me, Unwin. Oh yes, I love his work. Yeah. Yes. You was my lecturer back in the day when I was in Augsburg. So yeah, maybe you were doing that. Yes, yes. So I came across well an email that he sent to your Facebook, getting more out of graphics, which looks really interesting. So actually, I started playing around with that this morning and you're seeing, okay,

how doable is this? We're talking a lot of graphics. Let's just give it a try. So we're on 21, even though I said I would have stopped this month. But no, that's great. So and to the Unwin is the person who introduced me to the Grand World Graphic years and years ago. 25 years ago, I think it was. No, maybe not quite. Maybe it was more like 23 years ago, but it's still a while back. And this is part of what started me on this path to actually recognizing the importance of these structures, of these grammars, which we now, I would argue in many ways, the pretext work is related to this. The semanticness of it is, again, it's about trying to get that structure in a way which has meaning. And so this is really exciting that we're taking his book and putting it into pretext and potentially enabling that to bring added value in ways which could be unexpected and interesting. Oh, exciting. Yeah, no,

absolutely. Very exciting. And well, more specifically, I've got to say books because they are different books, but they cancel the same book. And so the one I want to discuss today is this, I guess, set of books, which are kind of learning statistics with, and then I say, I blank that because it was originally learning statistics with R, and then different people have covered along and made their variants. So now there's learning statistics with jazz, with Jermovie, with Coxbat. There's a French version, I believe. And there's all these different variants of the same book where this original book has been taken and formed into its own way. And we even now are talking about doing learning statistics in R in stat, which is another tool that you can use, another statistical tool that you could use, which we work with at items and that we've done many podcasts on before as well. So I won't go too much into that one here, I'm sure we will have our own one on that one another time. What I want to go into here is this kind of idea of these variants. Yeah. And this is really interesting because this is exactly what we

believe pretext can help us build a system to manage these multiple variants very differently. And to conceive books as having multiple variants, we've discussed this a bit in the past. And to have an open book which already has this concept emerging and to then try and discuss how can that be taken forward? How can we actually take what already exists and put it into a structure and actually try to understand what are the differences between these variants? How much has actually changed? What has changed? What hasn't changed? And to put them into a coherent, into a form of coherence, not into a coherent whole because each variant is its own entity, but actually being able to think of that more together. This is what we at items are really thinking deeply about. And it's a problem which comes up in all sorts of places, but this is a

beautiful illustration where each of those variants is owned and I'm putting owned as a sense of the people who have responsibility for it are different. So they're owned by separate entities or separate people or separate groups, but at the same time there is this sort of commonality around the open textbook and this is exactly what we want to imagine and encourage. How do we build systems which really enable this? This is the hard problem that we're trying to engage in. It's a very hard problem. It's one that scares me a lot and then I'm not sure even how I can't even... All I know is for now, I'm just taking a bit by bit, all I know is that this is something that there's clearly a need for. The people are creating these variants and therefore we want to add some ease to it if they change the base book, the original book. If we change this original book, then what updates do we do from that? And actually this is a book that has

changed. There are bits that have changed in their versions, which I've noticed. Okay, you've got this pack. Well, it's, you know, a small statement here or there. It's like, oh, you've got that one, but you've got that one. And I wonder if that was by choice that you chose to have that one? That's certain bullet point. The one in particular is Danielle refers to themselves as I throughout the book and there's one point in particular where they refer to themselves as a male in one version and in another version. They say, actually, I don't refer to myself that way anymore. And actually different versions of the book, different kind of variants of the book, some of them refer to the original and some of them then refer to, you know, having at this footnote saying, actually, this has changed, but for the sake of me not having to rewrite the whole book. But this is, you know, that's one specific piece. And it is exactly that idea of, well, how do we, when we have multiple variants, make those choices easily accessible? How do we make

that updating happen in other ways? This is a hard problem. And I'd love to actually be able to work with all the different partners on this of trying to say, and this is what we want to get towards, that we don't want to take over these things on the contrary. We're wanting to then gradually enable others to take their pieces of ownership and to find this added value that comes from working in a system which might support some of that ongoing maintenance. And you've been using AI for the translation, but one of the other things is, of course, what about an AI agent to act with the process of updating and that decision making? How do you do this in a way where it becomes conservative? And it's something which is made more accessible, but also at the same time where the human who is taking responsibility continues to be the person responsible. It's really interesting. There's a lot of scope. But as you say, it's not only relevant to these variants,

we could talk about it in different versions of stack questions and different versions of apps and different versions of courses, but just GitHub alone for anyone that uses GitHub. I guess now we're going to go into version control and that's definitely not a territory that I can have anything I'm used to say about other than that it seems very convoluted. I mean, it's complicated and it is really important. So the version control process is so important, but now being able to apply those to these textbook variants and versions is very nice. Yes, and here we have a really nice example where we can actually try, where we can actually play around and, you know, I'm more applied than you are. You're more theoretical. I'm more applied. I'd like to actually try it before I can understand it properly, whereas you can store it in your head. But no, I think that kind of part of the idea then and something that I alluded to is taking this learning statistics with our book and creating from this two versions, one version being learning statistics and that version in kind of our software agnostic version.

I think we've spoken about it on quite a few other podcasts about software agnostic and how having a version that doesn't rely on a specific software can then actually mean that the learning doesn't become about our or Python or that specific type, the learning instead, the emphasis gets to be on the actual statistics side itself. This is such a beautiful case to be working on, because I believe that the software agnostic variant can only really exist once you have at many different software specific variants. And so here we have a case where there is an R version, there's a Python version, there's a gem over and so on. And so these variants all exist already and therefore a software agnostic variant can refer to the others. And I want to come back and I'm not going to dig into this. It is something that has been discussed in previous episodes. But in a class, let's say a postgraduate class where you have students who have different

backgrounds, some of which might have an R background, others might have a Python background, others might have used Yarsport, your movie, or instead, to be able to have a textbook where they can use it and they can continue to use the software that they are comfortable in first, but be exposed to other software and where the class is a whole can be using a diversity of things. This isn't right in all contexts, but there are certainly contexts within which that is extremely valuable. The next year I'll be teaching at Ames, the African Institute of Mathematical Sciences. Next week, not next year. Oh, sorry, yes, next week. The African Institute of Mathematical Sciences on this doctoral training school and the profile of the different students just came in. And some of them are advanced R users, some are advanced Python users. There's this whole diversity and, of course, the course that I teach, which you've taught as well, is, of course, where it doesn't

matter what you've done. What matters is can you explore and investigate data and whatever tool you use, this is problem solving and statistics and data science, this is a course which I always get nervous teaching is, but I do love it because you never know what's going to happen because it depends which in the room, what they can do, what they can already do, what they then share, you know, billions of those skills and that awareness. But it's the perfect example of the software agnostic course, you know, I mean, and they say, what should I use for this? Whatever you want. And some people will use Excel and others will use R and others will use Python and others will use all sorts of other, you know, I think the most I ever had was six different software used in the same course by different people. It was great and they all struggled with different things. I love that. I love the fact that the tool does matter for what you're doing, but it doesn't matter which tool you're using. There's a contradiction in that, but it's really, really powerful. Well, and also to add, you could use multiple tools, you know, Excel has its

strengths for some things, pivot tables, for example, and filtering and things like that, but then actually if you want to get a real kind of power behind your graphics, for me, might go to with MBR or Insta, a kind of art-based tool. Yeah, absolutely. This is exactly right. I like to think of this as a language. If your software is a type of language, people who only speak one language is really hard to learn another, but some people who already speak multiple languages find it, they can more easily pick up other languages because they're used to sort of listening in a different way and listening and picking up. And this is the same with using different tools for digital analysis, which you're used to using multiple tools. It is easier to then pick up and add another tool to your toolbox. That's very powerful. So yeah, I'm really excited. I'm looking forward to teaching that course. I always enjoy it. It's probably going to be one of the last times I get to do this, simply because it's hard to spare the week. I haven't done this in a few years, and I'm doing it this time because John's finishing his PhD, and it's a good time to go into

Orlando and try and tie off things there. But I don't know when I'll next get the chance to do this. I do enjoy it. Definitely. I mean, as we've spoken about again on the podcast before James said, I taught it a couple of years ago, and it was very enjoyable and very interesting. But anyway, so these kind of two variants that will create these two variants, one being this software agnostic version. And that can then actually be by having these current variants in our Jamovie in Python and so forth, we can see, okay, we can work out from this how do we create what would look good in a software agnostic version. And then the other idea is to build a second variant, which is in our insert. And this is a good test to see. Is there something where there is work we need to do on our instinct to be able to cover the material? Are there gaps? Are there things within the content which are not currently prioritised, you know, my guess is the modeling there were the gaps? Yes, yes, I think that that's a pretty safe guess. But in the kind of prepare

side and in the describe side, you know, with graphics, and I'll be surprised if there were gaps. It'll be interesting. I'd love to know if there are gaps. I guess creating that kind of simple book first, and then even expanding in our own way, if we wanted to. I mean, I was talking to Roger who works on our own set as well about this. And he was digging into the data sets and he was saying he's a bit disappointed because a lot of the data sets are quite small that they use in a book. And well, we can use our own data sets. So we don't have to follow the books data sets and we can put in it ones that we enjoy and ones that we feel quite useful. Absolutely. And this is exactly where these open textbooks are so powerful at being able to sort of say, well, we don't need to change the textbook, but we could change or insert an example which actually could then go through. You could have a variant of the Jamofi variant, which uses that data set, which would be really exciting. This is something which these things can go in parallel in ways which are really interesting and exciting that these variants can then live alongside each other

and potentially combine in exciting and interesting ways. That's what I'm hoping can emerge from this process. You've kickstarted this with the amazing work you've done, getting 20 books into these technologies. But I think it's just the beginning of where I hope we would be able to get the community more engaged and involved in collaboratively contributing to these things more. What I love about the example of the learning statistics with our book is that the authors have been very positive about these variants. They refer to them in the book. They're really proud of the fact other people have taken their work and built on it. And that's what we want. The reusability is something which is often neglected within the open community where we really want to have that reusability where it's built in community around this. A whole community could emerge around the book which builds it and builds on it in interesting and exciting ways.

Definitely, the only thing that then came to mind as you were saying that was, it's a big task. This kind of creating a structure of developing different variants is not as big of a deal. But on top of that, we don't want it to become confusing or overwhelming for the receiver. If we have these different hypothetically, let's say that there's this version of the gemovivarion with these examples and the gemovivarion with those examples and then okay, which version are you using? Then that fear that we could create some discourse or just adding another layer of complexity for users. Creating this confusion, say, taking your gemovie example of having the gemovivarion that currently is and then the gemovivarion that then uses these new examples, say, then there's two gemovivarions and then just get confusing for the student, for the person reading it. If you're trying to compare books, if you're kind of talking to each other and you both know that you're using a gemovivarion,

but there's now two gemovivarions. So, which one are you using? Well, but this is where my hope is that confusion is desirable. In the sense that if we get to the stage where we always want uniformity, then we should have removed the diversity. I would love it that there are variants in the future where, for example, you have a variant which is more adapted to biology students or more adapted to psychology students. And then within that, what about Kenyan psychology students versus American psychology students? So, this desire for multiple variants, you could take this much further. And actually, I believe this is something where a single authoring perspective where you have the book is the reality we've lived in for so long. What does it look like when you have many

tailored variants, but where you actually have this element of yes, all these tailored variants, they are equivalent in terms of let's say the content which is covered. The structures and they're sort of all approved, whichever of these variants you're using, if you've gone through it, you should get the same statistical concepts of the same concepts in data science. That's what's so powerful. Then it doesn't matter which variant, and the fact that there are different variants might mean that different students might then engage with multiple variants because they might find that they actually it helps them to see the same concept presented in these different ways. So, all of this is something which is not possible to do unless you get a community behind it. But the bigger your community that's engaging in this process, the more variants you could potentially have and the more people you could be serving in ways which tailor them and hopefully also the more minority edge cases who could be deeply served.

So, you know, what if there was a variant which was very specifically, I don't know, related to pick your favourite narrow field? Marine biology, it's not that narrow, it's quite a big field, but it isn't something which is mainstream in other ways. Imagine you now have your marine biology variants and because of the way the structures work, well, they can do that with Jimovia, they can do that with R, they can do it with Python and so on. So, you still get that potentially these things can overlap and coexist in different ways. This is the dream that you can actually the marine biology community can be served without it then being tied in to let's say the choice of a particular person who's done that tailoring to a particular software choice. They might have chosen R because that's what they're comfortable with, but then their students might want to use Jimovia. Oh, I understand, or Python. This ability that the person creating the marine biology variant is not the person who has to be the expert at understanding all of the

software components, those can exist in parallel. That's the dream and now at the moment the structure is to do that and to maintain that and to build that ecosystem and to enable people to engage in those ways, it doesn't exist. But there has been work within the statistics education community on tailoring these books to specific audiences, on having the books tailored to specific software and these are things which have been shown to add value, but this adds complexity. And that's what I think we can actually dig into and really enable people to lean in to that diversity rather than everybody wanting the same textbook. Everybody having variants of a common textbook is an interesting alternative. Well, definitely. That's a very nice idea, a very kind of different future. And what you have demonstrated is that this is a future which is possible to imagine as a community effort enhanced by AI. And the key hit, and this is part of those

who have listened to lots of other episodes, it is this idea that the narrative, the dominant narrative around AI is about it taking jobs about it, taking over what people do, whereas this alternative narrative specifically around the generative AI possibilities that now exist is that they could and they are enabling us to bring communities together and to work more with communities of people in ways that would have been difficult otherwise. Because the AI can be part of bridging some of these gaps, the gaps of what, as you've already gone through, what was it written in? What was it authored in? And you converted that into a common authoring system using AI so that these things can now be more interoperable with one another. This is the sort of thing we're thinking of AI as being something which is enhancing human collaboration. That's what I

hope we're demonstrating within this, what you've started, a small textbook project, but a very exciting one. And it's not as small when you're in it, sorry. Oh, no, it's a fun little project, though, as you say. The only thing I want to add there about the AI is yes, it's very, very quick to get your, you know, our markdown version into a free text version, convert this RMD file into a PTX file. Okay, very easy to tell the robots to do that. They can enjoy changing the language and the tone. So the bit that's time consuming is checking. Absolutely. Doing it responsibly and being able to check through and so on. And this is where that human effort is still there and that need for the human expertise to be able to have that direction to take responsibility. This is critical. Well, thank you very much. This has been a very exciting conversation. It's great. The progress you're making is absolutely incredible. I really look forward to this moving forward in

interesting ways. And I'm sure we'll keep talking about this in the coming months as the next innovations happen. Yes. Thank you.

More episodes

More from The IDEMS Podcast

View all episodes →