Jane Austen Mini-LLM Teaches Us About AI Data Bias

A thought experiment imagines training a small language model only on Jane Austen's writing to explain, in plain terms, why AI output is statistical prediction shaped entirely by its training data, biases included.

10:33 video6 min readWatch on YouTube

Imagine training a language model on nothing but Jane Austen's collected works, no internet, no news, no modern slang, just her novels and letters. What would it write if you asked it to continue a story? This thought experiment is the engine behind a plain-language explanation of how large language models actually work, and why every one of them carries the fingerprints of whatever it was trained on.

A closed world, not a lookup table

The premise is deliberately narrow: a small language model trained exclusively on Jane Austen's body of work, a closed world by design. The first and most important claim is what such a model would not do. It would not memorize and regurgitate her actual sentences the way a database lookup would. Instead, it would learn her patterns of writing statistically, which words tend to appear near which other words, what phrasing and rhythm recur across her sentences. Ask it to write a story, and it predicts what Austen might plausibly have written based on the statistical structure of her existing corpus, not by retrieving something she actually wrote. That distinction, prediction versus memorization, is the single idea the entire video is built around, and it gets restated from several angles because it is easy to misunderstand.

What happens when you ask it something outside its world

The clearest way to test a closed-world model is to ask it something it has no basis to answer. A model trained purely on Austen's writing would have no idea what a laptop is, not because it is unintelligent, but because "laptop" simply does not exist anywhere in its training data. That is not a flaw in the model's reasoning; it is a direct consequence of what it was and was not exposed to. Every training dataset has this property. A model can only draw on the statistical patterns present in what it learned from, and anything outside that scope is, functionally, outside its world entirely.

The Wonder Woman example

To make the abstraction concrete, the video turns to an image generation example. Typing "a photo of Wonder Woman" into Midjourney reliably produces an image resembling Gal Gadot, the actress who played the character in recent films. It does not produce Lynda Carter, who played Wonder Woman decades earlier. That is not because Gal Gadot is objectively "more" Wonder Woman. It is because Midjourney's training data associates the Wonder Woman prompt more strongly with recent film imagery than with the earlier television portrayal, and the model's output reflects that imbalance directly. Ask for a drawing instead of a photo, and the output tends to shift toward a comic-book rendering, again following whatever associations dominate the training data for that particular phrasing. The term "statistical parrots" comes up here as a description some people use for these models, with the acknowledgment that the label undersells how genuinely impressive the prediction quality actually is, even while remaining accurate about the underlying mechanism.

Bias is built in, not bolted on

The throughline connecting the Austen thought experiment to the Wonder Woman example is that bias is not a bug introduced by carelessness. It is a structural property of any model trained on a finite, non-neutral dataset. A model trained only on Austen inherits her world's vocabulary and assumptions. A model trained on Midjourney's image corpus inherits whichever visual representations were most common in that corpus. Testing a model's grasp of a narrow domain, like asking whether Austen-generated text actually reads like Austen, is described as a job for genuine domain experts: people who know her full body of work well enough to say whether generated text captures her voice or misses it, which is essentially a domain-specific version of a Turing test.

Restricting a large model to a narrow world

An interesting wrinkle comes when the thought experiment is flipped: what happens if you take a full-scale large language model, trained on internet-wide data, and simply instruct it to restrict itself to Jane Austen's world? You can try this today in Claude, ChatGPT, or another chat assistant by asking it to answer only within Austen's universe. The model will do its best, but the video is candid about the limitation: things will creep in that do not fully belong, because a model trained on the entire internet cannot truly become a closed-world Austen model just by being asked to pretend. It might use vocabulary or ideas from far outside her era, because that knowledge is still sitting in its training data underneath the roleplay instruction. The only way to build a model that would never produce an anachronism like "laptop" is to actually train it, from the ground up, only on data from Austen's own time.

Personal bias, layered on top of model bias

The video is transparent about a second, more personal layer of bias sitting on top of the model's own: images generated of "what Jane Austen might have looked like" reflect not just Midjourney's training biases toward conventionally attractive, fashion-forward portrayals, but also the presenter's own choices in prompting, since nobody actually knows exactly what Austen looked like beyond a small number of sketches. Choosing a different image model trained on different data would produce a different look entirely, which is offered as a preview of a larger topic, how to choose the right model for a given task, that the video does not have time to fully unpack.

Why this matters for how you use AI at work

The practical payoff, and the reason this lesson sits inside an upskilling series rather than a purely academic one, is a workplace rule: the output of any language model is only as good as the instructions, meaning the prompt, that shape it, combined with the boundaries of what it was trained on. Understanding that a model predicts rather than recalls, and that its predictions are only as unbiased as its training data allows, is presented as the foundation for using AI tools like ChatGPT, Gemini, or Claude effectively as a high-speed assistant for drafts, data work, and routine tasks, rather than treating its output as neutral fact.

Key takeaways

  • A language model trained only on a narrow dataset, like Jane Austen's works, learns statistical patterns in that data rather than memorizing and repeating exact text.
  • Asking such a model something outside its training scope, like what a laptop is, produces no meaningful answer, because that knowledge was never present in its data.
  • Midjourney's tendency to render Wonder Woman as Gal Gadot rather than Lynda Carter illustrates how training data imbalance directly shapes model output.
  • All training data carries bias; it is a structural feature of any finite dataset, not a fixable defect.
  • Instructing a large, internet-trained model to restrict itself to a narrow world, like Austen's, only approximates a true closed-world model, since outside knowledge can still leak through.
  • Evaluating whether generated text authentically captures a specific style, like Austen's, requires genuine domain experts, functioning as a specialized version of a Turing test.

Who this is for

This lesson is part of the Humanitarians AI upskilling series aimed at helping people move past AI jargon and understand, at a practical level, how these models actually behave. It is built for anyone using AI tools like ChatGPT, Gemini, or Claude at work who wants to write better prompts by first understanding why a model's output is shaped entirely by prediction and training data rather than by lookup or true understanding.

Full transcript(auto-generated, with timestamps)

[0:00]Okay, we're going to understand large language models starting with a small language phone. So, we're going to just do a thought experiment. Uh I might do some simulations of this, but the thought experiment is going to be what if we just trained it not on the entire internet the way large language models are, but just on Jane Austin's writings or corpus or body of work. What will the output of that model be? So, this rest of this video is going to discuss that. We're going to try a little thought experiment. We're going to imagine that we trained a little language model only on Jane Austin's

[0:36]Works. And you could do this if you want. We may set up a lab doing this and see the results of it. But for now, we're just going to treat this as a thought experiment. What would this model output? So our parameters are this is a closed world. This is only Jane Austin's world. And so what it'll do is it won't memorize and regurgitate Jane Austin text. What it'll do is it'll learn her patterns of writing statistically. It will learn the words she uses, what words are related to other words. And if you ask it a question like write a story about something, it will sort of predict what she may

[1:20]Have written based on the statistics of you know her writing her corpus her works. This is not memorization. This is a statistical mapping. It looks at covariance and looks at patterns within her data in her data. It makes predictions. If you ask it something out of its scope, like what is a laptop, it doesn't know. It's not its world. It's not its training thing. It's not because it's stupid. It's because that's not its training data. And all training data has bias. If I go to Medjourney and I type in, you know, a photo of Wonder Woman, it shows me Galadub. Uh there is a Wonder Woman before Gail

[2:04]Gadub called Linda Carter. Doesn't show me Linda Carter. It shows me Gail GDO. What that means is Met Journey is trained on Gail Gdote, you know, uh, films because that's what it shows. It maps sort of Wonder Woman to Gail Gdote or if you do a, you know, a drawing of Wonder Woman, it will probably do a TC comic version of that. So, what these life models do is they predict that. They're just, some people call them statistical parrots. That may be demeaning a little bit about how amazing they actually are. It's remarkable what they do given how they work. They are making predictions. They

[2:45]Make very good predictions. It's it's sort of stunning how good they are at what they do given how they work because they're just making predictions. But they get better and better at making the predictions. But because they make predictions, they can make mistakes. They can predict things wrong. And we're going to get into that quite a bit. We're going to get into how do we know it did a good job. Sometimes is qualitative. Like here I made some images and we'll do this down the road. We'll do image prompting for midjourney image fronting for V3 for books for scientific research etc. This is what I imagine

[3:26]Jane Austin to look like. It probably looks nothing like Jane Austin maybe. uh has her I read she had I forgot what the thing called the thing on her head that is something I read uh I forgot what it's called language model would know what it's called she would know what it's called too because it's in her world in her world if he trained it on her world her world would not know what all laptop is so what these things are doing and they're amazing is there are pom ballistic machines if trained on Jane Austin data It gets sort of the probabilistic essence, the coariance of what Jane

[4:05]Austin might have written. There's not identity. We're not looking up and following her fact. We're sort of generating a novel thing that she may have written. And they're remarkably good at this. They're remarkably good at I'm not a super expert in Jane Austin. The way you would evaluate that can it write like Jane Austin is you get Jane Austin experts. You would generate text and ask them, well, does the same feel like Jane Austin? If they're real experts, they would know everything she wrote. So, they know that she did not write this. But if you sort of knew Jane Austin and knew her feel, a test would be could you

[4:44]Tell sort of generated Jane Austin from real and that's a turn test, which we'll talk about uh a bit, many turn tests. We'll get into that in this video. What the essence of this video is is these are statistical machines. They learn patterns. They don't do lookup. They don't do limmerization. And the our next video we're going to talk about okay this is just the Jane Austin world. What if we take them to the internet world a scale the entire internet and that's what these quote unquote large language models are. So when we use the word large we're talking about big big data like internet scale

[5:22]Data data which it cost you know 10 million 20 million to train just from the computing power to train the models but you could you could train your own Jane Austin models I don't know how much she's written but in computer world she hasn't written that much compared to what a computer can handle even if she written a thousand butts which I don't think so I think it's maybe like 10 text of whatever. I don't know. I'm not an expert in that. Um, you could definitely train a model on Jane Austin's corpus. We may do that as a experiment in this class. Technically, a project you could

[5:59]Do if you're a student in this class and it would be limited to the Jane Austin world. I certain that she'd never use the word laptop or hiphop or, you know, whatever or, you know, stuff even out of her world in England. Um, and so that's not in her world, but her patterns of writing, it's going to learn, it's going to learn her style of writing from the statistical coariance of the patterns. And the fake language models, the large language models, we'll do this at an internet wide scale. It'd be interesting. You could ask a large language model to restrict yourself just to Jane Austin's world. it will try its

[6:41]Best that it'll probably creep in some stuff that wasn't in her world. You couldn't really get a large angler model to truly be her world unless you only trained it on data from her time. You right now you can go to Claw Jedi Chat BT and say restrict yourself only to Jane Austin's universe and answer these questions. It'll do its best. It'll try, but things will creep in that it doesn't fully understand because it's just looking at statistical patterns and it may throw in the word laptop. It might have Jane Austin use the word laptop. If you trained it purely only on Jane Austin data from scratch, it just

[7:25]Doesn't know that exists. So, it's not going to predict that because it's not in its training data. So in our next one we're going to get into the next step scale large language models with the same idea feature in a data off fires and it'll have different kinds of biases. Uh I have another class where we talk about how to detect bias and what to do about bias. My bias here is this is my bias of what I think that's Emma her her most famous character and my bias of what I think Jane Austin looks like. They could be and probably is way way off.

[8:07]But um it's also my journey of bias because I use medjourney to generate these images. It clearly is trained on the left fashion images. I just had a critique of somebody who's signing, you know, Jane Austin looks like a supermodel. You know, she didn't look like a supermodel. I don't know anybody really knows exactly how she look because there's only sketches of drawings. Um, but it's true. The the midjourney bias is to make everybody look like they're in love. Use a different model, train on different data, it will have a different look in a different field. We'll get into that because choosing the right language

[8:42]Model for the advanced. This could be part of the prompt engineering experience, but we won't get into that. Now, I'm already over time. I think I'm over eight minutes and I want to keep these to five. Certainly less than 10. So, the major point of this is we're not memorizing. It's not if we train it on Jane Austin data, it would not just regurgitate and spit out like a database exact Jane Austin quotes. It would learn statistical patterns within Jane Austin's work and make a prediction of what the next word would be given her previous history of writing. And that same model is going to go to

[9:26]Basically large language models. It's going to do the same thing but a much greater scale. So, it's getting the patterns and rhythms. It's learning it statistically, the patterns and rhythms of of Jane Austin, just like it learns the patterns and rhythms of the internet. But those patterns are much more complicated because they're the internet. They're not just one writer. Uh, and so if you wanted to truly do a Jane Austin generator, you should probably limit it to Jane Austin data, which you can do. It's not that much. about the idea again the idea of this video is not memorization. It's not just looking up and spitting out something.

[10:04]It's looking at patterns then generating predicting novel sequences based on the patterns that it's learned. Okay, that's it. We're almost at 10 minutes. I think at 10 minutes right now. So I'm just going to cut this. The next one we're going to talk now about the next scale large language models. It should do the same thing but now at scale. Okay, take care.

More videos

Humanitarians AI Lyrical Literacy Project