How It Works
A training corpus is a collection with a history, not a sample of language
The text a model is built from was gathered, filtered and arranged by people making choices, and almost every property of the finished system traces back to those choices rather than to the architecture.
By Daniel Okonkwo4 min read

The material has to come from somewhere
Discussion of these systems dwells on architecture, on parameter counts and on the clever mechanism at the centre, and it skips lightly over the question of what was actually fed in. That is the wrong emphasis. Two models built to identical designs and trained on different collections of text will behave differently in ways that no amount of tuning afterwards fully corrects.
The bulk of the text in a general-purpose language model comes from the public web, gathered by crawlers that follow links and store what they find. Alongside that sit books, code repositories, reference works, transcripts, forum archives and whatever else a builder has managed to license or lawfully obtain. The proportions vary between projects and are frequently not disclosed.
None of this is exotic. It is the ordinary material of the internet, collected in bulk. But the ordinariness is the point: a system trained this way inherits whatever the internet happened to contain about a subject, including the parts written carelessly, the parts written to sell something, and the parts nobody has updated in a decade.
A crawl is not a sample in any statistical sense
A statistician drawing a sample of human writing would define the population first and then select from it in a way that made the sample representative. A web crawl does nothing of the kind. It collects what is reachable, what is not blocked, what sits in a format the pipeline can read, and what happens to exist in large enough quantity to survive the later filtering stages.
That produces predictable distortions. Subjects that generate a lot of online text are enormously over-represented relative to their importance, and subjects that people mostly discuss out loud rather than in writing are nearly absent. Languages with a large web presence supply orders of magnitude more material than languages with a small one, and the gap between them is not proportional to how many people speak each.
Time is skewed too. The web is not an even record of the past; it leans heavily towards the recent, towards whatever has been reposted often, and towards content that survived because somebody kept paying for the hosting. A model reflects that weighting whether or not anyone intended it to.
Cleaning is an editorial act carried out at industrial scale
Raw crawled text is largely unusable. It is full of navigation menus, cookie notices, machine-generated filler, duplicated pages, markup and spam. Before training begins a pipeline strips, deduplicates and filters, and every one of those steps embeds a judgement about what counts as good text.
Those judgements are rarely visible from outside and they are not neutral. A quality filter that scores documents by their resemblance to reference material will systematically favour prose written in a particular register — formal, edited, standardised — and quietly discard dialects, informal writing and the output of communities that write differently. The intent is to remove junk. The effect includes removing variety.
There is no way to avoid this. Some filtering is essential, because training on unfiltered web text produces a worse model on every measure anyone has proposed. But it is worth naming the activity honestly: this is editing, performed by rules, on a body of text no editor could ever read.
Not all of it was scraped, and that changes the argument
Alongside crawled material, builders increasingly use text they have paid for, text generated by other models, and text produced deliberately by contracted writers for the purpose. The last category matters more than its volume suggests, because it is targeted: written to cover a gap, to demonstrate a format, or to supply examples of a behaviour the general corpus does not contain enough of.
Training partly on machine-generated text is a genuinely contested practice. One camp holds that carefully filtered synthetic data is simply another form of curation and works well in domains where correctness can be checked automatically. Another argues that a model trained largely on the output of models risks narrowing towards its own tendencies, losing the tails of the distribution that made the original data valuable.
The evidence so far is mixed and highly dependent on how the synthetic data was produced and screened. Anyone stating confidently that it either ruins models or solves the data problem is running ahead of what is actually known.
Provenance is the argument that will not resolve quietly
Because the material was gathered rather than commissioned, questions of consent, compensation and legality attach to it, and those questions are being worked out in courts and legislatures in several jurisdictions at once. The outcomes differ by country and the reasoning differs even more, so any summary written here would age badly.
What can be said mechanically is that the disputes have consequences for how systems get built. If access to certain kinds of text becomes conditional on licensing, then the ability to train a competitive general model becomes partly a matter of who can strike those deals, which is a different bottleneck from having the hardware.
The deeper point holds regardless of how the law lands. A model is a compressed reflection of a specific collection assembled at a specific moment by particular people under particular constraints. Treating its output as a view from nowhere misreads what it is.
Common questions
Can you find out what a specific model was trained on?
Usually not in detail. Some projects publish the composition of their corpus and a few release the corpus itself, but most disclose only broad categories, and commercial and legal considerations push towards saying less rather than more. Independent researchers infer what they can from model behaviour, which is indirect evidence at best.
Does more data always help?
Only up to a point, and only if the additional data is not redundant with what is already there. Beyond a certain scale the composition and cleanliness of the collection matters more than its raw size, which is why so much engineering effort now goes into filtering rather than into gathering.
Is a document in the training data stored inside the model?
Not as a document. The text influences the parameter values during training and is then set aside, so the model holds a lossy statistical residue rather than a copy. That said, text repeated many times across a corpus can be reproduced closely, which is a real phenomenon with real implications and not merely a theoretical possibility.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





