Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Past a certain scale more text stops helping and cleaner text starts

The field spent years assuming the answer to every shortcoming was a bigger pile of data, and the evidence has been quietly pointing at composition instead.

By Naina Sethi3 min read

An IT professional operates a computer in a server room, managing network systems and connected devices.
Photograph by panumas nikhomkhai via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The volume story was true, and then it became less true

For a stretch of this field’s recent history the reliable way to improve a model was to make everything bigger — more parameters, more computation, more text to feed the process. That worked well enough and often enough to harden into an assumption, and the assumption outlived a good deal of the evidence for it.

What has emerged since is more awkward. Returns from additional data taper, and they taper faster when the additional data resembles what was already present. A hundred thousand near-identical product pages do not teach a model a hundred thousand things. They teach it one thing, emphatically, and they consume training budget doing it.

This does not mean scale stopped mattering. It means the interesting variable moved. Between two collections of the same size the difference in the resulting model can be substantial, and that difference comes from composition rather than from volume.

Duplication does something worse than waste a slot

Web text is enormously repetitive. Pages are syndicated, quoted, mirrored, scraped and republished, and boilerplate appears on millions of pages sharing a template. Before deduplication a naive crawl contains the same passages an absurd number of times.

The harm is not only inefficiency. A passage seen many times exerts an outsized pull on the parameters, so the model becomes disproportionately confident about it and more likely to reproduce it verbatim. How readily a model regurgitates an exact string correlates strongly with how often that string appeared, which makes deduplication both a quality measure and a mitigation.

Deduplication is fiddlier than it sounds, though. Exact matching catches only the easy cases; near-duplicates require fuzzy comparison across an enormous collection, and the threshold chosen decides whether you are removing genuine repetition or starting to delete legitimately similar documents.

Filtering for quality means choosing a definition of quality

The usual approach is to train a small classifier to distinguish text resembling some reference set — an encyclopedia, a curated shelf of books — from ordinary crawled pages, then keep the documents scoring highly. It is efficient and it demonstrably improves the resulting model on standard evaluations.

It also imposes the reference set’s idea of good writing across the entire corpus. Text in a non-standard dialect, informal registers, communities whose conventions differ from edited prose, and anything unusual enough to look anomalous will score lower and be thinned out. The model that results is better at sounding like the reference material and worse at everything the filter treated as noise.

Nobody has a clean solution. The trade is real: less filtering yields a messier, more varied model that performs worse on the tests everyone uses for comparison. That those tests reward the same register as the filter is a circularity the field discusses less often than it might.

Order and proportion turn out to matter as well

Data is not simply poured in. The proportions of the mixture get chosen — how much code, how much reference material, how much conversational text — and those proportions have effects reaching well beyond the obvious. Including a substantial quantity of source code, for instance, has been repeatedly associated with improvements on tasks involving structured multi-step reasoning rather than programming.

Whether that is because code enforces explicit sequential structure, because it is unusually well-formed, or because of some correlation nobody has isolated, is not settled. The observation is fairly robust; the explanation is not, and confident causal accounts of it should be read as hypotheses.

The sequence in which material is presented can matter too, particularly in later stages where a smaller quantity of carefully chosen material shapes behaviour. That final stage is not where a model learns most of what it knows, but it is where a great deal of what people notice about it gets set.

Where this leaves the argument about running out

A recurring claim holds that high-quality text is a finite resource approaching exhaustion. It is a real concern and it is also frequently overstated, because the estimates depend entirely on what gets counted as usable, and that definition has shifted repeatedly as filtering methods improved.

There are several plausible responses and none is proven. Extracting more from the same data through better training procedures; drawing on non-text sources; generating material synthetically under verification; or simply accepting slower progress. Which matters most is exactly the sort of question where forecasts have a poor track record, so treat any confident answer, including a pessimistic one, as speculation.

The durable point is narrower and safer. Data is not an undifferentiated commodity, the pile is not the product, and deciding what belongs in a corpus is closer to editing than to collecting.

Common questions

Would training on the entire web produce the best model?

No, and this has been tested repeatedly. Unfiltered web text carries so much spam, boilerplate and machine-generated filler that models trained on it perform worse than models trained on a smaller filtered subset. Some filtering is not optional.

Does deduplication remove things a reader would want?

It can. Aggressive near-duplicate removal deletes documents that are similar for legitimate reasons, such as multiple accounts of the same event, or standard legal and technical text that is repetitive by nature. Where the threshold sits is a judgement call, and it is made once for the whole corpus.

Is a small, carefully chosen corpus enough?

For a general-purpose model, no. Curation improves what you have but cannot substitute for breadth, because a model can only be reliable about material it encountered. The finding is that quality matters more than the totals suggest, not that scale is irrelevant.

How It Worksdata qualitydeduplicationfilteringscaling
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.