Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Training a model on another model’s output is a real technique with a known trap

Manufactured training material solves problems that collecting more real material cannot, and it carries a distinct failure that only becomes visible after several rounds of it.

By Zoya Rahman4 min read

Professional woman working on laptop in a server room, showcasing technology and remote work.
Photograph by Christina Morillo via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Some of the data you need was never going to exist

The usual picture of training has a system fed on text and images that people made for their own reasons and that somebody later gathered up. For a great many tasks that works, and for a stubborn minority it does not, because the examples required were never produced in the first place. Nobody writes down the ordinary.

Consider a system meant to recognise a rare mechanical fault, or one meant to handle a request phrased in a way that only a handful of people ever use. The events are real, they matter, and there is almost no record of them. Collecting harder does not help when the scarcity is a property of the world rather than of your collection effort.

Manufactured data is the response. It covers everything from a photograph rotated and re-lit to widen a small set, through scenes rendered in a simulator, to text written by one model in order to train another. These are quite different practices wearing one label, and the differences decide whether the result is sound.

What manufacturing buys, when it works

The first thing it buys is coverage. If you can generate examples deliberately, you can generate them where you need them — at the edges, in the awkward combinations, in the cases your real archive happens to be thin on. Real data arrives in whatever proportions the world supplied, and those proportions are almost never the ones a trainer would choose.

The second is labels that come free and come correct. When a scene is rendered rather than photographed, the system that drew it already knows where every object is, how far away it sits and what it is called. That removes the annotation step entirely, along with the disagreement and expense that step normally carries with it.

The third is a way past a permission problem. Material that cannot be used for legal or ethical reasons can sometimes be replaced with manufactured material carrying the same statistical shape and none of the same people. Whether that replacement is genuinely clean is contested, since a generator trained on restricted material may carry traces of it forward.

The trap is a narrowing, and it compounds

A model trained on data produced by another model inherits that model’s distribution rather than the world’s. Whatever the parent under-represented, the child sees less of. Whatever the parent found unlikely, the child may not see at all. Run this arrangement for a generation or two and the tails of the distribution thin out.

The mechanism is not mysterious. Generation samples from a learned distribution, sampling is biased towards the centre of that distribution by construction, and training on the samples re-fits the centre more tightly. The result looks fluent and increasingly average, with the unusual cases quietly absent rather than visibly wrong.

Researchers have demonstrated this degradation in controlled settings, and the concern that it might happen at scale across the open web — where generated text now sits alongside written text, unlabelled — is genuine. It is also a projection rather than an observation, and worth marking as such.

Why the collapse is avoidable rather than inevitable

The degradation depends on a closed loop with nothing entering it. Break the loop and the arithmetic changes. Keeping real data in the mixture, rather than replacing it, is the most obvious remedy and appears to be effective; the manufactured material then widens coverage without dictating the shape of the whole.

Filtering matters at least as much. If generated examples are checked against something external — a solver that can verify an answer, a simulator that obeys physics, a person who rejects the bad ones — then the material entering training is better than the material leaving the generator. Verification, not generation, is doing the work in those setups.

This is why manufactured data has been most convincing in domains with a cheap and honest test. Where an answer can be confirmed, generation becomes a search for good examples rather than a copy of an existing distribution, and the failure described above does not get started.

Reading a claim about it sensibly

The useful question is never whether a system used manufactured data, since almost everything does in some form. It is which kind, in what proportion, and against what check. A claim that omits all three is not telling you much, and the word alone covers practices with very different risk profiles.

It is also worth separating two arguments that frequently get muddled. One is technical, about whether training on generated material degrades a model. The other is about provenance and permission, and would remain live even if the technical problem were solved tomorrow. They have different answers.

The honest summary is that this is a working technique with a documented failure mode and known mitigations, not a shortcut and not a dead end. Most of the disagreement in the field is about proportion rather than principle.

Common questions

Is augmenting real photographs the same thing as generating them?

Not really, though both are called synthetic. Rotating, cropping or re-lighting a real image keeps the underlying content and teaches a model that certain changes should not matter. Generating an image from scratch produces content that was never observed, which is a much stronger claim about what the world contains.

Does manufactured data solve the problem of running out of text?

It is sometimes offered that way and the argument is weak on its own terms, because generated text carries no information the generator did not already have. Where it genuinely helps is in restructuring existing knowledge into forms that train better, which is a narrower and more defensible claim.

Can you tell from the outside whether a model was trained this way?

Generally no. Training composition is rarely disclosed in detail, and the artefacts of manufactured data are not reliably detectable in a finished model. This is one reason the debate proceeds largely on argument rather than on measurement.

How It Workssynthetic datatrainingdata
Zoya Rahman
Deputy editor, AI Worth Knowing

Zoya joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and would rather show the working than assert the conclusion.