Jargon
Distillation trains a small model to imitate a larger one’s behaviour
The technique produces compact models that punch well above their size, and it works for a reason that says something interesting about what a large model actually contains.
By Manish Trivedi3 min read

A student fitted to a teacher rather than to the world
The ordinary way to build a model is to train it on data with correct answers attached. Distillation changes the target: a smaller model is trained to reproduce the outputs of a larger one, using the larger model’s responses rather than the original labels.
The vocabulary is teacher and student, and it is more apt than most metaphors in this field. The student is not learning the task from first principles. It is learning to behave the way something more capable behaves, on whatever inputs it is shown.
The result is frequently a small model that performs far better than a model of the same size trained conventionally. That gap is the whole reason the technique exists, and explaining it took a while. It also implies something about large models that is easy to miss: a great deal of their capability can be transferred into a much smaller container, which suggests the original size was needed to find the behaviour rather than to hold it.
The extra information is in the shape of the answer
A correct label carries one bit of guidance: this is the answer. A teacher model’s output carries a full distribution across possible answers, including which wrong answers it considered nearly right and by how much.
That structure is enormously informative. A teacher that assigns a small but non-trivial probability to a related but incorrect option is expressing something about similarity that no single label contains, and a student fitted to the whole distribution absorbs that structure rather than just the winner.
This also explains why distillation can outperform training on the original data. The teacher has already done the work of smoothing away the noise and idiosyncrasy in the raw labels, so the student learns from a cleaner and more informative signal than the one the teacher had.
The same word now covers a looser practice
The current usage has broadened considerably. Generating a large volume of responses from a capable model and training another model on that text is also called distillation, even though it uses only the produced text rather than the underlying distributions.
This is closer to learning from a curriculum written by the teacher than to the original technique, and it works well enough to be widespread. Its dependence on the teacher’s quality is total, though, and the student inherits the teacher’s errors, biases and characteristic phrasings without any mechanism for noticing them.
There is a compounding concern where models are trained on the output of models that were themselves trained this way. Whether that meaningfully degrades quality over successive generations is actively studied and not settled, and the answer appears to depend heavily on how much original data remains in the mixture.
What it does not transfer
A distilled model imitates behaviour on the kinds of input it was distilled over. Outside that range, the imitation has nothing to go on, and the student can diverge from the teacher sharply in ways that are hard to anticipate.
It also cannot exceed its own capacity. A small model has fewer parameters and therefore less room, so distillation moves it towards the teacher’s behaviour rather than reproducing it. The gap that remains tends to be concentrated in exactly the demanding cases where capacity matters.
And it copies behaviour, not the reasons for it. Where the teacher is reliable for principled reasons, the student may be reliable only where it happened to be shown, which is a difference that ordinary evaluation on similar material will not reveal.
The commercial and legal argument around it
Because a capable model can be accessed through an interface, its outputs can be collected in quantity and used to train a competitor. Terms of service for commercial systems commonly forbid this, and the enforceability of such terms is being tested in various places with no settled answer.
Detecting it is difficult. A distilled model may retain characteristic phrasings, refusal patterns or errors that suggest its origin, and none of that constitutes proof, since models trained on overlapping public data converge on similar behaviour anyway.
This sits uncomfortably alongside the industry’s own position on training data gathered from the open web. Whether the two situations are meaningfully different is exactly the argument being had, and it is one where the parties have obvious interests and the principles are genuinely unclear.
Common questions
Is a distilled model just a compressed version of the teacher?
Not in the sense that quantisation is. Quantisation keeps the same network and represents it more coarsely, while distillation builds a different and usually smaller network from scratch and trains it to behave similarly. The internals need not resemble the teacher at all.
Why not train the small model on the original data instead?
You can, and it usually performs worse. The teacher’s outputs carry more information per example than a bare label, and they have already filtered out much of the noise in the raw data. That richer signal is what lets a small student outperform its conventionally trained equivalent.
Does distillation copy the teacher’s flaws?
Yes, fairly reliably. Biases, characteristic errors and refusal behaviour all transfer, because they are part of the behaviour being imitated. A student is unlikely to be safer or more accurate than its teacher except where deliberate filtering was applied to the training material.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





