Jargon
Overfitting is memorisation wearing the costume of success
A model that scores perfectly on the data it was trained on has told you almost nothing, and the concept explaining why is the oldest genuinely useful idea in the field.
By Samar Bhatia3 min read

Fitting the noise as well as the signal
Any dataset contains two things: the underlying pattern you want, and the accidents of that particular sample. Measurement error, unusual cases, coincidences and whatever happened to be true of the specific examples collected. A model with enough flexibility will fit both, because the training procedure cannot distinguish between them.
Overfitting is the name for having fitted too much of the second kind. The result performs superbly on the data it was trained on and poorly on anything new, because the accidents it learned do not recur. It has learned the sample rather than the phenomenon.
The classic illustration is drawing a curve through scattered points. A straight line misses much of the detail; a curve wiggling through every single point captures them all perfectly and predicts the next point worse than the straight line does. Both errors have names and the second one is this.
The only way to detect it is to hold data back
Because an overfitted model looks excellent on its training data, performance on that data is worthless as an indicator. The standard practice is therefore to split the available data before training and evaluate on a portion the model never saw, which is the single most important methodological convention in the field.
In careful work there are three portions rather than two. One for training, one for the many decisions made during development, and a third touched only at the very end, because a set used repeatedly to guide choices becomes contaminated by those choices even though the model never trained on it directly.
The characteristic signature is easy to spot once you look for it: error on the training data continues to fall while error on held-out data flattens and then begins to rise. That divergence is the whole of the concept in one picture, and stopping training when it appears is a legitimate and widely used technique.
The remedies all amount to restricting freedom
Every mitigation works by making it harder for the model to fit noise. Using fewer parameters is the direct approach. Penalising large weights during training discourages the extreme values that sharp wiggles require. Randomly disabling parts of the network during training prevents any single path from becoming essential.
Adding data is the most effective remedy of all, and the least available. With more examples, accidents of the sample average out and the underlying pattern becomes proportionally stronger, so the same flexible model overfits less. Artificially expanding a dataset by transforming existing examples is a widely used approximation of this.
These are collectively called regularisation, and every one of them trades some capacity to fit real structure for reduced capacity to fit noise. The trade is unavoidable, and where to set it is judgement informed by measurement rather than a calculation.
Large models complicated the tidy story
Classical theory predicts a clean trade-off: as a model gains capacity, held-out error falls to a minimum and then rises as overfitting takes over. Very large networks were observed to break this, with error falling, rising, and then falling again as capacity increased far past the point where the model could memorise the training data entirely.
This behaviour has been reproduced across many settings and it is not seriously in dispute. Explaining it is another matter. Competing accounts appeal to properties of the training procedure, to structure in real datasets, and to the geometry of very high-dimensional spaces, and no explanation commands consensus.
The practical consequence is that parameter count alone no longer predicts whether a model will overfit, and that the classical intuition can mislead in the regime where modern systems operate. The underlying concept has not been overturned. The rule of thumb attached to it has.
Why the term keeps escaping the field
Overfitting has become a useful metaphor outside machine learning, and for once the metaphor is reasonably faithful. A policy tuned precisely to the last crisis, a strategy built from a small number of vivid cases, a hiring process calibrated to the people who happened to work out before — all of these fit their sample and generalise badly.
The transferable insight is that performance measured on the same information used to build something is not evidence. It is a description. Evidence requires exposure to a case the construction did not get to see, and arranging that is inconvenient in exactly the situations where it matters most.
That is worth carrying away even if none of the rest of this is ever relevant. The field’s most durable contribution to general reasoning may turn out to be the insistence on holding some data back.
Common questions
What is underfitting?
The opposite failure: a model too constrained to capture the real pattern, so it performs poorly on both training and held-out data. It is easier to detect, because the training error itself stays high, and it is generally the less dangerous of the two since nothing about it looks like success.
Can a model overfit a test set?
Indirectly and commonly. If results on a test set guide which model to keep, which settings to use and when to stop, information from that set leaks into the final system across many small decisions. This is why a genuinely untouched final evaluation set is standard practice in careful work.
Do large language models memorise their training data?
Some of it, demonstrably. Passages that appear many times or are highly distinctive can be reproduced, and this has been shown by direct experiment. How much is memorised versus generalised is an active research question with significant legal and privacy implications, and the answer varies with model size and data duplication.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





