Jargon
A scaling law is a curve fitted to past runs, not a promise about the next one
The observed relationships between computation, data, model size and prediction error are empirical regularities that have held over a wide range, and reading them as a guarantee mistakes a measurement for a mechanism.
By Imran Sheikh3 min read

What was actually observed
Researchers training many models of different sizes on different quantities of data noticed that the resulting prediction error fell in a strikingly regular way as each input increased. Plotted appropriately, the relationship looked like a straight line over several orders of magnitude, which is unusual enough in an empirical science to be worth a great deal of attention.
The practical consequence was enormous. If error can be predicted from the resources spent, then the outcome of a very expensive training run can be estimated from a series of cheap ones, and a budget can be allocated on evidence rather than on hope. Committing to a large run stopped being a gamble.
That’s the substance of the finding: a reliable extrapolation procedure. It says nothing about why the relationship holds, and there is no accepted theoretical account of it, which is a point often skipped when the term is used rhetorically.
The allocation result was the more consequential one
A second line of work asked how a fixed computational budget should be divided between making the model larger and training it on more data. The answer revised prevailing practice: models in common use had been larger than was efficient for the quantity of data they were trained on, and a smaller model trained on more data would do better for the same cost.
This reordered priorities across the field, and it has a consequence people frequently miss. A model that is efficient to train is not necessarily the one you want to deploy, because a smaller model costs less every time it is used. Training efficiency and serving efficiency point in the same direction here, which is a happy accident rather than a principle.
These relationships have been revisited as methods changed, and the specific coefficients are properties of a setup rather than constants of nature. Quoting them as fixed numbers is a misreading of what was published.
The measured quantity is not the interesting quantity
What the curves predict is the loss — a technical measure of how well the model predicts held-out text. What anybody cares about is whether the system can do something useful, and the relationship between the two is loose, indirect and not described by any comparable regularity.
Improvements in loss have generally accompanied improvements in capability, which is the empirical basis for treating the curves as meaningful. But the mapping is not smooth: performance on a specific task can be flat while loss falls steadily, then change quickly. Whether such changes are genuine discontinuities or artefacts of how the task is scored is disputed.
This gap is where most of the overreach happens. A well-supported claim about a technical error measure gets restated as a claim about intelligence, and the extrapolation that was justified for the first is carried over to the second without justification.
Extrapolating a fitted curve is a distinct claim
Every such relationship was fitted over a range that had been tried. Assuming it continues beyond that range is an assumption, and empirical regularities in other fields have a long record of holding until they do not. Nothing in the data can establish that a curve continues past the last point on it.
There are also arguments about approaching limits — of available text, of manufacturable hardware, of power supply — each of which would bend the curve for reasons having nothing to do with the relationship itself. These are contested in their timing and not in their existence.
The counter-position deserves stating: the regularity has held across a remarkably wide range and through several changes of method, which is more than most empirical relationships manage. Dismissing it as mere curve-fitting understates how much it has already survived.
How to read the term when you meet it
Treat it as a description of what has been measured and a planning tool, which is what it was proposed as. Treat any use of it to predict future capability as an argument with an unstated premise, and look for whether the premise is acknowledged.
Be alert to the substitution of loss for capability, which is where a defensible statement usually becomes an indefensible one. And note that specific numbers attached to the term are setup-dependent and age quickly.
The word law is the problem. These are regularities, in the sense that an economic relationship is a regularity, and the vocabulary of physics attached to them conveys a certainty that nobody claiming it would defend if asked directly.
Common questions
Do these relationships mean bigger is always better?
They describe returns that continue while diminishing, in a specific technical measure, over the range tested. Whether continuing to spend is worthwhile depends on what the improvement in that measure is worth to you, which is an economic question rather than a scientific one.
Have the relationships broken down?
No breakdown has been established publicly, and the most capable current systems are not documented in enough detail for anyone outside to check. Claims in either direction rest on limited public evidence.
Why is there no theory explaining them?
Several partial explanations have been proposed, drawing on the structure of language and on how models fit data of varying difficulty, and none has become generally accepted. The regularity is better established than any account of it, which is not unusual early in the study of a phenomenon.
Editor, AI Worth Knowing
Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.





