Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Parameters are not knowledge and the count is not a score

The number quoted alongside every model describes how many adjustable values it contains, which is a measure of capacity rather than of capability and has been treated as the opposite for years.

By Daniel Okonkwo3 min read

A detailed close-up of a Bible page showing religious scripture in focus.
Photograph by Brett Jordan via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What one parameter actually is

A parameter is a single number stored inside a model that gets adjusted during training. Most of them are weights, describing how strongly one internal value influences another. A smaller share are biases, which shift a result up or down regardless of the input. That is the whole inventory.

The word comes from statistics, where a parameter is a quantity of a model that is estimated from data rather than fixed in advance. The usage here is the same usage, applied to a model with rather more of them than a statistician of an earlier generation would have contemplated.

Nothing about a parameter is meaningful on its own. Pick one at random and read its value and you learn nothing, because what it does depends entirely on every other parameter it interacts with. The information is distributed, and it is distributed in a way that resists being localised.

Capacity is the right word for what it measures

The count tells you how much a model could in principle represent. More parameters mean a more flexible function, capable of fitting more complicated relationships and holding more distinctions. It is an upper bound on expressiveness, and upper bounds are not achievements.

Whether that capacity gets used well depends on the training data, the objective, the architecture, the optimisation procedure and the amount of computation spent. Two models with identical counts can differ enormously, and a well-trained smaller model regularly outperforms a poorly trained larger one on the same task.

This is not a minor caveat. A great deal of published work has shown that for a fixed computing budget there is a trade-off between model size and quantity of training data, and that historically the field sat on the wrong side of it, building models larger than the data being fed to them justified.

The count became a marketing figure

Parameter counts are easy to state, comparable across systems and impressive when large, which made them irresistible as a headline number. For a period the size of a model was reported the way engine displacement used to be reported for cars, and with about as much predictive value for how the thing performs in traffic.

The figure has become less informative over time rather than more, for several reasons. Numerical precision varies, so two models with the same count can occupy very different amounts of memory. Some architectures activate only a fraction of their parameters for any given input, which means the count and the work done per query have come apart entirely.

Several organisations have stopped publishing the number at all. That is sometimes read as concealment and is at least as likely to reflect that the figure no longer answers the question people are asking with it.

Where the knowledge is, if not in individual parameters

It is tempting to imagine specific facts living at specific addresses, and this is mostly wrong. Representations are distributed across many parameters, and a single parameter participates in many unrelated representations. This is why you cannot delete a fact from a model by finding it and removing it.

The picture is not entirely uniform, though. Interpretability research has located structures that behave like recognisable concepts and has demonstrated that intervening on them changes behaviour in predictable ways, which is a genuine result rather than a metaphor. Certain layers appear to specialise in retrieving factual associations.

How far this localisation goes is disputed, and there is a real argument about whether the identified structures are properties of the model or artefacts of how they were searched for. Anyone claiming that a particular capability lives in a particular place should be asked how that was established.

Reading the number sensibly

A parameter count is one input to a rough judgement and it does tell you something real: about how much memory the model needs, roughly what class of hardware can run it, and what order of magnitude of expense was involved. Those are useful facts.

What it does not tell you is how good the model is at anything, how current its information is, how it behaves under pressure, or whether it suits your task. None of those correlate reliably with size once you are comparing systems built with comparable care.

The habit worth adopting is to treat the figure the way you would treat the length of a book. It constrains what could be in there. It says almost nothing about whether it is any good.

Common questions

Do more parameters always mean better performance?

No. Holding everything else equal it helps, but everything else is never equal — training data quantity and quality, computation spent and architectural choices all matter as much or more. Smaller models trained longer on better data routinely beat larger ones, which is now a well-established result rather than a curiosity.

What are active parameters?

In architectures that route each input to a subset of the model rather than through all of it, only some parameters participate in any given computation. The total count then describes the storage required while the active count describes the work per query, and quoting only one of the two is misleading in opposite directions.

Can parameters be removed after training?

Yes, and this is standard practice. Pruning removes connections contributing little, and reducing numerical precision shrinks each stored value. Both cut memory and cost with some loss of quality, and how much is lost varies by technique, by model and by task in ways that are still being mapped.

Jargonparametersweightscapacityterminology
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.