Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Foundation model was a naming decision, and it was argued about immediately

The term was proposed to describe a genuine shift in how systems are built, the objections to it were raised at the time, and the words the field settles on now determine what rules can be written later.

By Samar Bhatia3 min read

Close-up of words 'poem' and 'poet' magnified in a dictionary
Photograph by Nothing Ahead via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The shift the word was coined to name

Something real changed in how systems get built. For decades a model was trained for a task: one for sentiment, one for translation, one for recognising faces. Then it became normal to train one very large model on a broad corpus and adapt it to many tasks afterwards, with the adaptation being comparatively cheap.

That arrangement needed a name, because the older vocabulary did not distinguish it. A term proposed by a research group to describe such models was chosen deliberately to emphasise that they act as a base others build upon, and it spread quickly through both technical and policy writing.

The naming was accompanied by an argument that the shift brought a new kind of risk: if many downstream systems rest on the same base, then a defect in the base propagates everywhere at once. That observation was the substantive point, and it holds regardless of what the thing is called.

The objections were raised at the time

Critics noted that the word carries connotations of solidity and completeness that the systems do not earn, and that describing something as a foundation invites treating it as settled rather than as a component with known defects. Language shapes expectation, particularly among people reading at a distance from the technology.

A second objection was that the term implied a sharper break from earlier practice than the record supports. Training a general model and adapting it was not new; it had been standard in computer vision and had precursors in language work. The novelty was scale and breadth rather than the idea itself.

A third was simpler: the alternatives were available and more descriptive. Large pretrained model says what it is. General-purpose model says what it is for. Neither is elegant, and both resist being read as an endorsement.

A second term arrived for the policy conversation

As regulation began to be drafted, another phrase came into use for the most capable systems, distinguishing them from the general population of large models. It is used in policy documents to mark the category warranting additional scrutiny.

The difficulty is defining it. Definitions have been attempted by training compute, which is measurable and only loosely connected to what a system can do, and by capability, which is closer to what anybody cares about and cannot be measured in a way that different parties would agree on.

Both approaches will age. A compute threshold set today becomes ordinary as efficiency improves, and a capability threshold depends on evaluations that are themselves disputed. Any definition attached to a fixed number is a definition with an expiry date, which drafters generally know and cannot entirely avoid.

Once an obligation is written against a category, the boundary of that category becomes a place where money and liability sit, and everybody with an interest starts arguing about which side of it they fall on. The vaguer the term, the more of the eventual dispute is about classification rather than conduct.

This is not unique to this field. It is what happened with categories in employment law, in financial regulation and in medical devices, and the pattern is well documented: a term coined for descriptive convenience becomes load-bearing, and its imprecision is inherited by everything built on it.

It also creates an incentive to be described one way rather than another, which is worth remembering when reading how organisations characterise their own systems in submissions and documentation.

A reasonable position on the vocabulary

The underlying distinction is genuine and worth preserving: some models are built to be adapted and others are built for a task, and the first kind concentrates risk in a way the second does not. That is the observation the terminology was meant to capture.

Whether this particular word was the right vehicle is a fair question, and it is largely moot now, since terms are settled by use rather than by argument and this one is in use. Objecting to established vocabulary is a poor use of anybody’s energy.

What remains useful is precision about what is being claimed. When a system is described this way, the informative questions are what it was trained on, how it is adapted, what rests on it, and what happens downstream if it is wrong. The label answers none of those, which is true of most labels.

Common questions

Is every large language model a foundation model?

Under the usual definition, one built to be adapted to many downstream uses qualifies, and a large model trained for a single narrow purpose does not. The boundary is fuzzy because adaptability is a matter of degree, and usage has drifted towards applying the term to any large model.

Why do policy documents avoid naming specific technologies?

Because a rule naming a specific architecture becomes obsolete when the architecture changes, which in this field can happen within a couple of years. Drafters prefer categories defined by function, scale or risk, which age better and are correspondingly harder to apply to a particular system.

Does the terminology affect how systems are built?

Indirectly. Where obligations attach above a threshold, there is an incentive to stay below it or to structure a system so it falls into a lighter category. Whether that produces genuinely safer design or merely creative classification depends on how the threshold was drawn.

Jargonfoundation modelterminologypretrainingpolicy
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.