Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

The cost of training decides who gets to build, and that has already narrowed

A field that ran for decades on university budgets now has a frontier accessible to a handful of organisations, and the second-order effects of that concentration reach into what gets researched at all.

By Samar Bhatia3 min read

A robotic arm engages in a chess match, showcasing AI and robotics in a studio setup.
Photograph by Pavel Danilyuk via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A discipline changed its cost structure

For most of its history, artificial intelligence research was affordable. The important ideas of several decades were developed on hardware that a department could buy, and a graduate student with a good idea could test it against anyone else’s. That equality of access shaped the field’s culture, its publication norms and its self-image.

Training a frontier model today is a capital project. The computation involved is rented or owned at a scale that no ordinary research budget reaches, and the surrounding costs — data acquisition, engineering staff, evaluation, infrastructure — are of the same order. This is not a difference of degree that patience can overcome.

The consequence is straightforward and rarely stated plainly. A specific and important class of experiment can now only be run by organisations that can commit very large sums, and everybody else works downstream of what those organisations choose to build and release.

Concentration changes which questions get asked

Research directions follow the resources available to test them. When the frontier requires enormous computation, hypotheses that can only be evaluated at that scale become the property of the few groups able to evaluate them, and everybody else redirects towards questions answerable with less.

That redirection is not purely a loss. A great deal of valuable work — efficiency, evaluation, interpretability, applications, small models — is done precisely because the frontier is closed, and some of it has turned out to matter more than another increment of scale would have.

But there is a real cost in what does not get examined. Negative results at scale are expensive and rarely published. Alternative architectures cannot be given a fair test, because a fair test means the same computational budget as the incumbent, and nobody will spend that on a long shot. The field may be more path-dependent than it appears.

Independent verification becomes structurally difficult

Science depends on somebody being able to repeat the experiment. When the experiment costs more than most institutions have, replication stops functioning as a corrective, and claims are assessed on the credibility of the claimant rather than on independent confirmation.

This is not an accusation of dishonesty. It is a description of how evidence works when reproduction is unaffordable, and it applies to any field with expensive apparatus — high-energy physics has lived with a version of it for a long time, and has evolved elaborate norms in response.

The AI field has not yet evolved comparable norms. Evaluations are frequently run by the organisation whose system is being evaluated, using tests it selected, on infrastructure nobody outside can inspect. Everybody involved knows this is unsatisfactory and no accepted alternative has emerged.

The counter-argument deserves a fair hearing

Concentration at the frontier has coincided with unusually wide access downstream. Capable models are available to anyone with a network connection at prices that would have been implausible a few years earlier, and released weights have put serious capability into the hands of researchers who could never have trained it themselves.

It is also not obvious that the gap between the frontier and what is freely available has widened rather than narrowed. Techniques diffuse quickly, and capability that required extraordinary resources in one year has repeatedly become reproducible with far less in the next. Whether that pattern continues is genuinely uncertain and worth marking as such.

So the honest summary is mixed: building the most capable systems has concentrated sharply, while using capable systems has democratised sharply, and which of those matters more depends on what you think research and industry actually need.

Public capacity is the variable to watch

Several governments have begun funding shared computing capacity for academic and public-interest research, on the reasoning that a field of this consequence should not be assessable only by its own vendors. Whether these initiatives reach a scale that matters is an open question and the record of such programmes is mixed.

The mechanism is at least plausible. If independent evaluation is the missing corrective, then funding the apparatus for independent evaluation addresses the actual gap, in a way that publication requirements and voluntary commitments do not.

It would be foolish to predict the outcome. The point is narrower: the distribution of computing capacity is not a technical detail but a determinant of who can produce knowledge about these systems, and that makes it a political question whether or not anybody treats it as one.

Common questions

Can a small team still do meaningful research?

Yes, and a great deal of the most useful recent work has come from small groups — evaluation methods, interpretability, efficient training, domain applications. What has become impractical for a small team is training a frontier general model from scratch, which is one important line of work rather than the whole field.

Does more money reliably produce a better model?

It helps considerably and it does not settle the matter. Data composition, training procedure and engineering quality vary enormously between projects with similar budgets, and there are well-known cases of expensive efforts producing disappointing results. Capital is necessary at the frontier without being sufficient.

Why does concentration matter if the models are available cheaply?

Because availability and accountability are different things. Being able to use a system does not let you check how it was built, verify claims about it, or study alternatives that were never funded. Cheap access addresses the first problem and leaves the others untouched.

In The Worldeconomicsconcentrationresearchcapital
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.