Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Training a model and running one are two different kinds of work

One is a construction project with a fixed budget and an end date; the other is a utility bill that arrives every time somebody presses a key.

By Zoya Rahman4 min read

Professional woman standing confidently in a data center, surrounded by glowing servers.
Photograph by Christina Morillo via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two activities, one confusing word

People speak about the cost of artificial intelligence as though it were a single figure, and it is not. Building a model and using a model are separate engineering problems with separate bottlenecks, separate hardware preferences and separate economics. Conflating them produces most of the bad reasoning you will encounter about who can afford to do what.

Training is a one-off capital event. A quantity of computation is bought or rented, run for a period measured in weeks rather than seconds, and at the end there is an artefact: a file of numbers that can be copied indefinitely at almost no cost. The expenditure is enormous, concentrated, and then finished.

Inference is the opposite shape. Each individual run is cheap, but it happens again for every request, forever, and it scales directly with how many people are using the thing. A model that becomes popular gets more expensive to operate, whereas a model that becomes popular does not need retraining to serve them.

The forward pass is the cheap half of the arithmetic

Running a model means pushing an input through the layers once, in one direction, and reading what falls out. That is called a forward pass, and it is straightforward multiplication and addition on a very large scale — expensive by ordinary computing standards, but bounded and predictable.

Training does the forward pass too, and then does considerably more. It computes the loss, sweeps backwards through every layer to work out how each parameter contributed to the error, and then applies an update. The backward sweep costs roughly as much arithmetic again as the forward one, sometimes more, and it has to happen for every batch of examples in the entire run.

The multiplier that really matters is repetition. A single training run passes over an immense quantity of data many times over, so the total work is the per-step cost multiplied by an extremely large number of steps. That product, rather than any single clever operation, is where the electricity goes.

Memory tells the two apart more sharply than speed does

To compute how much each parameter contributed to the error, a training system must keep the intermediate results of the forward pass in memory until the backward sweep reaches them. Inference can discard each layer’s output the moment the next layer has consumed it. Training cannot, and this difference dominates the hardware requirements.

On top of that, most training procedures hold additional numbers alongside every parameter — running averages that smooth the updates and stop the process oscillating. So the memory footprint during training is several times the size of the finished model, while the footprint during inference is roughly the model itself plus whatever the current conversation occupies.

This is why a model that comfortably runs on modest hardware may have required a data centre to produce. The asymmetry is structural rather than a matter of optimisation, and no amount of engineering removes it entirely.

The two jobs want different machines

Training is a throughput problem. Nobody cares whether any individual example takes an extra millisecond; what matters is total work completed per hour, so the work is batched aggressively and spread across many processors that must communicate constantly. Interconnect bandwidth between chips becomes a limiting factor in a way it never is for a single query.

Inference is often a latency problem instead. A person waiting for a reply notices delay directly, and text is generated one piece at a time with each piece depending on the last, so a large part of the work cannot be parallelised away. Operators batch many users’ requests together to recover efficiency, which trades individual responsiveness against overall cost.

The result is two quite different procurement conversations happening under one heading. Chips optimised for one job are frequently mediocre at the other.

Why the distinction changes who can participate

Because a trained model is just a file, the ability to run one is far more widely distributed than the ability to produce one. Adapting an existing model to a narrower purpose costs a small fraction of building it from nothing, which is the practical reason a great deal of applied work starts from something somebody else trained.

How that balance develops is genuinely contested. One view holds that the aggregate cost of serving models will come to dwarf the cost of building them, since serving grows with users while building happens occasionally. The opposing view is that model builders keep raising the bar of what a frontier system costs, so construction stays the dominant expense. Both arguments are plausible, both depend on efficiency improvements nobody can schedule, and anybody stating either as settled fact is guessing.

What is not speculative is the shape of the two curves. One is a lump; the other is a meter that never stops running.

Common questions

Is inference always cheaper than training?

Per run, overwhelmingly yes. In aggregate, not necessarily — a widely used model can accumulate more total computation serving requests over its lifetime than it consumed being built. Which side wins depends on usage volume, and that varies so much between systems that a general answer is not available.

Why does a longer conversation seem to cost more?

Because the model re-reads its context on every step, and the work involved in relating each piece of text to every other piece grows faster than the length itself. Various engineering techniques reduce the repeated effort considerably, but the underlying relationship still makes long inputs disproportionately expensive.

Can a model be made cheaper to run after training?

Yes, and this is an active area. Reducing the numerical precision of the stored parameters, pruning connections that contribute little, or training a smaller model to imitate a larger one all cut running costs. Each involves some loss of capability, and how much is very much disputed and depends on the task.

How It Workstraininginferencecomputehardware
Zoya Rahman
Deputy editor, AI Worth Knowing

Zoya joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and would rather show the working than assert the conclusion.