Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Inference is a borrowed word for the part that runs every single time

The term sounds like deduction and means something much more mundane, and the confusion it causes is worth clearing up because most of the operational vocabulary hangs off it.

By Daniel Okonkwo3 min read

Dictionary page close-up with a magnifier highlighting the word 'discrimination.'
Photograph by Nothing Ahead via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two meanings collided and the wrong one stuck

In logic, inference is the drawing of a conclusion from premises. In statistics, inference is the estimation of properties of a population from a sample. In machine learning as it is practised, inference is neither: it means running a trained model on an input and reading the output. That is all.

The word arrived through the statistical tradition, where using a fitted model to produce an estimate for a new case is a reasonable thing to call inference. It then detached from that context and became the standard industrial term for the operation of a deployed system, at which point the connotation of reasoning came along for free.

This causes real confusion outside the field, since it makes the operational sentence about inference costs sound as though somebody is buying reasoning by the hour. What is being bought is arithmetic. Whether reasoning is occurring is a separate and much harder question.

What the operation consists of

An input arrives, is converted into numbers, and passes forward through the layers once. For a text model this produces a probability distribution over possible next fragments, one is selected, appended to the input, and the whole thing runs again for the next fragment. A paragraph of output is many passes through the model.

Nothing is stored between requests unless something outside the model stores it. The parameters do not change, no record of the conversation persists in them, and two identical requests are two entirely independent events. Any apparent memory is text being supplied again as input.

That statelessness is why serving these systems is a tractable engineering problem at all. It also explains a good deal about their behaviour that people attribute to design decisions.

The vocabulary of serving is mostly about queueing

Latency is the delay before an answer, and it splits usefully into two parts: how long until the first fragment appears, and how quickly the rest follows. These are governed by different bottlenecks, which is why a system can begin responding promptly and then produce text slowly, or pause and then deliver everything at once.

Throughput is total work completed per unit of time across all users, and it usually trades against latency. Operators combine many requests into batches to use the hardware efficiently, and a request that waits for its batch to fill has been made slower to make the system cheaper. Almost every complaint about response times traces back to some position on that trade.

The rest of the vocabulary follows the same logic. Caching stores the intermediate results for text already processed so it need not be recomputed. Quantisation reduces the precision of the stored parameters to fit more model into less memory. These are optimisations of one operation performed extremely often.

Cost behaves differently from ordinary software

Conventional software has an approximately fixed cost per request that falls with scale, since the expensive part was writing it. Inference does not behave that way. Each request consumes real computation proportional to how much text goes in and comes out, and serving twice as many users costs approximately twice as much.

This is why services are billed by volume of text rather than by number of requests, and why a long document costs more to process than a short one. It also explains why free tiers are limited in ways that free tiers of conventional software are not, and why usage caps exist at all.

The comparison worth holding is with electricity rather than with software. There is a meter, the meter runs while the thing is in use, and nothing about having many customers makes any individual query cheaper in the way that distributing a program does.

Why the naming is worth objecting to

Words shape expectations, and this one imports a claim. Describing the operation as inference suggests something is being worked out, which is precisely the disputed question about what these systems do. A neutral term — evaluation, or simply running the model — would leave the question open where it belongs.

Several other pieces of the field’s vocabulary do the same work. Learning, attention, hallucination, understanding and reasoning were all borrowed from human experience for mechanisms that resemble their namesakes loosely at best, and each one smuggles in an assumption that then has to be argued back out.

This is not a plea for jargon reform, which never works. It is a suggestion for reading. When a technical term describes a mental activity, check whether the mechanism earns the word, because in this field it frequently does not, and the borrowed word is doing persuasive work its author may never have intended.

Common questions

Why is the first word slower to arrive than the rest?

Because the entire input must be processed before any output can begin, and that work grows with input length. Once generation starts, each subsequent fragment requires much less work because the results for earlier text have been computed and retained for the duration of the request.

Does running a model on my own machine avoid these costs?

It moves them rather than removing them. You pay in hardware, electricity and the effort of running the thing, and typically in capability, since the models that fit on ordinary hardware are smaller. What you gain is that your inputs stay local, which for some purposes is the entire point.

Is inference the same as prediction?

They are used interchangeably in much of the field, with prediction the older statistical term for producing an estimate for a new case. Inference has become the operational word, particularly when discussing infrastructure, while prediction tends to be used when discussing what the model is doing mathematically.

Jargoninferenceservinglatencyterminology
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.