Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Ask the same question twice and the answers differ, for two separate reasons

Repetition producing different text is the expected behaviour of a system that samples from a distribution, and underneath that sits a second source of variation which is much harder to remove.

By Samar Bhatia3 min read

Detailed shot of Ethernet cables connected to server ports highlighting technology infrastructure.
Photograph by Brett Sayles via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The model doesn’t produce a sentence, it produces a distribution

At every step a language model outputs a score for every fragment in its vocabulary — tens of thousands of numbers, converted into probabilities that add up to one. That is the entire output of the network. Turning it into a piece of text requires a separate decision, and that decision is not part of the model.

The component making it is called the decoder, and it is a comparatively simple piece of code. It can take the highest-scoring fragment every time, or it can draw at random with each fragment’s chance proportional to its probability, or it can do something in between the two.

The chosen fragment is then appended to the input and the process repeats. So a single random draw early on redirects everything that follows, because the next step is now conditioned on a different sentence than it would otherwise have been. Small divergences don’t stay small.

The settings that shape the draw are doing something specific

Temperature rescales the scores before they become probabilities. Below one it exaggerates the gap between likely and unlikely options and the output becomes more predictable; above one it flattens the distribution and rare fragments start appearing. At zero the draw collapses into always taking the top score.

Other settings truncate rather than rescale. One keeps only the few highest-scoring options; another keeps however many are needed to cover a set share of the probability mass, which adapts to how confident the model happens to be at that particular step. Both exist to stop the long tail of implausible fragments being sampled at all.

None of this changes what the model produced. The distribution is fixed by the parameters and the input; the settings govern only how it is consumed. That separation is easy to lose sight of, because the visible effect of changing them is so dramatic that it feels like a change to the model.

Always taking the most likely fragment is not the safe option

It sounds as though taking the top score at every step should give the best answer, since each choice is locally optimal. It does not, and the reason is that a sequence of locally optimal choices is not a globally optimal sequence. Text generated this way tends to loop and to flatten out.

The behaviour is well documented: repeated phrases, circling back to the same construction, prose that reads as though it is afraid to commit. Some randomness demonstrably produces text that people judge better, which is an odd fact when stated plainly. Deliberate noise improves the output.

For tasks with one correct answer the calculation differs, and deterministic decoding is often preferred there because reproducibility is worth more than variety. It is a trade rather than a ranking, and which side it falls on depends entirely on what the output is for.

Even with the randomness switched off, runs can differ

This is the second source, and it catches people out. Fix the settings to be fully deterministic and identical inputs can still produce different outputs on a large serving system. The cause is not the model and not the decoder. It is arithmetic on real hardware.

Adding floating-point numbers is not associative: group the same values differently and the last digits can disagree. Large models split their arithmetic across many processors and combine partial results in whatever order those results arrive, so tiny discrepancies appear. Usually they vanish. Occasionally one lands where two fragments were nearly tied, and flips the choice.

Batching makes this more likely, because the grouping depends on which other requests happened to arrive at the same moment. A request processed alongside twenty others may follow a different arithmetic path than the same request processed alone. Nothing about the person’s input changed at all.

Variation is a measurement problem before it is a user problem

For everyday use this matters little, since two phrasings of a good answer are both good answers. For evaluation it matters enormously, because comparing two systems on a handful of outputs cannot distinguish a real difference between them from the spread of a single system against itself.

Careful evaluation therefore runs many samples and reports the variation, which is ordinary practice in other empirical fields and has been adopted unevenly here. Comparisons published without it should be read as suggestive. That is not an accusation against anyone in particular; it is a consequence of how expensive repeated runs are.

Debugging inherits the same difficulty. A failure that appears once in twenty runs is genuinely present and genuinely hard to reproduce, and an engineer who cannot make it happen again may reasonably conclude that it never did. It did.

Common questions

Does setting temperature to zero make a model deterministic?

It removes the deliberate randomness, which accounts for most of the variation. It doesn’t guarantee identical output, because the arithmetic underneath can still resolve near-ties differently depending on how a request was batched and which hardware ran it. Repeated runs agree very often, not invariably.

If randomness helps the writing, why not turn it up further?

Because the same setting that produces a surprising phrase also produces an unsupported claim. Flattening the distribution makes low-probability fragments accessible, and low probability is where both originality and error live. No setting admits one while excluding the other.

Is the variation a sign that the model is unsure?

Not reliably. The spread of the distribution carries some information about the model’s confidence, but that confidence is poorly calibrated in general, and a system can be uniformly confident about something false. Differing answers are evidence worth noticing rather than a measurement of doubt.

How It Workssamplingdecodingdeterminismtemperature
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.