Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Correlation carries these systems a long way and then stops at the edges

Statistical association is enough to produce behaviour that looks like understanding across most of the range, and the difference only becomes visible where the training data ran out.

By Manish Trivedi3 min read

Close-up of a dark room with a curved monitor showing the ChatGPT interface on screen.
Photograph by Matheus Bertelli via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The interior and the boundary behave differently

A trained model has effectively drawn a surface through the examples it was shown. Inside the region those examples cover, the surface is well supported and predictions are good, and this region is often far larger and stranger in shape than intuition suggests. Outside it, the surface continues, because a mathematical function does not stop at the edge of the data, but nothing is holding it up.

The system gives no indication of which side of that boundary a given input falls on. It produces an output with the same apparent assurance either way, and it has no representation of the boundary itself to consult.

Most discussion of what these systems can and cannot do is really a disagreement about where that boundary sits, conducted by people generalising from a handful of examples on either side of it.

Shortcuts are learned in preference to explanations

A model minimising error will use whatever feature predicts the answer, regardless of whether that feature has anything to do with the underlying phenomenon. If a particular background appears in most images of a category, the background becomes a feature. If a certain phrasing appears in most positive examples, the phrasing becomes the signal.

This has been demonstrated repeatedly across domains, and the pattern is always the same: the model achieves excellent test performance, the test set shares the same quirk as the training set, and the system fails the moment it meets data where the quirk is absent. The failure looks sudden and inexplicable from outside, and is entirely predictable from inside.

The important part is that this is not a bug being fixed. It is the optimisation working correctly. A shortcut that predicts well is a good solution by the only criterion the training process has, and preferring the deeper explanation would require some pressure that the objective does not supply.

Causal structure is not recoverable from observation alone

There is a well-established result underneath all of this, older than modern machine learning. Data that only records what co-occurred cannot, in general, distinguish between competing causal explanations of those co-occurrences. Two quite different accounts of what causes what can produce identical observational data.

Distinguishing them requires intervention — changing one thing deliberately and seeing what follows — or assumptions strong enough to rule out the alternatives. A system trained on a fixed corpus intervenes in nothing. It can absorb causal claims that appear in the text, which is genuinely useful, but that is inheriting somebody else’s causal knowledge rather than deriving it.

This is why the same system can explain a mechanism impeccably and then reason badly about a counterfactual involving it. Reciting the relationship and using it to answer what would have happened if things were different are different operations, and only the first is well supported by text statistics.

Small changes that should not matter sometimes do

A characteristic symptom is instability under irrelevant variation. Rephrasing a question, changing the order of options, adding a sentence that has no bearing on the problem, or altering names and numbers in a word problem can all shift the answer, sometimes from right to wrong.

A system that had extracted the structure of the problem would be indifferent to these things. A system that has learned a strong association between surface patterns and answers is not, and the degree of sensitivity is one of the better available indicators of which is happening in a given case.

Sensitivity has decreased noticeably as systems have grown, and this is the strongest evidence for the view that scale genuinely does push towards more structural solutions. Whether it approaches zero or plateaus at some level is exactly the point in dispute, and current evidence does not settle it.

Neither triumphalism nor dismissal survives the evidence

The dismissive reading — that this is only correlation and therefore not really doing anything — has aged poorly, repeatedly. Capabilities that were confidently declared impossible for statistical systems have arrived, and the people who declared them impossible were often reasoning from the same argument sketched above.

The triumphalist reading has aged poorly too. Failures at the edges have proven stubborn, they recur in new forms with each generation of system, and the pattern of a capability appearing to work and then collapsing under mild perturbation has repeated often enough to be a genre.

The defensible position is unsatisfying and probably correct: correlation over a sufficiently vast and varied corpus produces something far more capable than the classical objection anticipated, and something meaningfully short of what its performance in the well-supported region suggests. Both halves matter, and most public argument holds only one.

Common questions

Is there a way to tell whether a model is using a shortcut?

Not from the output alone. The practical test is to construct examples where the shortcut and the correct answer come apart, which requires guessing what the shortcut might be. This is a standard technique in evaluation and it works, but it only finds the shortcuts somebody thought to look for.

Do these systems ever learn genuine causal relationships?

They learn to reproduce causal claims stated in their training data, which functions well for anything commonly written about. Whether anything in the internal representation corresponds to causal structure rather than a rich association is unresolved, and researchers examining model internals report evidence that is suggestive in both directions.

Why does giving worked reasoning improve answers?

Producing intermediate steps changes the computation, since each step conditions what follows, and it makes the model spend more of its processing on the problem before committing. It genuinely helps on multi-step problems. It does not guarantee the steps are the reasoning that produced the answer, which is a separate question.

Limits & Risksgeneralisationcausationshortcutsrobustness
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.