Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Being confidently wrong is the default behaviour, not a malfunction

Nothing in the way these systems are built distinguishes a supported statement from an unsupported one, so the surprising thing is not the false answers but how many true ones there are.

By Manish Trivedi4 min read

Close-up of HTML code with syntax highlighting on a computer monitor.
Photograph by Bibek ghosh via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The system was never given a concept of truth

A language model is fitted to produce text that resembles the text it was trained on. At no point in that process is there a separate check asking whether a statement corresponds to anything in the world. There is no database being consulted, no fact table, no flag distinguishing a well-attested claim from one the model has assembled out of the general shape of similar claims.

What the model holds is a very large set of statistical regularities about how words follow other words. Many of those regularities happen to encode facts, because factual text is internally consistent in ways that fiction and error are not, and consistency is exactly what a statistical fit picks up. But the encoding is incidental.

Once you accept that, the puzzle inverts. The interesting question is not why these systems produce false statements. It is why a mechanism with no representation of truth produces correct answers as often as it does.

Plausibility and accuracy come apart at the edges

For well-covered material, plausible and accurate largely coincide. A fact repeated across thousands of documents leaves a deep, consistent groove in the statistics, and the most probable continuation is the true one. This is why these systems are reliable on the general and unreliable on the specific.

Move to something thinly covered — a minor figure, an obscure regulation, a specific page number, an exact date in a niche field — and the groove is shallow or absent. The model still produces a continuation, because producing a continuation is the only thing it does, and that continuation will have the correct shape. A citation will look like a citation. A statute number will look like a statute number.

This is why fabricated references are such a characteristic failure. The format of a reference is enormously well represented in the training data while any individual reference may appear once or not at all, so the system reproduces the pattern with confident precision and invents the content.

Hedging was trained down as well as up

Written text overwhelmingly makes claims rather than declining to. Encyclopaedias, textbooks, manuals and articles are full of assertions and comparatively empty of passages where the author stops and says they do not know. A model fitted to that corpus inherits the distribution of confidence in it.

Later training stages using human feedback adjust this, and they can push in either direction. A rater comparing two answers tends to prefer the helpful and specific one over the cautious one, and preferences aggregated across many raters therefore reward assurance. Deliberate effort goes into counteracting this, with mixed and much-debated success.

There is a genuine tension here that nobody has resolved cleanly. A system that refuses whenever it is uncertain becomes useless, since it is always somewhat uncertain. A system that never refuses is dangerous in exactly the situations where it matters most. Every deployment picks a point on that trade-off and no point is comfortable.

What retrieval fixes and what it does not

The most effective mitigation is to stop asking the model to recall and start giving it the source material, retrieving relevant documents and asking it to answer from those. This substantially reduces fabrication on questions where a good document exists, and it makes answers checkable, which matters more than the accuracy improvement itself.

It introduces its own failure modes. The retrieval step can fetch the wrong document, or a document that is outdated, or one that is confidently incorrect, and the model will answer from it faithfully. Where the retrieved material is ambiguous or partially relevant, the system may blend it with its own prior statistics and produce something that is neither.

So retrieval changes the question from whether the model knows something to whether the corpus does and whether the search found it. That is a better question to be dealing with, because it can be inspected. It is not the same as solving the problem.

The honest framing for using any of this

The useful mental model is not an assistant that occasionally errs. It is a system that produces well-formed output regardless of whether it has any support for the content, where the presentation is identical either way. Fluency carries no information about reliability, and the temptation to read it as though it does is the actual hazard.

That is why domains split so cleanly on whether these tools are appropriate. Where an error is cheap and visible — drafting, summarising something you will read anyway, generating options you will evaluate — the failure mode is tolerable. Where an error is expensive and invisible, and where the person receiving the output cannot check it, the same tool becomes a liability.

Nothing about this is likely to be fully fixed by a better model, since it follows from what the training objective is rather than from how well it was optimised. Substantial improvement is entirely possible and has already happened. Elimination would require a different kind of system, and whether anyone is building one is a matter of opinion rather than record.

Common questions

Why does asking the model to double-check sometimes work?

Because a second pass conditions on the first answer as text, and evaluating a stated claim is a slightly different statistical task from generating one. It genuinely catches some errors. It also produces false corrections and can talk the system out of a right answer, so it is a weak filter rather than a verification step.

Are some subjects worse than others for this?

Yes, and predictably so. Anything where the answer is specific, thinly documented and formatted regularly is high risk — citations, case numbers, product specifications, quotations, precise dates and figures. Broad conceptual explanations of well-covered topics are considerably safer.

Does the word hallucination describe this well?

Many researchers think not, and there is an ongoing argument about the term. It implies a departure from some normal perceiving state, when in fact the same process produces the true answers and the false ones. Alternatives like confabulation or fabrication have been proposed, without much agreement.

Limits & Riskshallucinationreliabilitytruthfailure modes
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.