Jargon
Ground truth is a confident name for something people had to agree on
The phrase borrowed from surveying implies a fact against which everything else is measured, and in most machine learning the thing it names is a recorded human judgement with all the softness that implies.
By Samar Bhatia3 min read

The metaphor came from measurement, and it does not travel well
The term originates in remote sensing and surveying, where an observation from a distance is checked against a measurement taken on the ground. There the metaphor is exact: somebody stood in the field and measured the thing, and the aerial reading either matches or does not.
Transplanted into machine learning, it names whatever is being treated as the correct answer for training and evaluation purposes. Sometimes that is a genuine measurement. Frequently it is a person’s opinion recorded under time pressure, and the phrase carries the authority of the former into the latter.
This isn’t a pedantic complaint about wording. Systems are assessed against these labels, decisions are made on the strength of the assessment, and the confidence attaching to the phrase transfers to the conclusion without anybody examining what was actually recorded.
Annotators disagree, and the disagreement is often the signal
On any task involving judgement, two people given the same item and the same instructions will sometimes label it differently. Agreement rates are measured, reported in careful work, and are frequently far from perfect on exactly the tasks that matter most.
The standard response is to take a majority of several annotators, which produces a single label and discards the information that the item was contested. An item where three people split evenly and an item where all agreed become identical in the training data, having told you quite different things.
A body of research argues that this discarding is a mistake, and that disagreement should be preserved and modelled rather than resolved. The counter-argument is practical: most training machinery expects a single answer, and preserving disagreement complicates everything downstream. The debate is genuine and ongoing.
The recorded label is often a proxy for what you wanted
A great many labels are not judgements at all but traces of something else. Whether somebody clicked stands in for whether they found it useful. Whether a case was flagged stands in for whether it was fraudulent. Whether a patient was diagnosed stands in for whether they had the condition.
Each substitution imports whatever produced the trace. A diagnosis label records who was tested and by whom, not who was ill. A flag records what the previous system caught, not what occurred. Training against these teaches a model to reproduce the recording process rather than the phenomenon.
This is one of the most common and least visible sources of failure in applied work, because the label looks like data and the substitution happened before anybody in the modelling team was involved.
The tell is usually a label that was cheap to obtain. Anything recorded automatically at scale was recorded because it was easy to record, and ease of recording is unrelated to how closely a trace resembles the thing somebody actually wanted to know about.
Gold, silver and the honest hierarchy
Careful practice distinguishes levels. A gold set is labelled with unusual care, often by experts, often with adjudication of disagreements, and is small because that is expensive. A silver set is labelled more cheaply, perhaps automatically or by a single annotator, and is large.
The sensible arrangement uses the large cheap set for training and the small careful set for final evaluation, on the reasoning that training tolerates noise better than measurement does. Where the same cheap process supplies both, the evaluation inherits every systematic error in the labelling and cannot detect it.
Automatically generated labels have made this more pressing, since a set produced by another model is cheap, plentiful and shares that model’s blind spots exactly. Evaluating against it measures agreement with the labeller rather than correctness.
A better habit than accepting the phrase
When a system is reported to reach some level of accuracy, the informative question is what it was accurate against. Who produced the reference answers, under what instructions, with what agreement between them, and how were disagreements handled. Careful papers report this and many do not.
It is also worth asking whether the task has a correct answer in principle. Some do, and for those the phrase is defensible. Many involve judgement on which reasonable people differ, and for those a perfect score would indicate agreement with a particular set of annotators rather than correctness.
The term is not going away, and it is unlikely to be replaced by anything better. Reading it as agreed reference rather than as truth costs nothing and prevents a specific and common category of overconfidence.
Common questions
Are there tasks with genuine ground truth?
Yes, and they are the ones where an external process can confirm the answer: a measured physical quantity, the outcome of a game under fixed rules, whether a program compiles. Those cases justify the phrase, and they are a minority of what these systems are trained on.
Does more annotators per item fix the problem?
It reduces random variation and does nothing about systematic disagreement, since a larger group drawn from the same population shares its assumptions. Improving the reference usually means changing who annotates and how the instructions are written, not simply adding people.
Why do published accuracy figures rarely mention this?
Partly convention and partly because agreement statistics complicate a headline number. It is more commonly reported in work where annotation was a substantial part of the effort, which is a reasonable heuristic for how carefully it was done.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





