Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

What a benchmark score measures and what it quietly does not

Every claim about a system being better than another rests on a test, and the properties of those tests are less well understood than the numbers they produce.

By Samar Bhatia3 min read

Elderly man at computer with termination notice, facing unemployment.
Photograph by Ron Lach via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A benchmark is a proxy for a proxy

To claim that one system is better than another you need a measurement, and measurement requires a fixed set of questions with known answers. That is what a benchmark is: a collection of problems, a scoring rule, and a number at the end that can be put in a table.

The chain of reasoning underneath is longer than it looks. You care about some capability. Somebody constructs tasks intended to require that capability. The tasks are scored automatically, which constrains them to have checkable answers. The resulting number is then reported as evidence about the original capability, several steps removed.

Each step in that chain can hold or fail, and whether the tasks genuinely require the capability they are named after is the step most often skipped. That question has a name in measurement theory — construct validity — and it is asked far less often than it should be.

Contamination is the structural problem

These systems are trained on very large collections of text gathered from the public internet. Benchmarks are published on the public internet, along with discussion of them, worked solutions, and datasets mirrored across repositories. The overlap is not hypothetical and it is difficult to rule out.

If test questions appeared in training data, a high score may reflect memorisation rather than capability, and the two are indistinguishable from the outside. Researchers take this seriously and use various defences — held-out test sets that are never published, canary strings designed to detect inclusion, freshly written problems, and evaluation on material created after a model was trained.

None of these fully settle it. A held-out set stops being held out the moment results are reported often enough for the community to optimise against it, and problems written after a training cutoff test only the models that predate them. Contamination is managed rather than eliminated, and any single score should be read with that in mind.

A measure that becomes a target stops measuring

When a benchmark becomes the accepted way of demonstrating progress, effort flows towards it. This is not cheating; it is what people do when a number determines whether work is recognised, and the same dynamic operates in every field with a league table.

The effect is that scores rise faster than the underlying capability, because some of the improvement is in fitting the particular format, phrasing and answer distribution of the test. A model tuned on many benchmark-style problems gets better at benchmark-style problems, and how much of that transfers is exactly what the benchmark can no longer tell you.

This is why benchmarks age. They get built, they discriminate usefully for a while, scores climb until nearly everything scores near the top, and they stop carrying information. Saturation is the normal end of a benchmark’s life, and the field has cycled through several generations of them.

Aggregation hides more than it shows

Most reported figures are averages over many items of varying difficulty and kind. An average conceals the shape of the distribution: a system that handles most items easily and fails completely on a specific category can produce the same headline as a system that is mediocre throughout.

It also conceals variance. Run the same evaluation twice with different sampling and the score moves, sometimes by more than the difference being claimed between two systems. Reporting a single number without any indication of that spread is common and it makes small differences uninterpretable.

Where the tasks come from matters too. A benchmark assembled from one source, one language and one professional domain measures competence in that context and gets cited as though it measured competence in general. That gap is rarely stated in the table.

What benchmarks are still good for

None of this makes them worthless, and the reflexive dismissal of all benchmark results is as unhelpful as taking them literally. They provide a shared reference point, they catch gross regressions, they let independent parties check claims, and without them comparison would rest entirely on demonstrations chosen by the party making the claim.

The useful posture is to read them as weak evidence about a narrow thing. A large gap on a well-constructed, uncontaminated benchmark tells you something real. A small gap tells you almost nothing. A single aggregate figure tells you less than the same result broken down by category.

The field is aware of all this and disagrees about the remedy. Some favour continuously refreshed private test sets, some favour head-to-head human comparison, some favour evaluating on real deployed tasks with real users. Each approach trades reproducibility against realism, and there is no consensus about the right trade.

Common questions

Why do two published scores for the same system differ?

Because evaluation involves many unreported choices — how the question was phrased, how the answer was extracted, how many attempts were allowed, what sampling settings were used. These details move results substantially, which is why reproducing a published evaluation is harder than it sounds.

Is human preference comparison a better measure?

It captures things automated scoring misses, particularly around usefulness and tone. It also measures what raters like, which correlates with length, formatting and confidence as much as with correctness. It is a different measurement with different blind spots, not a strictly superior one.

Should a buyer trust benchmark results at all?

As a filter, yes; as a decision, no. The only evaluation that reliably predicts performance on your work is an evaluation built from your work, with your data and your definition of a good answer. That is more effort than reading a table, which is precisely why tables get read instead.

Limits & Risksbenchmarksevaluationmeasurementcontamination
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.