Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Using one model to grade another is convenient, and it inherits every bias of the grader

Judging open-ended output automatically has become standard practice because the alternative is expensive, and the shortcut carries systematic distortions that are only partly understood.

By Daniel Okonkwo4 min read

Close-up of JavaScript code on a computer screen, showing web development programming.
Photograph by Marek Prášil via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The measurement problem that created the practice

Anything with a single correct answer can be scored mechanically. Open-ended output cannot: there is no list of acceptable summaries of a document, no enumeration of good replies to a difficult message. For years this was handled by paying people to read output and rate it, which is slow, costly and hard to keep consistent.

Development cycles do not tolerate that. A team changing something wants to know within hours whether the change helped, across thousands of cases, repeatedly. Human evaluation cannot supply that rhythm at any affordable price, and the gap between what development needs and what evaluation can provide had become the binding constraint.

So the field did the obvious thing and asked a capable model to do the rating. Give it the input, the output and a description of what good looks like, and collect a score. It’s fast, it’s cheap, it runs whenever you like, and it correlates with human ratings well enough in published comparisons to have become normal practice.

What the arrangement quietly assumes

The assumption is that evaluating output is easier than producing it, which is often true for people and is not obviously true here. A grader model has exactly the same limitations as a generator model, because in most setups it is one. It has no access to ground truth it did not already have, and no mechanism for noticing that it is wrong.

A grader that cannot verify a factual claim will grade on the qualities it can assess, which are the surface ones: fluency, structure, confidence, apparent thoroughness. Those correlate with quality across ordinary cases and come apart exactly where an evaluation is most needed, which is on the difficult and unusual items.

This is the same failure the underlying technology has everywhere else, applied to the instrument being used to detect it. Nothing about relabelling a model as a judge changes what it is doing.

The distortions that have been documented

Position effects are well attested: when two candidate outputs are compared, which one appears first influences the verdict, and controlling for it requires running the comparison both ways. That anybody had to discover this is itself informative about how the technique was adopted.

Length is a persistent one. Longer output tends to be scored more highly than shorter output of equal or better substance, which pushes any system tuned against such a grader towards verbosity. Anyone who has noticed assistants padding straightforward answers is observing an optimisation pressure with a plausible origin here.

There is also evidence of models rating output from their own family more favourably, which matters a great deal when the same organisation builds both the system and the evaluation. Formatting, confident phrasing and a familiar register all appear to shift scores independently of content.

Optimising against a flawed grader is worse than using one

Measuring with an imperfect instrument introduces noise. Training against that instrument introduces something worse: a systematic search for whatever the instrument rewards. Any consistent quirk in the grader becomes a target, and the resulting system is genuinely better at satisfying the grader while not necessarily better at anything else.

This is the general problem of a measure becoming a target, in a form where the measure is cheap enough to be applied millions of times. Scale is what makes it dangerous. A biased human rater produces a small drift; a biased automatic rater applied throughout a training loop produces a strong and specific pull.

The effect is difficult to detect from inside the loop, because the metric being used to check for problems is the one with the problem. It becomes visible when somebody runs a genuinely independent evaluation, which is exactly the expensive step the arrangement was designed to avoid.

What makes it more defensible

The practice works considerably better when the judgement is decomposed. Asking whether a specific claim appears in a supplied source is a narrow question with a checkable answer; asking whether a response is good is not. Structured rubrics with concrete criteria produce more stable results than a general quality score.

Grounding helps for the same reason. A grader given reference material, or a known correct answer, or the output of a program that can verify something, is doing a comparison rather than exercising taste. That is a much more reliable use, and it is available more often than teams assume.

Periodic human calibration is the other half. Sampling a portion of automatic verdicts and having people check them establishes whether the correlation still holds, which is not something to be assumed once and forgotten as systems change.

An honest statement of where this stands

The technique is useful and it is not a measurement in the sense the word usually carries. Treating it as a fast, biased indicator that catches large regressions is reasonable. Treating a small difference in an automatic score as evidence that one system is better than another is not, and published comparisons do this routinely.

The field is aware of the problem and there is genuine disagreement about severity. One position holds that the correlations with human judgement are strong enough for practical purposes and that the biases are correctable. Another holds that the entire arrangement is circular in a way no amount of correction fixes.

The unglamorous conclusion is that a claim of improvement means something quite different depending on how it was measured, and the measurement method is often the part of a result least examined.

Common questions

Does using a different model as the judge solve the bias?

It removes the self-preference effect and leaves the others, since position, length and style effects appear across models rather than being specific to one. Using several judges from different families and looking at their disagreement is a stronger arrangement than switching to a single alternative.

Is human evaluation reliable by comparison?

Not straightforwardly. Human raters disagree with each other, are affected by fatigue and presentation, and are influenced by length and confidence too. The argument for human evaluation is that its biases differ from the system being tested, not that it is unbiased.

Why not just use benchmarks with correct answers?

Where a task has correct answers, that is the better instrument and it is generally used. Automatic judging exists for the large category of tasks where no such answer set can be written down, which is most of what these systems are actually asked to do.

Limits & Risksevaluationvalidationbias
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.