Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Recognition systems perform unevenly and the unevenness is not random

Speech and image systems are described by a single accuracy figure, and that figure conceals which people the system works for and which people it does not.

By Imran Sheikh3 min read

Detailed view of robotics components and tools on a workbench, showcasing innovation and technology.
Photograph by Tanha Tamanna Syed via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

An average is a summary of a population

When a recognition system is described as accurate to some percentage, that number is the result of running it over a test set and counting mistakes. It is an average, and an average tells you nothing about how the errors were distributed among the people who made up the test.

A system can post an excellent overall figure while failing substantially for a subgroup, provided that subgroup is a small share of the test data. Nothing about the headline number reveals this, and until relatively recently very few published evaluations broke results down at all.

This is not a subtle statistical trap. It is the same reason a national average income tells you little about any particular household, and it applies with full force the moment a system is deployed somewhere whose population does not match the test set.

Speech carries far more variation than the training data does

A speech recogniser learns the mapping from sound to text from recordings, and recordings differ along more dimensions than most people consider. Accent and dialect are the obvious ones. Speaking rate, pitch range, age, whether the speaker has a cold, whether they are reading or talking, background noise, room acoustics and microphone quality all matter substantially.

Training corpora historically over-represented certain kinds of speaker, certain recording conditions and certain speaking styles, because those were the recordings that were easy to gather. The resulting systems worked best on speech that resembled the corpus and degraded on speech that did not, in a pattern that turns out to track social categories quite closely.

Vocabulary compounds this. A recogniser has strong prior expectations about what words are likely, and unusual names, regional terms and technical vocabulary are all improbable by default. The system does not fail loudly on these; it substitutes something more common and moves on.

Image systems inherit assumptions from the camera onwards

The equivalent story in vision starts before any model is involved. Camera exposure and colour processing have historically been tuned against reference standards that did not represent everyone equally, so the input itself is not neutral. A model trained on those images inherits the problem and then adds its own.

Lighting conditions, image resolution, angle and occlusion all shift performance, and again the shifts are not evenly distributed across the people being photographed. Evaluations conducted under controlled conditions consistently overstate performance in the field, where nothing is controlled.

The consequences depend enormously on what the output is used for. A photo application that mislabels something is an annoyance. The same underlying error rate feeding an access control system, an identification process or an enforcement decision is a different matter entirely, and the technology cannot tell which situation it is in.

Deployment multiplies whatever the error rate is

Two factors turn a modest error rate into a serious problem. The first is base rate: when the thing being searched for is rare, even a very low false-positive rate produces far more false alarms than genuine finds, simply because the system is applied to so many negatives. This arithmetic surprises people repeatedly and it is not intuitive.

The second is the threshold. Every recognition system produces a score and somebody chooses the cut-off, trading false positives against false negatives. That choice is a policy decision dressed as a configuration setting, and it is usually made by whoever installed the system rather than by anyone accountable for the consequences.

Add automation bias — the well-documented tendency of people to defer to a machine output they were nominally supposed to check — and the human review step that justified the deployment stops functioning as a safeguard.

What has changed and what has not

Serious work has gone into this. Disaggregated evaluation is now expected rather than exceptional in research, datasets have been rebuilt with deliberate attention to composition, and several procurement frameworks require subgroup reporting before a system can be bought. Measured gaps have narrowed in a number of published evaluations.

Narrowed is not closed, and the underlying dynamic persists: performance tracks representation in the training data, and representation tracks who was easy to record. Every new domain, language and deployment context reopens the question rather than inheriting the fix.

There is also a real argument about whether some applications should exist at whatever accuracy. Improving a system and deciding whether to deploy it are separate questions, and it is possible to hold that a technology has become much better and that certain uses of it remain unjustified.

Common questions

Why does a system work well in a demonstration and badly in use?

Demonstrations happen in conditions that resemble the training data — decent microphone, quiet room, cooperative speaker, good lighting. Real deployment introduces variation the system never saw. This gap between benchmark conditions and field conditions is one of the most consistent findings in applied machine learning.

Does more training data fix uneven performance?

More data of the right kind helps considerably. More data of the same kind mostly does not, and can entrench the imbalance further, since the majority pattern gets even better represented. Composition matters more than volume once you are past a certain size.

Is a human reviewer an adequate safeguard?

Only if the reviewer has the time, the information and the standing to disagree. Research on automated decision aids consistently finds that people accept machine outputs more readily than the accuracy warrants, particularly under time pressure. A review step that exists on paper is not the same as one that functions.

In The Worldspeechrecognitiondatasetsaccuracy
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.