Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Recognising speech took fifty years and three complete changes of method

The task was attempted almost as early as anything in the field, resisted every approach tried on it, and its eventual solution arrived by discarding the machinery that had made partial progress possible.

By Imran Sheikh3 min read

Close-up of a vintage IBM circuit board showcasing electronic components and retro technology.
Photograph by Nicolas Foster via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The problem is harder than it sounds and was assumed to be easier

Speech looks like an obvious target. People do it effortlessly, the signal is available, and the output is text, which computers already handle. Early efforts accordingly promised a great deal, and the systems produced could recognise a small vocabulary from a single cooperative speaker in a quiet room.

Everything about that description is a constraint being hidden. Words in continuous speech are not separated by silence; the boundaries a listener perceives are largely constructed. The same word is acoustically different depending on what surrounds it, who is speaking, how fast, and what they are doing with their face at the time.

Background noise, overlapping speakers and unfamiliar accents each defeat approaches that work in the laboratory. The gap between a demonstration and a usable system was larger here than almost anywhere, and it took decades to close even partially.

Template matching gave way to a statistical framework

The first approaches compared an incoming signal against stored patterns, with elaborate techniques for stretching and compressing time so that a slowly spoken word could match a quickly spoken template. This works for isolated words and small vocabularies and does not extend, because the number of templates required grows unmanageably.

The replacement was a probabilistic model treating speech as a sequence passing through hidden states, where the states correspond roughly to speech sounds and the observed signal is generated from them with some uncertainty. Combined with a separate model of which word sequences are likely in the language, this framework dominated for roughly two decades.

It succeeded because it handled variation as a first-class property rather than as noise to be eliminated. It also required careful engineering of what features to extract from the audio, and much of the practical progress in this era came from that engineering rather than from the statistical framework itself.

Organised evaluation shaped the field more than any single method

Government-funded evaluation programmes, particularly in the United States, ran common tests on common data and published the comparison. Participation was a condition of funding for many groups, which meant methods were compared directly and repeatedly on tasks nobody could tune in private.

This arrangement is now standard across machine learning and it was unusual then. It produced steady, measurable improvement and it also produced the characteristic distortion of any shared benchmark: effort concentrated on the specified task, and progress on unmeasured aspects of the problem was slower and less visible.

The tasks were also chosen for tractability, which meant read speech before conversational speech, and clean recordings before realistic ones. Each successive evaluation moved the target closer to reality, deliberately and slowly.

Neural methods entered gradually and then replaced everything

Neural networks were tried on speech repeatedly across the decades and generally did not beat the statistical framework. What changed the picture was using them for part of the job — estimating the probabilities within the existing framework — which produced a substantial improvement while keeping the surrounding machinery.

That hybrid arrangement was itself displaced by end-to-end systems mapping audio directly to text with no explicit representation of speech sounds, no separately engineered features and no hand-built pronunciation dictionary. Components that had absorbed decades of expertise were removed, and results improved.

This has a familiar shape and it should be stated carefully. The hand-built components were not useless; they carried the field for decades and made the data-hungry approach possible by establishing what the task was and how to measure it. They were replaced when enough data existed to learn what they encoded.

What was achieved and what the achievement conceals

Contemporary systems transcribe clear speech in widely spoken languages at a quality that made dictation, captioning and voice interfaces ordinary. That is a genuine and hard-won result, arrived at through three distinct methodological eras rather than through one insight.

The remaining difficulties are the ones that were always hardest. Heavily accented speech, children, atypical speech, several people talking at once, poor recordings and languages without large transcribed collections are all handled substantially worse, and the single headline accuracy figure conceals all of it.

The pattern across the whole history is worth noting: each era solved the version of the problem it could measure, and the parts left over were the parts nobody had built a benchmark for. That is not a criticism of the researchers so much as a description of how organised measurement steers a field.

Common questions

Why did speech take so much longer than text processing?

Because the input is a continuous signal with no natural boundaries and enormous variation between speakers, whereas text arrives already segmented into discrete symbols. Converting sound into something with the properties of text was itself most of the problem.

Did the older statistical framework disappear entirely?

From the frontier, largely yes, though the vocabulary and the evaluation practices it established remain in use. It also survives in constrained settings where computation is limited and a small vocabulary is sufficient.

Is transcription now a solved problem?

For clear speech in well-resourced languages it is good enough for most practical purposes, which is not the same as solved. Performance degrades sharply outside those conditions, and the degradation is not visible in the figures usually quoted.

Historyspeech recognitionhistorystatistical methods
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.