Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Sound and images reached the same toolkit by completely different roads

Speech recognition and computer vision were separate disciplines with separate assumptions for decades, and how each converted its raw signal into something learnable still explains what they do differently.

By Manish Trivedi4 min read

A complex network of cables in a data center with a monitor in the foreground.
Photograph by panumas nikhomkhai via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two problems that look alike only from a distance

It is tempting to treat hearing and seeing as one problem, since both involve a machine making sense of a messy physical signal. The research communities did not see it that way. For most of their history speech and vision had different conferences, different mathematics, different benchmark tasks and largely different people, and the assumptions each built in are still visible in how the systems behave.

The underlying signals genuinely differ. Sound is a one-dimensional pressure wave unfolding strictly in time, where order is everything and nothing can be revisited. An image is two-dimensional and simultaneous, with no inherent direction, and its meaningful relationships are spatial rather than sequential.

Those differences dictated the early engineering, and the early engineering shaped decades of work afterwards. Only relatively recently did both fields end up using broadly the same machinery, which is a convergence worth understanding rather than assuming.

Speech had to be turned into a picture of itself first

A raw audio waveform is a very long sequence of amplitude measurements, and almost nothing that matters perceptually is legible in that form. So the standard first move was to chop the signal into short overlapping windows and compute, for each window, how much energy sat in each frequency band. The result is a two-dimensional map of frequency against time.

That representation was then compressed further using transformations designed around known properties of human hearing, particularly the fact that we discriminate pitch far more finely at low frequencies than at high ones. The resulting features dominated speech systems for a very long time, and they were engineered rather than learned.

What followed was statistical modelling of how sounds succeed one another, since speech is a sequence with strong ordering constraints. This is part of why speech recognition became an early and decisive proving ground for probabilistic methods: the problem had temporal structure that statistics handled naturally and hand-written rules did not.

Vision spent decades hand-designing what to look for

The equivalent move in vision was to design detectors for image features a person could describe — edges, corners, gradients, textures, regions of consistent colour — and then build recognition on top of those descriptors. Enormous ingenuity went into features that stayed stable when an object was rotated, rescaled or lit differently.

This was a reasonable programme and it produced genuinely useful systems. Its limitation was that each new task tended to need new features, designed by somebody who understood both the mathematics and the specific domain, and that expertise did not transfer freely between problems.

The shift came from letting the features be learned instead, using layers that apply the same small detector at every position in the image. That sharing encodes a real assumption about the world: that a pattern means roughly the same thing wherever it appears in the frame. For photographs the assumption holds well. It is not universally true, and where absolute position carries meaning it can be a handicap.

Convergence happened, but it is less complete than it looks

Both fields now largely use architectures built around the same sequence-modelling mechanism, applied to audio slices in one case and image patches in the other. The engineering vocabulary has merged, people move between the areas, and a single system can increasingly take text, sound and pictures together.

Underneath, the differences persist. Speech remains strictly sequential and must handle timing, overlap and the fact that identical words carry different meaning depending on delivery. Vision must handle occlusion, perspective and the flattening of a three-dimensional world before the model ever sees it. Neither of those problems is the other one.

There is also a practical asymmetry in data. Images with useful text descriptions exist in vast quantity because people caption photographs online. Transcribed speech is comparatively scarce and expensive to produce, and that shapes what each field can attempt.

What the shared toolkit did not equalise

Because the two lineages converged late, their failure modes remain distinct. Vision systems fail on unusual viewpoints, unfamiliar lighting and objects outside the range they were shown. Speech systems fail on accents, overlapping speakers, background noise and the sort of disfluent real conversation that clean transcripts under-represent.

Both inherit the character of their training material, which is a general property of the technology rather than a quirk of these two domains. But the material differs so much in kind that the resulting weaknesses do not line up, and a benchmark designed for one gives you no information about the other.

The useful takeaway is a caution about the word multimodal. A single model accepting several kinds of input is not a single model that understands them equally well. It is several lineages of engineering sharing a chassis, and their separate histories are still doing work.

Common questions

Why not feed raw audio straight to a model?

Some systems now do, and it works, but it is expensive: raw audio has an enormous number of samples per second and most of the variation is perceptually irrelevant. Converting to a frequency representation first discards a great deal that does not matter and makes the learning problem substantially smaller.

Is a model that handles text, images and audio really one model?

Partly. Typically each input type is converted into the same kind of numerical representation by its own component, after which a shared core processes them together. That sharing is genuine and does produce useful cross-modal behaviour, but the front ends remain specialised and so do their weaknesses.

Why do speech systems struggle with accents they were not trained on?

Because the acoustic patterns of an unfamiliar accent sit outside the range the model was fitted to, and nothing in the system flags that it is operating outside its evidence. The result is a confident transcription that happens to be wrong, which is the same failure shape seen throughout the field.

How It Worksspeechcomputer visionsignalsarchitecture
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.