Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A network for reading handwriting worked commercially long before anybody called it a breakthrough

The architecture behind the vision boom was designed, trained and deployed in a working industrial system decades before the field decided the approach was viable.

By Imran Sheikh4 min read

Vintage office environment with a man working on a classic CRT monitor and desktop computer setup.
Photograph by MART PRODUCTION via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

An idea taken from the anatomy of seeing

Work in the middle of the last century on how the visual cortex of a cat responds to stimuli found cells reacting to simple features in a small patch of the field of view, with later cells combining those responses into something less local. The organisation was hierarchical, and each stage looked at a limited region.

That description was taken up by researchers building artificial systems, and an architecture published around 1980 by a Japanese researcher implemented the idea directly: layers of local feature detectors, arranged so that a feature recognised in one place could be recognised in another. It anticipated most of what would matter later.

What it lacked was a way to train itself from examples. The features had to be arranged largely by hand, which limited how far it could be pushed. The missing ingredient was a training method that could assign credit through many layers, and that method existed and had not yet been connected to this problem.

Weight sharing was the decisive design choice

The insight that made the architecture practical was that a detector for an edge is the same detector wherever the edge appears, so the same small set of parameters can be applied across the whole image rather than learned separately at every position. This is called weight sharing and it does two things at once.

It reduces the number of parameters enormously, which matters when data and computation are scarce. And it builds in an assumption about the problem: that position should not change identity. A network that begins with this assumption does not have to spend data discovering it, which is why it learns from far fewer examples than an unstructured network.

This is an unusually clear example of an architecture encoding knowledge about a domain. Nothing about the training procedure changed; what changed was the shape of the thing being trained, and the shape carried a fact about pictures.

It went into production, and the field did not follow

By the 1990s a network of this design was reading handwritten digits well enough to be used in the automated processing of cheques, which is a demanding application: the input is genuinely messy, the volume is large, and an error costs money. It handled a meaningful share of that work in the United States.

This was a deployed, commercially consequential neural network operating at a time when the prevailing view held that neural networks were an interesting failure. The result was known, published and largely treated as a special case rather than as evidence about the approach in general.

The explanation for that reception is not that anybody was foolish. Scaling the same architecture to harder problems did not work with the data and hardware available, other methods with firmer theoretical grounding were performing better on most benchmarks, and a technique that works on one narrow task is a weak argument for a research programme.

What eventually changed was supply, not insight

The revival came when three things arrived together: a large labelled image collection assembled through distributed annotation, processors designed for graphics that happened to suit the arithmetic, and a set of practical training refinements that made deeper networks trainable. None of these was a new idea about vision.

When those were combined and the architecture was scaled up, results on a shared image benchmark improved by a margin large enough that the field reorganised around it within a couple of years. The architecture being scaled was recognisably the one from the cheque readers.

This is why the episode is told as a story about persistence, which is only partly right. The approach was not vindicated by argument or by anybody holding their nerve; it was vindicated when the inputs it needed became available, and those became available for reasons that had nothing to do with vision research.

The lesson usually drawn is too comforting

The version told at conferences is that a neglected idea was correct all along and the field should be more open to unfashionable work. That contains something true and it flatters everybody currently working on an unfashionable idea, most of whom are working on something that will not turn out to be correct.

A more careful reading is that an approach can be right in principle and unusable in practice for decades, and that distinguishing the two from the inside is extremely hard. The evidence available in the intervening years genuinely did not support the approach; it was not being ignored irrationally.

The uncomfortable implication is that the current set of neglected ideas contains some that will look obvious later and many more that will not, and no method exists for telling them apart in advance. The history is a caution against confidence in both directions.

Common questions

Was this architecture the first neural network used commercially?

It is among the earliest with a well-documented industrial deployment at scale, though smaller applications of neural methods existed in various industries earlier. The claim worth making is that it was a substantial working system, not that it was uniquely first.

Why did the same design not work on harder images immediately?

Larger and more varied images require a deeper network, deeper networks were difficult to train with the methods of the time, and the labelled collections needed to train them did not exist. Each of those obstacles was eventually removed by separate developments.

Are these networks still the standard for images?

They remain widely used and have been partly displaced by architectures adapted from language work, which scale better with very large datasets. The convolutional design retains advantages where data is limited, because the assumption it builds in is genuinely useful.

Historyneural networkscomputer visionhistory
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.