History
The training algorithm everyone uses was discovered more than once
The method for assigning credit through the layers of a network arrived independently in several fields over roughly two decades, was ignored each time, and only mattered when somebody paired it with a demonstration.
By Naina Sethi3 min read

The idea is a bookkeeping trick, not a theory
To improve a layered network you need to know how much each parameter contributed to the error. The chain rule of calculus gives the answer, and the only difficulty is doing it efficiently for a system with a very large number of parameters rather than one at a time.
The efficient version works backwards from the output, reusing the calculations from each layer as it goes, so the total cost is comparable to running the network forwards. Stated that way it is a technique for organising a computation, and the mathematics involved is taught to undergraduates.
That modesty is exactly why the history is tangled. Nobody looking at the method in isolation would identify it as the key to anything, and several people who described it did not present it as important. It reads as a procedure for arranging derivatives conveniently, which is precisely what it is, and the significance lies entirely in what becomes possible once the arrangement is cheap enough to run millions of times.
It arrived through control theory first
Optimising a system with many stages, where each stage feeds the next, is a standard problem in optimal control, and methods for propagating sensitivities backwards through the stages were developed there in the 1960s. Aerospace trajectory optimisation was the motivating application.
A more general formulation appeared in the early 1970s in work on automatic differentiation, describing how to compute derivatives of a composed function efficiently in reverse. This is the same algorithm, stated in a form that has nothing to do with neural networks and applies to any composition of differentiable operations.
A doctoral thesis in the mid 1970s made the connection explicitly, applying the method to networks and to problems in forecasting. Almost nobody in the neural network community read it at the time, partly because there was barely a neural network community to read it.
Nobody connected the fields
Control theorists, statisticians and the small group interested in learning machines published in different venues, used different notation, and had little reason to read each other. A result stated as an efficiency improvement for computing derivatives does not announce itself to somebody thinking about brains.
The other obstacle was equipment. The method needs to be run for many iterations on a network large enough to do something interesting, and the hardware of the period made that unaffordable. An idea that cannot be demonstrated persuades very few people, whatever its merits on paper.
This is a recurring shape in the field’s history. Ideas sit dormant not because anybody rejected them but because the conditions for showing they work do not yet exist, and the conditions usually arrive from an unrelated industry.
The paper that made it matter did something different
The version that finally landed appeared in the mid 1980s in a prominent general science journal, presented alongside demonstrations that networks trained this way developed useful internal representations of the problem — not merely that the method worked, but that something interpretable had been learned.
That framing was the contribution. It answered the objection that had shelved the field a generation earlier by showing that hidden layers could be trained at all, and it did so with examples that other researchers could reproduce and build on.
It also arrived into a community ready to receive it, with a growing interest in distributed models of cognition. Timing did a great deal of the work, as it usually does when an old idea suddenly takes.
What the credit disputes actually show
Arguments about who deserves priority here have been long-running and occasionally sharp, conducted in public, and they are unlikely to be settled. All the claims have documentary support. What differs is whether one counts describing a method, applying it to this class of problem, or persuading a field to adopt it.
A reasonable view is that these are three distinct contributions and that conflating them is the source of the disagreement. Independent discovery is normal in mathematics, and it is more common the closer an idea sits to a standard technique. This one sits very close.
The practical lesson is worth more than the scorekeeping. A method can be correct, published and available for years without having any effect, because effect requires a demonstration, an audience and hardware. Being first is not the same as being the reason something happened.
Common questions
Is backpropagation how biological brains learn?
Almost certainly not in this form. The algorithm requires an exact backward pass using the same connection strengths as the forward pass, which does not correspond to anything known in neural anatomy. Researchers have proposed biologically plausible approximations, and none is established as what actually happens.
Has the method changed much since?
The core calculation has not. What surrounds it has changed considerably: how the update size is chosen, how gradients are stabilised across many layers, how the work is distributed across processors, and how numerical precision is managed. The chain rule underneath is the same.
Why do modern frameworks make it invisible?
Because they implement the general reverse-mode differentiation the method belongs to, applied automatically to whatever computation you write. That generality is the reason a researcher can define a new architecture without deriving any of the derivatives by hand, which changed the pace of experimentation substantially.
Features writer, AI Worth Knowing
Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.





