How It Works
Learning, in a neural network, means adjusting numbers until the errors shrink
The word was borrowed from human experience and it flatters the mechanism; what actually happens is a very large curve being fitted through a very large number of points.
By Manish Trivedi4 min read

The word was borrowed and it carries baggage
Nothing inside a neural network resembles what a person does when they learn a language or memorise a poem. The word was borrowed early in the field’s history, it stuck, and it has been quietly misleading people ever since. What happens instead is closer to fitting a curve — an enormous curve, through an enormous number of points.
A network is a long chain of arithmetic. Numbers arrive at one end, each layer multiplies them against a stored table of other numbers, applies a simple non-linear squash so the layers cannot collapse into one, and passes the result along. Nothing in that description is intelligent and nothing in it is even mathematically difficult; what is remarkable is only the scale at which it is carried out.
The stored numbers are the parameters, and at the start of training they are essentially noise. Training is the process of nudging them, repeatedly and by very small amounts, until the output at the far end stops being wrong so often. That is the whole mechanism. Everything else is engineering built around it.
Being wrong has to be measured before it can be reduced
A network cannot improve without some definition of failure, and that definition is a single number called the loss. For a system predicting the next piece of text, the loss is broadly a measure of how much probability the model assigned to the piece that actually turned up. Confident and correct scores well. Confident and wrong scores appallingly, and that asymmetry shapes more of the resulting behaviour than most people assume.
The loss is the only account of quality the system has. It knows nothing about truth, usefulness, tone or harm except in so far as those things happen to correlate with the number. If you want a model to be cautious, you have to make carelessness expensive in that score, and finding a way to do that is far harder than it sounds.
This is why so many arguments about model behaviour are really arguments about objectives. Change what gets measured and you change what emerges — but the relationship between the two is loose, indirect, and regularly surprises the people who designed it.
Descent is a very local kind of navigation
Given a loss, the question becomes which way to move each parameter. Calculus supplies the answer: work out, for every parameter, how much the loss would change if that parameter were nudged slightly upward, then move it the other way. Do that for all of them simultaneously and you have taken one step downhill.
Backpropagation is the bookkeeping that makes this affordable. It computes all those sensitivities in a single sweep backwards through the network rather than testing each parameter on its own, which would be hopelessly slow. The idea is neither new nor mathematically deep. Its importance is entirely practical.
The word descent is apt in one respect and misleading in another. Each step genuinely is downhill, but the system can only feel the slope directly beneath its feet. It has no map of the wider landscape, no memory of the ground it has already crossed, and no way of knowing whether a far better region sits somewhere it never walked.
Generalisation is the part nobody has fully explained
A network that merely reproduced its training examples would be useless, and a lookup table would be cheaper and faster. The interesting property is that networks trained on enough data behave sensibly on inputs they have never encountered. They generalise, and exactly why they do it as well as they do remains genuinely unsettled.
Classical statistical intuition says a model carrying vastly more parameters than it has training examples ought to memorise and then fail badly on anything new. Large networks visibly break that intuition, and several competing explanations circulate: implicit regularisation from the training procedure itself, deep structure in the data, or something about the geometry of the loss surface in very high dimensions. Researchers disagree about which matters most, and they do not always disagree politely.
It is worth sitting with that for a moment. The central useful behaviour of the technology is not fully accounted for by theory, which is an odd position for an engineering discipline to occupy while shipping products.
What the metaphor costs
Calling this learning invites two opposite errors. The first is imagining comprehension where there is correlation — assuming that a system which produces a correct explanation must therefore hold the concept the explanation describes. The second is the mirror image: dismissing the whole thing as mere statistics, as though statistics gathered over a sufficiently large corpus could not produce something genuinely capable.
Both errors come from treating the word as a claim rather than a label. A more honest phrasing would be that the parameters were fitted, which sounds bureaucratic, and that is precisely why nobody says it.
The practical consequence is modest but real. When a model behaves strangely, the useful question is almost never what it believes. It is what the training signal rewarded, what the data actually contained, and where the fitted surface stops being supported by any examples at all.
Common questions
Does a model keep learning while people use it?
Not by default. The parameters are frozen once training finishes, and a conversation does not change them — anything the system appears to remember within a session is text being fed back in as input. Providers may later train new versions using collected data, but that is a separate, deliberate process rather than something happening live.
If the loss goes down, is the model getting better?
Better at the thing the loss measures, which is not the same as better overall. A falling loss on the training data while performance on unseen data worsens is the classic signature of overfitting, and a model can also improve its score by learning shortcuts that happen to work on the data it was shown.
Why start the parameters as random noise?
Because if every parameter began identical, every unit in a layer would receive the same update and stay identical forever, so the network could never differentiate. Random starting values break that symmetry. The specific scheme used matters more than it sounds, since a poor one makes the early stages of training unstable.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





