Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Learning from consequences is a different arrangement from learning from examples

One family of methods never sees a correct answer at all, only a score arriving after the fact, and everything awkward about it follows from that single missing ingredient.

By Naina Sethi3 min read

A priest leading altar servers during a church ceremony with congregation present.
Photograph by Huynh Van via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

No answer key, only a verdict

The familiar arrangement supplies a model with inputs paired with correct outputs, and training pushes the outputs towards the answers. There is a second arrangement in which no correct output is ever provided. A system acts, something happens, and a number arrives saying roughly how well that went. It must work out the rest by itself.

The vocabulary is deliberately abstract because the setting is general. Something takes actions, the world responds with a new situation and a scalar reward, and the goal is to accumulate as much reward as possible over time. That description covers a board game, a warehouse robot, and a system deciding how much cooling a building needs.

The absence of an answer key is not a minor inconvenience. It means the system cannot be told that its action was wrong and what would have been right; it can only be told the outcome was poor. Discovering what would have been better is its own problem, and it is the hard one.

Rewards arrive late, and that is the central difficulty

In most interesting problems the consequence of an action is separated from the action by a long interval. A move made early in a game determines the loss, and the loss is only observed at the end. Working out which of hundreds of decisions deserves the blame is called credit assignment, and it is what the field is really about.

The standard machinery for this learns a running estimate of how good each situation is, and updates that estimate by comparing it with what actually happened next. Value propagates backwards through experience: a position becomes known as bad because it tends to lead to positions already known as bad. It’s slow, and it works.

The delay also means that a system can learn a correlation that holds in its own experience and nowhere else, having never been shown the counterfactual. Nothing corrects that except more varied experience, which brings its own problem.

Exploration has a cost that supervised learning never pays

To find a better strategy, a system has to try something other than what it currently believes is best, and by definition that trial will usually be worse. Every unit of exploration is paid for in performance. Balancing this against exploiting what is already known is a genuine dilemma with no universally right answer.

In a simulator the cost is trivial: run it a million times more. In a physical system, or one interacting with people, exploration means occasionally doing something bad on purpose to find out what happens. That is where reinforcement learning stops being a technical question and becomes a question about what is acceptable.

This is also why sample efficiency is such a preoccupation. These methods typically need enormously more experience than a supervised method needs examples, because most of what they try teaches them very little.

Games were the natural home for structural reasons

A game supplies everything the arrangement demands. The rules are exact, the outcome is unambiguous, the reward is defined by the rules rather than by anybody’s judgement, and a game can be replayed at whatever speed a computer permits. Self-play adds a further gift: an opponent that improves at exactly the rate you do.

The successes there were real and they were also the easiest possible case. Every property that made games tractable is a property the world lacks. Rules are unstated, outcomes are contested, resets are unavailable, and there is rarely an opponent obliging enough to grow alongside you.

This is why transferring the results outward has been slower than the demonstrations implied. The methods are sound; the conditions are not reproduced.

Specifying the reward is the part that goes wrong

Because the system optimises the number it is given, a reward that imperfectly captures the intent produces behaviour that satisfies the number and violates the intent. This has been observed repeatedly in controlled settings, with systems finding strategies that score well by exploiting the definition rather than by doing the task.

Writing a reward that resists this is much harder than it appears, because the failures are creative and only obvious afterwards. One response is to learn the reward from human judgement rather than writing it down, which moves the problem rather than removing it, since human judgement is itself an imperfect and narrow signal.

The honest position is that this family of methods is powerful, well understood in theory, and awkward in practice for reasons that are structural rather than temporary. Where the environment can be simulated cheaply and the objective can be stated exactly, it is unmatched. Elsewhere it is a research problem.

Common questions

Is this how language models are trained?

Not the main part. The bulk of the training is ordinary prediction over text. A reinforcement-style step is used afterwards to shape behaviour, and it is a heavily constrained special case of what is described here rather than the general method.

Why is it described as trial and error if there is a model involved?

Because the source of information genuinely is trial and error — the system finds out by doing. The learned model is what accumulates and generalises across those trials, so that the same mistake need not be made in every situation that resembles it.

Can a system learn from recorded experience rather than acting?

There is an active line of work on exactly that, using logs of decisions somebody else made. It sidesteps the cost of exploration and inherits a hard problem in exchange: the data only shows what happened after the actions that were actually taken.

How It Worksreinforcement learningtrainingreward
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.