Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Chain enough reliable steps together and the chain becomes unreliable

Systems that plan and act over many steps inherit an arithmetic problem: per-step accuracy that sounds excellent produces end-to-end results that are not, and the errors do not stay where they started.

By Samar Bhatia3 min read

Senior man looks away from an unemployment notice on a computer screen indoors.
Photograph by Ron Lach via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The arithmetic is unforgiving and rarely stated

Consider a task decomposed into a sequence of steps, where each step must succeed for the final result to be correct. If the steps fail independently, the probability of overall success is the per-step reliability multiplied by itself once for every step, and that product falls away much faster than intuition suggests.

A rate that reads as impressive in isolation becomes ordinary across five steps and poor across twenty. This is not a criticism of any particular system. It is a property of chains, familiar to anyone who has designed a manufacturing line or a distributed system, and it applies regardless of how the individual steps are implemented.

The reason it catches people out here is that demonstrations show single steps. A model calling a tool correctly is convincing. Thirty such calls in sequence, each depending on the last, is a different proposition entirely.

Errors in a chain do not merely accumulate, they propagate

The independence assumption above is generous. In practice a mistake at step four becomes the input to step five, so the later steps are now operating on a false premise and doing so confidently. Their outputs are not merely wrong; they are elaborations of the original error, made plausible by the surrounding structure.

This is why failures in multi-step systems can look so strange. The system does not stop and report confusion. It continues reasoning fluently from a wrong starting point, generating justifications along the way, and the resulting trace can be internally coherent from the point of divergence onward.

There is a specific hazard where a step involves the system reading something it produced earlier. Any error is then re-ingested as apparent evidence, which is a positive feedback loop of exactly the kind engineers usually design carefully to avoid.

Recovery requires knowing that something went wrong

Robust systems handle failure by detecting it. A network request that fails returns an error; a database transaction that cannot complete rolls back. The step reports its own failure, and the surrounding logic can respond.

A language model producing a wrong answer reports nothing. The output is well-formed and confident whether or not it is correct, so there is no signal for the orchestration layer to act on. This is the same problem as unwarranted confidence, but its consequences multiply when the output feeds another step rather than reaching a person who might notice.

Where an external check exists the situation improves dramatically. Code that must compile, a query that must return rows, an arithmetic result that can be recomputed — these give a step a genuine pass or fail, and reliability improves accordingly. That is why tasks with a hard verifier work so much better than tasks without one.

The mitigations help and none of them is a solution

Shorter chains are the most effective response and the least discussed, since the appeal of these systems lies in long autonomous runs. Checkpoints where a person confirms before proceeding work well and reintroduce the human bottleneck the automation was meant to remove. Redundant attempts with a vote reduce random error while doing nothing about a systematic misunderstanding, because the repeats share the mistake.

A separate model reviewing the work is the most popular approach and the most oversold. It catches some errors and it shares the underlying tendencies of the model it is checking, so the errors it misses are precisely the ones both find plausible. Correlated reviewers provide much less assurance than their number suggests.

Constraining actions so that irreversible ones require confirmation does not improve reliability at all. It limits the damage from unreliability, which is a different and often more sensible goal.

What this predicts about where autonomous systems will work

The pattern is reasonably clear from first principles. Long chains work where steps can be verified cheaply and where errors are recoverable — drafting, exploring, generating candidates that a person or a test will filter. They work poorly where verification is expensive and mistakes are expensive too.

This is a claim about the current shape of the technology rather than a permanent one. If per-step reliability improves substantially, the arithmetic becomes friendlier and longer chains become practical. The counter-case is that improvements in per-step reliability have been steady rather than dramatic, and the length of the chains people want has grown at least as quickly.

The safer conclusion is a question rather than a forecast: before trusting a long autonomous run, ask what would detect a failure at step seven, and how much the answer costs.

Common questions

Why do demonstrations look so much better than deployments?

Demonstrations usually show short chains on familiar tasks, often after several attempts, with a knowledgeable person steering. Deployment involves long chains on unfamiliar variations with nobody watching each step. The gap is not deception; it is the difference between a single step and a product of many.

Does a bigger model fix compounding error?

It raises per-step reliability, which helps because the effect is multiplicative, but it does not change the structure of the problem. A better model chained across more steps can easily end up in the same place, and the ambition for these systems has grown alongside their capability.

Is a review step by another model worth adding?

Often yes, provided the expectations are right. It catches a useful fraction of errors, particularly obvious ones, and it should not be treated as independent verification because the reviewing model shares training, tendencies and blind spots with the one it reviews.

Limits & Risksagentsreliabilitycompoundingautomation
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.