Limits & Risks
The human reviewing the output is doing a harder job than the design assumes
Almost every deployment plan places a person in the loop as the safeguard, and almost none accounts for what happens to that person after a few thousand correct outputs in a row.
By Naina Sethi3 min read

The safeguard everyone reaches for
Ask how an automated system will avoid causing harm and the answer is usually that a qualified person reviews the output before it takes effect. It is a reassuring answer, it satisfies most governance requirements, and it is frequently the least examined part of the design.
The assumption underneath is that the reviewer will catch what the system got wrong. That assumption requires the reviewer to be able to detect the error, to have the time to look, to have the standing to overrule, and to remain attentive over long stretches of routine. Each of those has been studied in other high-automation settings, and none of them holds automatically.
This is not a claim about human weakness. It is a claim about what the arrangement asks of a person, which is often close to the hardest possible version of a monitoring task.
Reliability makes vigilance harder, not easier
A system that is wrong frequently keeps its reviewers sharp, because errors turn up often enough to sustain the expectation of finding one. A system that is right almost all the time trains its reviewers to approve, and the training is not a lapse in discipline but a rational response to overwhelming evidence.
Sustained attention to a task where nothing usually happens is a well-documented difficulty across aviation, industrial monitoring and quality inspection, and the literature is consistent about the direction of the effect. Performance on rare-event detection falls off over a period of watching, and it does not recover simply by trying harder.
There is an uncomfortable implication. Improving the accuracy of an automated system can weaken the effectiveness of the human check placed on top of it, so overall reliability does not improve as much as the component figures suggest.
A confident suggestion anchors the person reviewing it
Reviewing an existing answer is a different cognitive task from producing one. The presented answer supplies a starting point, and the natural mode of engagement becomes looking for reasons it might be right rather than working the problem independently and comparing.
This shows up as two distinct errors that have been documented in various automated settings: accepting an incorrect recommendation, and failing to act when the system stayed silent about something that mattered. The second is easier to miss in audits because there is no artefact to inspect.
Fluency compounds it. An output that is well-organised, appropriately hedged and written with apparent care reads as considered work, and the reviewer’s impression of quality is drawn from characteristics that have no connection to whether the content is correct.
The role often carries responsibility without authority
In many deployments the reviewer is measured on throughput, has less time per item than a genuine review would take, and is aware that overruling the system requires justification while agreeing with it does not. That asymmetry quietly determines behaviour regardless of what the policy says.
There is also the accountability question. When the arrangement exists partly so that a named person is responsible for the outcome, the human in the loop can become a way of placing liability rather than a way of catching errors. The two purposes are compatible in principle and frequently in tension in practice.
It is worth asking of any such design how often the reviewer actually disagrees. If the answer is almost never, the review may be a formality that has been counted as a control.
What tends to work better
Designs that fare better generally reduce what the reviewer has to do. Routing only uncertain or high-stakes cases to a person, rather than everything, concentrates attention where it can matter. Presenting the evidence before the recommendation, or withholding the recommendation until the reviewer has formed a view, reduces the anchoring effect.
Giving reviewers occasional cases with known answers provides an actual measurement of whether the review is functioning, which most deployments do not have. Measuring override rates and investigating when they fall towards zero is cheap and rarely done.
None of this is settled science and the effect sizes vary by domain, so these are directions rather than guarantees. The general point is more robust than any specific remedy: a person placed in a loop is a component with characteristics, and designing as though that component were an infallible backstop is the most common error in the whole area.
Common questions
Does human review satisfy regulatory requirements?
Requirements vary by jurisdiction and sector, and several frameworks ask for meaningful oversight rather than the mere presence of a reviewer. Whether a given arrangement qualifies is a legal question that depends on the details and on where you are, and it is worth getting proper advice rather than assuming a checkbox has been ticked.
Is it better to have no human in the loop?
Usually not, but the choice is rarely that stark. The productive question is what the reviewer is realistically able to catch, and then designing the workload, the presentation and the incentives around that. A review that cannot function is worse than none, because it creates confidence that nothing supports.
Why does fluency affect reviewers so strongly?
Because in ordinary life, care in presentation correlates with care in preparation, and that heuristic serves people well. Generated text breaks the correlation: presentation quality is now independent of accuracy, and a lifetime of reasonable inference stops applying without anyone noticing that it has.
Features writer, AI Worth Knowing
Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.





