Limits & Risks
The expertise needed to supervise a system is built by doing the work it replaced
Handing a task to a tool removes the practice that produced the judgement required to check the tool, and the loss is invisible for as long as nothing unusual happens.
By Imran Sheikh4 min read

Judgement is a by-product of repetition
Somebody who has done a task many times acquires something that cannot be conveyed in a description of the task. They notice when a case is unusual before they can say why. They know which parts of a problem are commonly got wrong. They have a sense of what a reasonable answer looks like, calibrated against several thousand actual ones.
That sense is what makes review possible. Checking work requires having a prior expectation to compare against, and the expectation comes from having produced the work yourself often enough for the pattern to settle. It’s not transferable by instruction, which is why apprenticeship survives in fields that have tried to replace it.
Automating the task removes the repetition and keeps the requirement. The person is still asked to check, and the mechanism by which they became able to check has been withdrawn. This is not a hypothetical concern; it is the standard finding from decades of studying automation in other settings.
The pattern was described long before this technology
Research into automated systems in aviation and process control identified this in the last century and named the resulting arrangement an irony: the more reliable the automation, the more the human role narrows to handling the rare cases, and the less prepared for those cases the human becomes.
The response in those industries was not to abandon automation, which would be absurd, but to build the maintenance of skill into the job. Practising manual operation deliberately, running scenarios that do not occur in normal service, and treating currency as something requiring active upkeep rather than something possessed once.
That response has costs and it is defended because the alternative was demonstrated to be worse. It is worth noting that these are industries where failure is dramatic and investigated thoroughly, which is how the effect became well documented rather than merely suspected.
Where the cost lands first
The effect is unevenly distributed and it falls hardest on people who have not yet acquired the judgement. An experienced practitioner using a tool retains what they already built. Somebody starting out, whose formative years consist of reviewing generated output rather than producing their own, may never build it.
This is a specific worry in fields with a defined path from novice to expert, where the early work being automated is precisely the routine work that trained people. The tedious tasks were not only tedious; they were the mechanism by which competence accumulated, and nobody designed them for that purpose.
What replaces them is not obvious. Deliberately having people do work a machine could do, purely as training, is expensive and organisations rarely sustain it. Whether the profession adapts by developing new formative practices is a genuine open question rather than something to be assumed.
Two effects that are easy to confuse
One is short-run: attention degrades during a session of reviewing output that is nearly always correct, so errors are missed. The other is long-run: the underlying ability to evaluate the work atrophies, or never develops, so errors are missed even when attention is perfect.
They call for different remedies. The first responds to workload design, rotation and sampling. The second responds only to practice, and no amount of interface design or procedural care substitutes for it. Confusing the two produces plans that address the easier problem and leave the harder one untouched.
The long-run effect is also much harder to observe, because it emerges over years and is confounded with everything else changing in a profession over the same period. Absence of evidence here should be read carefully.
The counter-case is real and should not be waved away
Every technology that removed a skill was accused of this, and in most cases the skill genuinely was lost and the loss turned out not to matter. Very few people can navigate by the stars or perform long division at speed, and neither absence has proved consequential. Skills become obsolete, and mourning all of them equally is not a serious position.
There is also a case that these tools raise the floor more than they lower the ceiling: a less experienced person supported by a good tool may outperform their unaided self substantially, and studies of individual tasks have found effects in that direction. Access to competence is a real benefit, not a consolation.
The distinction that matters is whether the skill is needed to supervise the automation. Navigation by stars is not required to check a satellite fix. Clinical judgement is required to check a clinical recommendation. Where the skill is the check, losing it removes the safeguard the whole arrangement depends on.
What follows from taking it seriously
It suggests that the question to ask about a deployment is not only whether the tool performs well, but what happens to the capability of the people around it over several years. That is a slower and less satisfying question than a benchmark score, and it is the one that determines whether the safeguard still exists later.
It also suggests scepticism about arrangements that place a nominal reviewer in a role they have no way to grow into. A checking step performed by somebody who could not have produced the work is a formality, and describing it as human oversight in a governance document does not make it one.
None of this argues against the tools. It argues that skill is a maintained asset rather than a permanent one.
Common questions
Is there hard evidence of this happening with current tools?
There is well-established evidence from earlier automation in other industries and early studies in some professional settings, and it is too soon for the long-run effect to have been measured here. The argument rests on a documented general pattern rather than on direct measurement of this case.
Does using a tool as an assistant rather than a replacement avoid it?
It helps, because the person remains engaged in producing the work rather than only in checking it. The boundary is slippery in practice, since assistance tends to expand towards replacement wherever the output is acceptable.
Who should be responsible for maintaining the skill?
In industries that took the problem seriously it became an employer and regulator responsibility rather than an individual one, because an individual has no incentive to practise a task their organisation has automated. How that translates elsewhere is unsettled.
Editor, AI Worth Knowing
Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.





