In The World
A regulated setting asks questions a demonstration never has to answer
Medicine, aviation, finance and public administration each built approval processes around tools that do not change, and a system whose behaviour was fitted to data sits awkwardly against every one of them.
By Zoya Rahman3 min read

Approval procedures assume a fixed device
The machinery of certification in safety-critical fields grew up around instruments, drugs and mechanical parts. You specify what the thing does, you test whether it does that, and then it keeps doing it. The specification is the anchor, and everything downstream — testing, labelling, liability — hangs off it.
A learned system frustrates this at the first step, because nobody can write a complete specification of its behaviour. What exists instead is a description of how it was built and a record of how it performed on some collection of cases. Those are evidence, and they are not a specification.
Regulators have responded by treating such systems as devices whose evidence is statistical, which works after a fashion and pushes all the difficulty into the question of which evidence counts. That question turns out to be where most of the argument lives.
Performance on old cases is not performance on new ones
The usual evidence is retrospective: the system was run over historical records with known outcomes and its answers were compared against them. This is cheap, it can cover many cases, and it systematically flatters the result, because those records were assembled under conditions that no longer hold exactly.
Prospective evaluation runs the system alongside normal operation and measures what happens next. It is slow, expensive and much more informative, and it sometimes shows a system that looked strong retrospectively performing considerably worse in place. The reasons vary: different equipment, different populations, different workflow, different rates of the thing being detected.
The rate matters more than people expect. A system evaluated on a curated set containing many positive cases behaves very differently in a setting where positives are rare, because the same detection ability produces a far higher share of false alarms when the underlying prevalence drops.
Updating the model restarts the argument
Traditional approval assumes the approved thing stays the same. A model that is retrained on newer data is, in a meaningful sense, a different model, and if every update requires the full evidence package again then updates become rare and the system drifts away from current conditions.
This is why some frameworks have moved towards approving a change procedure rather than a fixed artefact: the operator declares in advance what kinds of update they will make, how they will validate them and what would trigger a withdrawal. Whether this is sufficient oversight or a loophole is disputed among people who take the question seriously.
The alternative — freezing the model permanently — trades one risk for another. A frozen system is predictable and it decays as the world changes around it, which is a failure mode that produces no alert and is easy to miss for years.
Somebody has to be answerable, and the chain is unclear
When an automated decision goes wrong, liability has to land somewhere: the operator who deployed it, the organisation that built it, the professional who accepted its output, or the institution that set the policy. Existing law was not drafted with a statistical intermediary in mind, and jurisdictions are resolving it differently and slowly.
Several regimes require that a person affected by a consequential automated decision can obtain an explanation or a human review. Whether an explanation generated after the fact satisfies that requirement, given that such explanations are reconstructions rather than records, is exactly the sort of question that will be settled case by case.
None of this constitutes advice about any particular obligation. The rules differ by country, by sector and by the category of decision, and they are being rewritten as this is read; anyone with a real duty here needs to take it to somebody qualified in their jurisdiction.
The disagreement is real and both sides have a case
One argument holds that approval processes are too slow, that they were designed for a different kind of artefact, and that the delay itself causes harm by keeping useful tools out of use. There are documented instances of tools performing well in evaluation and taking years to reach anybody.
The other holds that the evidence standards are, if anything, too permissive — that retrospective evaluation is accepted too readily, that post-deployment monitoring is weak, and that a system distributed widely can cause correlated harm at a scale no individual practitioner ever could.
Both observations can be true simultaneously, which is the uncomfortable part. The processes may be simultaneously too slow at the front and too loose at the back, which is a design problem rather than a matter of choosing a side.
Common questions
Why would a system approved in one country not be available in another?
Because approval is jurisdictional and the evidence standards differ. A regulator may also require evidence drawn from a population resembling its own, since performance can vary between populations, equipment types and record-keeping practices in ways that a study conducted elsewhere does not establish.
Does a high accuracy figure mean a system is ready for deployment?
Not on its own. The figure depends heavily on which cases were tested and how common the target condition was in that set. The same system can look excellent on a curated evaluation and generate an unworkable volume of false alarms in a setting where the condition is rare.
What is post-market monitoring meant to catch?
Deterioration that only shows up in use: shifts in the population being served, changes in equipment or process, and gradual divergence between the world the model was fitted to and the world it now operates in. It is the mechanism intended to catch drift, and how rigorously it is done varies a great deal.
Deputy editor, AI Worth Knowing
Zoya joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and would rather show the working than assert the conclusion.





