Limits & Risks
Accurate and useful are separate properties and they come apart often
A system can produce correct answers and change nothing about the decision it was bought to improve, because usefulness depends on timing, on what the errors cost and on what people were already doing.
By Naina Sethi3 min read

Right answer, wrong question
A model that predicts something accurately is only valuable if the prediction feeds a decision somebody can act on. Plenty of accurate systems predict things nobody can change, or predict them at a moment when the decision has already been taken, or predict them for a case that was never in doubt.
The classic shape is a model that identifies risk beautifully among cases where the answer was obvious anyway, and adds nothing in the ambiguous middle where a person actually needed help. The overall score looks strong because the obvious cases are numerous. The contribution is close to zero.
This is not a flaw in the model. It is a mismatch between what was optimised and what the organisation needed, and it usually originates in the framing of the problem months before any training took place.
Mistakes do not cost the same amount
Accuracy treats every error identically, and almost no real setting does. Missing a genuine case and raising a false alarm have different consequences, different costs and often different people bearing them. A single averaged figure blends the two and hides which one the system is choosing to make.
Once you attach costs, the ranking of systems can change entirely. A model with a lower headline score but a tendency to fail in the cheaper direction may be plainly better for the purpose, and the only way to see that is to state the costs explicitly. Organisations frequently decline to, because stating them is politically awkward.
The awkwardness is the point. Writing down what a missed case costs relative to a false alarm forces a conversation about values that everyone would rather leave implicit. Deploying a system without having that conversation does not avoid the decision; it just makes it silently and badly.
Usefulness is measured against what was already happening
The relevant comparison is never against random guessing. It is against the arrangement currently in place, which might be an experienced person, a simple rule written years ago, or an existing model. Those baselines are often stronger than expected, and they are frequently omitted from evaluation entirely.
A simple rule can be a formidable competitor because it is understood, cheap, stable, and easy to override when it is obviously wrong. Beating it by a small margin at much greater cost and much less transparency is a poor trade, and it happens regularly because the comparison was never run.
Where an existing process is genuinely weak, a modest model can be transformative. The size of the gain depends on the gap rather than on the sophistication of the method, which is why the same technique produces a revolution in one setting and a rounding error in another.
Timing and actionability decide most of the value
A prediction delivered after the decision point is worthless regardless of accuracy. A prediction delivered early enough to act on, but attached to nothing anybody can do, is nearly as worthless. Value comes from the combination of a correct answer, arriving in time, next to an available action.
This is why so much of the effort in deploying a model turns out to be workflow design rather than modelling. Where the output appears, who sees it, what it interrupts, what it lets a person do next — those choices routinely have more effect on the outcome than several points of accuracy would.
It also explains a familiar disappointment. A pilot demonstrates strong performance on historical cases, the system is installed, and nothing measurable improves. The model was fine. It was inserted at a point in the process where the answer could not be used.
Offline scores and real outcomes are different measurements
Evaluation on a held-out dataset asks whether the model reproduces recorded answers. Deployment asks whether the world got better. These come apart because acting on a prediction changes what happens next, so the recorded data no longer describes the situation the model is now operating in.
A prediction that leads to an intervention makes itself wrong when the intervention works, which is a genuine measurement problem rather than a paradox. Systems that flag risk and prompt action will look inaccurate precisely where they succeeded, and untangling that requires a comparison group rather than a scoreboard.
The serious way to settle it is a controlled comparison in the live setting, which is expensive, slow and sometimes not permissible. Where it has been done, the results have often been more modest than the offline figures implied. That gap is not evidence of dishonesty. It is what the two measurements mean.
Common questions
Can a system be worth deploying without beating the existing process?
Sometimes, if it delivers comparable results faster, more cheaply, more consistently, or at hours when nobody is available. Those are legitimate reasons and they should be stated as the reason, rather than dressed up as an accuracy improvement that the evidence does not support.
Why do pilots so often outperform the eventual deployment?
Pilots run on selected cases, with engaged staff, close attention and fresh data. Deployment runs on everything, with ordinary staff, no special attention and data that ages. Almost every factor that differs between the two runs in the same direction, so a drop is the expected outcome rather than a surprise.
Is a lower-accuracy model ever the better choice?
Frequently. If it fails in the direction that costs less, is faster to run, is easier to explain when challenged, or can be overridden sensibly by the person using it, those properties can outweigh a higher score. Accuracy is one input to the decision and rarely the decisive one.
Features writer, AI Worth Knowing
Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.





