Limits & Risks
A deployed model decays because the world moves and the parameters do not
Performance measured at launch describes a relationship between a fixed model and a moment that has already passed, and the gap widens quietly in ways nothing in the system reports.
By Samar Bhatia3 min read

Fitted to a world that will not hold still
A model is fitted to a particular distribution: the range of inputs it saw and the relationships between those inputs and the outcomes it was trained to predict. That distribution described the world during a specific period, and the world was under no obligation to stay that way.
Shifts happen for banal reasons. A form is redesigned and a field arrives in a new format. A supplier changes. Customer behaviour moves in response to a competitor. A regulation alters what people can do, and therefore what they do. None of this involves anything dramatic and all of it moves the ground under a fitted model.
The distinction worth holding is between the inputs changing and the relationship changing. A model can adapt reasonably to slightly unfamiliar inputs. It cannot adapt at all to a world where the thing it learned to predict now depends on something else.
The failure is silent by construction
Conventional software fails loudly. A service that cannot reach its database raises an error and somebody is paged. A model whose accuracy has fallen keeps producing well-formed outputs at the same rate, in the same format, with the same apparent confidence. Nothing in the pipeline registers a fault.
Detecting the decay requires comparing predictions against outcomes, and outcomes often arrive late or not at all. Whether a rejected loan application would have been repaid is never observed. Whether a flagged transaction was actually fraudulent may take months to establish, and only for the ones that were investigated.
So teams monitor proxies instead: the statistical profile of incoming data, the distribution of outputs, the rate of human overrides. These are early warnings rather than measurements, and they can miss a change that alters the relationship without altering the inputs at all.
Deployment itself changes the distribution
A subtler mechanism operates whenever a model influences the process that generates its future data. A screening system that rejects certain applications ensures that no outcome is ever observed for them, so the record used for future training is systematically incomplete in precisely the region where the model was making decisions.
The same effect appears wherever people adapt to being scored. If a measure becomes consequential, behaviour reorganises around it, and the historical relationship between the measure and the underlying quality weakens. The model was fitted to a world in which nobody was optimising against it.
This is why evaluating a deployed system against its own accumulated data is unreliable. That data is not an independent sample of the world; it is a record of the world after the system started acting on it.
Retraining is maintenance, not a fix
The standard response is periodic retraining on recent data. This works, up to a point, and it introduces its own difficulties. Each retrained version behaves slightly differently, so anything downstream that depended on the old behaviour may break, and errors in recent labelled data propagate directly into the new model.
Retraining also cannot solve the observation problem. If outcomes are missing for the cases the model rejected, then training on recent data trains on a filtered view, and the filtering was performed by the previous version of the model. Correcting for that requires deliberate effort — holding out a fraction of decisions from automation, for example — which costs money and looks like inefficiency to anyone not tracking why it exists.
A retraining schedule is therefore an operating cost that continues for as long as the system is used. Projects budgeted as though a model were a deliverable rather than a maintained asset run into this a year or two after launch, with predictable results.
Rare events are where the decay hurts most
Systems trained on ordinary conditions have seen little of the unusual, which is exactly when their outputs matter most and when the distribution has moved furthest. A model performing well in normal times may be at its least reliable precisely during a disruption, and there is no signal to say so.
This argues for keeping a non-automated path available for unusual situations, and for treating a model’s confidence as least informative when conditions are least familiar. Both are unfashionable, because both cost something during the long stretches when nothing unusual is happening.
None of this makes deployment unwise. It makes deployment a commitment: something to be monitored, periodically re-measured against reality, and retired when the world it was fitted to has gone. The technology is not the hard part of that. The organisational habit is.
Common questions
How quickly does a model degrade?
It depends entirely on how fast the underlying domain changes. A model of physical measurements may remain valid for years, while one fitted to consumer behaviour or fraud patterns can deteriorate within months because the environment includes people actively adapting. There is no general rate, which is why monitoring rather than a schedule is the answer.
Can a system detect its own drift?
Only indirectly. It can compare current input statistics against those seen during training and raise a flag when they diverge, which catches some cases. It cannot detect a change in the relationship between inputs and outcomes without observing outcomes, and those are often the very thing that is unavailable.
Is continuous automatic retraining the answer?
It is one approach and it carries risks. Automatic retraining on live data can lock in feedback effects, absorb bad labels without review, and change behaviour without anyone deciding it should change. Where it is used, it is normally paired with held-out evaluation and the ability to roll back quickly.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





