Limits & Risks
Finding the rare case is hard for a reason arithmetic makes unavoidable
Detecting fraud, faults and intrusions means searching for events that almost never happen, and the scarcity that makes them worth finding is exactly what makes them difficult to learn.
By Samar Bhatia3 min read

You cannot learn from examples you do not have
Ordinary supervised learning wants many examples of each category. Anomaly detection wants to find the category with hardly any examples at all, and often with none that resemble the next one. Fraud schemes change, machines fail in new ways, and intrusions are designed specifically to look unlike previous intrusions.
So the usual framing is inverted. Rather than learning what the rare event looks like, the system learns what ordinary behaviour looks like and flags whatever sits far from it. That reframing is elegant and it quietly substitutes a different question: unusual is not the same as bad.
Most unusual things are harmless. A new customer with an odd purchasing pattern, a machine running warm because the weather changed, a login from an airport. A detector built on distance from normal will surface all of them, and it has no way to tell which of them anybody cares about.
The arithmetic of rarity defeats intuition
Imagine a detector that correctly flags almost every genuine case and only rarely flags an innocent one. Now apply it to a population where genuine cases are one in many thousands. The small proportion of innocent cases flagged is drawn from an enormous number, and it swamps the genuine ones completely.
The consequence is that most alerts are false even when the detector is excellent by any conventional measure. This is not a failure of the model and it isn’t fixable by improving the model modestly. It follows from the ratio between how common the event is and how often the detector is wrong.
Anyone evaluating such a system on a balanced test set, with equal numbers of normal and abnormal cases, will therefore see a figure that has almost no bearing on what happens in operation. This is one of the most common and most costly mistakes in applied work.
Alerts nobody can process are worse than no alerts
A team can investigate a certain number of cases per day, and that number is the real constraint on the system. Producing more alerts than can be examined does not increase detection; it means alerts are ignored, sampled arbitrarily, or handled so quickly that the examination is nominal.
The measure that matters is therefore not accuracy but yield within the budget: of the cases we could actually look at, how many were worth looking at. That reframes the design entirely, from a detector to a ranking system whose job is to fill a fixed queue as well as possible.
It also has a human cost that compounds. A queue that is mostly noise trains the people working it to expect noise, and the eventual genuine case arrives to an audience that has stopped expecting one. The failure then looks like inattention and is really a design decision made upstream.
Normal is a moving target, sometimes moved deliberately
What counts as ordinary behaviour changes with seasons, product launches, holidays, equipment replacement and shifts in who the customers are. A detector calibrated against last quarter will treat this quarter as anomalous, producing a wave of alerts that reflects the calendar rather than anything wrong.
Retraining on recent data handles gradual change and introduces a different weakness: if the recent data contains the very behaviour you want to catch, the system learns it as normal. Slow-moving abuse can be absorbed into the baseline precisely because it persisted long enough to look like the background.
Where an adversary is involved the problem becomes strategic. Someone who suspects a threshold exists will probe below it, and behaviour will migrate to whatever the detector considers ordinary. Detection then shapes the thing it is detecting, which is not a situation that any static evaluation captures.
What tends to work is a system rather than a detector
Deployments that hold up usually combine several weak signals with explicit rules, route cases into tiers by cost, and accept that the highest tier gets human attention while the rest is handled by cheaper measures such as friction, delay or a request for confirmation.
They also instrument the feedback loop deliberately, sampling cases that were not flagged in order to estimate how much is being missed. Without that, the only cases anybody sees are the ones the system already caught, and the record silently confirms whatever the model believes.
None of that is glamorous and it is where the results come from. The pattern repeats across the field: the interesting part is rarely the model, and the parts that determine whether it works are the ones nobody writes about.
Common questions
Why is a detector with excellent accuracy still mostly wrong in practice?
Because accuracy is measured against the test set and the alert stream is drawn from the real population, where the target event is rare. A small error rate applied to a very large number of ordinary cases produces more false alerts than there are genuine ones, however good the detector looks in evaluation.
Would more training data solve it?
It helps with defining ordinary behaviour and it does not supply what is missing, which is examples of rare events. Collecting more data mostly collects more of the common case. Improvements tend to come from better signals and better handling of the alert queue rather than from volume.
Is a rule-based system ever preferable here?
Often, at least as one layer. Rules are explicit, auditable, quick to change when a new pattern appears, and defensible when challenged. Their weakness is that they only catch what somebody anticipated, which is why the durable designs use both rather than choosing between them.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





