Limits & Risks
A model can be corrupted during training in a way that shows up only on a chosen signal
Interfering with training data or with a distributed model file can plant behaviour that lies dormant through every ordinary test, which makes provenance a security property rather than an administrative one.
By Manish Trivedi4 min read

Attacking the training is different from attacking the running system
Most discussion of adversarial behaviour concerns inputs crafted at the moment of use to produce a wrong answer. There is an earlier point of attack. If somebody can influence what a model is trained on, they can influence what the finished model does, and the influence persists in every copy of it thereafter.
The general name is poisoning, and it splits into two aims. One is degradation: corrupt enough of the data and the model is simply worse, which is noisy, requires volume, and is usually detected because performance drops. The other is far more interesting and far more troubling.
The second aim is a backdoor. The model behaves normally on everything except inputs carrying a particular marker, and on those it does what the attacker chose. Ordinary accuracy is unaffected, so ordinary testing finds nothing, because the test set does not contain the marker.
Why a small amount of corrupted data can be enough
The intuition that a handful of bad examples cannot matter among millions is reasonable and wrong in this particular case. The reason is that the trigger occupies a region of the input space that nothing else occupies. There is no competing evidence about what a bizarre and specific pattern should mean, so the few examples defining it face no opposition.
Research demonstrations have established that this works across image classifiers, text models and speech systems, using triggers ranging from a small visual patch to an unusual phrase. The general finding is that the required proportion of poisoned examples is much smaller than intuition suggests, and that it does not grow proportionally with the dataset.
That last point deserves emphasis. A defence based on the idea that large datasets dilute the problem does not follow from what has been observed, and the scale of modern corpora is not by itself protective.
The opportunity comes from how corpora are assembled
Large training sets are built by gathering material from many sources, and a substantial portion of any web-scale corpus consists of pages that anybody can edit or publish. Anything that can be written can potentially be trained on, and the collection process rarely establishes who wrote what.
There is a further mechanism specific to gathered corpora. A dataset is often distributed as a list of locations rather than as the material itself, and material at a location can change after the list was made. Somebody acquiring control of a small fraction of those locations acquires a channel into every model trained from the list afterwards.
None of this requires sophistication or privileged access, which is what distinguishes it from most attacks on machine learning. It requires patience and the ability to publish, which almost everybody has.
The distributed file is a second surface
A model obtained from a repository is a file that somebody produced, and the recipient generally cannot verify how it was produced. A backdoor planted deliberately by whoever trained it is indistinguishable from a backdoor introduced through poisoned data, and neither is visible by reading the parameters.
Some serialisation formats compound this by permitting arbitrary code to run when a file is loaded, which is a conventional software vulnerability rather than a machine learning one. Safer formats exist and are increasingly the default, and the older ones remain in circulation.
Checksums and signatures establish that a file has not been altered since publication. They say nothing about what it did when it was published, which is the harder question and the one with no good technical answer at present.
Detection is difficult for a structural reason
Finding a backdoor means finding an input pattern that changes behaviour, without knowing what pattern to look for, in a space of possible inputs far too large to search. Techniques exist that attempt to reconstruct likely triggers or to identify parameters implicated in unusual behaviour, and they work against known styles of attack.
The pattern that follows is the familiar one from security research generally: a defence is published, an attack adapted to it is published shortly afterwards, and the defence is understood to cover a narrower case than first claimed. This has happened enough times that broad robustness claims are treated cautiously.
What does help is entirely conventional. Knowing where training material came from, keeping records of it, retaining the ability to retrain, obtaining models from parties with something to lose, and testing against the specific behaviours that would matter in your deployment rather than against general benchmarks.
How seriously to take this
Most of the evidence is from research settings rather than from documented attacks in the wild, and it is fair to note that a demonstrated vulnerability is not a demonstrated threat. Someone attacking a system usually has cheaper options than a multi-year campaign to influence a training corpus.
The counter-argument is that this class of attack is by construction difficult to observe, so the absence of public incidents is weak evidence. It is also the kind of attack that becomes worth the effort precisely as these systems are placed in positions where their decisions matter.
The proportionate response is to treat model provenance the way software supply chains are already treated, which is to say seriously, imperfectly, and with an accepted residual risk that is stated rather than assumed away.
Common questions
Can a poisoned model be cleaned by further training?
Additional training on clean data reduces the effect and has not been shown to remove it reliably, and implanted behaviour has been observed to persist through adaptation in research settings. Retraining from verified data is the dependable answer and is rarely practical.
Is this the same as a jailbreak?
No. A jailbreak manipulates a finished model through its input to get behaviour its operators did not intend. A backdoor was built into the model before anybody used it, and would be present even if every input were benign.
Does using a model through a service avoid the problem?
It transfers the problem to whoever runs the service, who faces the same questions about their own training data and dependencies. It does mean the risk is being managed by an organisation with an incentive to manage it, which is not nothing.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





