Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A system that can be fooled on purpose is a different problem from one that errs

Ordinary mistakes are distributed by chance and an attacker searching for a failure is not, which turns a tolerable error rate into a security property that the underlying methods were never designed to have.

By Imran Sheikh3 min read

A person interacting with ChatGPT interface on a computer screen in a dimly lit room.
Photograph by Alberlan Barros via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Random error and adversarial error are different categories

A system with a modest error rate is perfectly usable when the mistakes fall where chance puts them. The reasoning changes completely when somebody is deliberately hunting for inputs that fail, because the relevant question is no longer how often the system errs but whether a failing input exists and can be found.

For essentially every model in use, such inputs exist. Finding them is a search problem, and search problems with a measurable objective are exactly what modern optimisation is good at. The same machinery that trains a model can be pointed at it to locate the places where it breaks.

This is why security people and machine learning people often talk past one another about reliability. One group means the average case. The other means the worst case an intelligent opponent can construct, and those two numbers can be extraordinarily far apart.

The perturbation attack and why it works at all

The best-studied version modifies an image by a tiny amount, spread across many pixels, invisible to a person and calculated to push the input across a decision boundary. The result is a picture that looks unchanged and is classified as something entirely different, with high confidence.

The uncomfortable explanation is that models behave close to linearly over the space of inputs, and in a space with very many dimensions a large number of imperceptible nudges accumulate into a substantial movement. The boundary between categories runs closer to every ordinary input than intuition suggests.

A related account holds that these systems latch onto patterns that are genuinely predictive in the data but meaningless to human perception, so an attack is not fooling the model so much as exploiting features it was reasonable to learn. If that is right, the vulnerability is a property of the data rather than a bug in the method, which is a considerably less comfortable conclusion.

Attacks travel between systems that never met

An input crafted against one model frequently fools a different model trained by different people on different data. Transfer is not universal and it is common enough to matter, and it removes the assumption that keeping a model private protects it.

The consequence for defence is severe. An attacker with no access to a deployed system can build their own approximation, attack that at leisure, and carry the result across. Secrecy about weights and architecture provides much less protection than it appears to, and confidence based on secrecy has repeatedly proved misplaced.

Attacks have also been demonstrated in physical form — printed patterns, stickers, particular arrangements of light — which removes the assumption that an attacker must have digital access to the input at all.

Text systems have a version that requires no mathematics

Language models read everything they are given as a single stream, with no structural distinction between instructions from the operator and content from elsewhere. Text arriving inside a document, a web page or a message can therefore contain something that reads as an instruction, and the model has no reliable basis for treating it differently.

This matters most where a system has been given the ability to act — to fetch pages, call other services or write somewhere. The attack surface then extends to every piece of content the system might encounter, which in practice means the open internet.

The underlying issue is architectural rather than a mistake anyone made in configuration. Separating data from instructions is a solved problem in conventional software because the boundary is enforced by the format. Here there is no format, only text, and the boundary has to be inferred by the same system being attacked.

Defence has been an arms race, and opinions differ on the ending

Many proposed defences have been broken shortly after publication, often by adapting the attack to the defence rather than by anything ingenious. This has happened enough times that the field has developed a professional wariness about claims of robustness, and evaluation standards for such claims have tightened considerably.

Approaches with a firmer footing include training explicitly against attacks, which improves robustness at some cost in ordinary accuracy, and methods that provide a mathematical guarantee within a small region around each input. Guarantees are meaningful and the regions they cover are narrow relative to what an attacker can do.

One camp regards this as a temporary state of engineering immaturity that better methods will resolve. Another holds that susceptibility follows from fitting a high-dimensional surface to data and cannot be eliminated, only managed, which would make defence a matter of layered mitigation rather than a fix. The argument is unresolved and worth watching rather than settling.

Common questions

Does adversarial vulnerability matter if nobody is attacking my system?

Less, but the exposure grows with the stakes. Anywhere a model gates access, money, moderation or eligibility, someone has an incentive to find the input that gets through. Systems processing content from strangers should be assumed to be attacked eventually, even if nothing has happened yet.

Would a more accurate model be harder to fool?

Not reliably. Robustness against deliberate attack and accuracy on ordinary inputs are separate properties, and improving one has sometimes been observed to cost a little of the other. A model at the top of a benchmark can be as vulnerable as a mediocre one.

Can filtering the input solve the text version?

Filters catch known patterns and are routinely circumvented by rephrasing, encoding or indirection, because the space of ways to express an instruction is unbounded. Filtering is a useful layer rather than a solution, and designs that limit what the system is permitted to do tend to help more than designs that try to detect bad input.

Limits & Risksadversarialsecurityrobustnessattacks
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.