Limits & Risks
Interpretability has produced real findings and not the thing people wanted
The effort to work out what a trained network is actually doing has become a serious research programme with genuine results, and it is still a long way from the explanation that most people mean when they ask for one.
By Naina Sethi3 min read

The question is older than the current wave
As soon as trained models began outperforming hand-written rules, the objection arrived: nobody can say why it produced that output. The objection has never been answered satisfactorily, and it has been reframed several times as the systems changed shape and the stakes rose.
It is worth being precise about what is being asked, because at least three different questions travel under the same heading. Which parts of this input mattered. What internal process produced this output. And what would have had to be different for the answer to change.
Those questions have different answers, different methods and different degrees of difficulty. Conversations about explainability often founder because two people are asking two of them and each assumes the other means theirs.
Highlighting the input was the first serious attempt
A family of methods marks which parts of an input most influenced the output — which pixels, which words. The output is intuitive, easy to display, and was adopted quickly in settings where somebody had to justify a decision to a person affected by it.
Then came the checks. Researchers found that some of these methods produced similar-looking highlights even when the model’s parameters had been randomised, which is a devastating result for a technique meant to describe what a specific trained model was doing. Not every method failed, and the episode changed how the field validates such tools.
The broader lesson stuck. An explanation that looks plausible to a person is not thereby faithful to the computation, and plausibility is precisely what these methods are optimised to produce. Testing for faithfulness separately became standard practice rather than an afterthought.
The mechanistic turn asks what the parts are doing
A newer programme tries to reverse-engineer the computation directly: identifying what individual components respond to, tracing how information moves between layers, and describing small circuits that implement a specific behaviour. Some of these accounts have been detailed enough to make predictions that were then checked, which is the standard one would want.
The main obstacle discovered along the way is that individual units do not correspond to individual concepts. A single unit responds to several unrelated things, apparently because a network with limited width packs more distinct features into its dimensions than it has dimensions, overlapping them in a way that is efficient and illegible.
Techniques for pulling those overlapping features apart into a larger set of cleaner ones have become a substantial research direction. Early results are genuinely interesting. Whether they scale to accounting for the behaviour of a large model, rather than isolated behaviours within one, remains open.
A faithful account and a usable one are not the same
Suppose the mechanistic programme succeeded entirely and produced a complete description of the computation. It would be enormous — millions of interacting components — and reading it would not obviously help a person appealing against a decision, or a clinician deciding whether to trust an output.
What people usually want is a short, causal, human-sized reason. Those exist for simple models and may not exist for complex ones, in the same way that no short account explains why a particular economy grew last year. Compression is what makes an explanation useful and it is also what makes it partly false.
This is why some researchers argue for building models that are interpretable by construction in high-stakes settings, accepting some loss of performance in exchange for a system whose reasoning can be stated. The counter-argument is that the loss is real and sometimes large, and that a more accurate opaque system may serve people better. The disagreement is substantive and ongoing.
An honest scorecard
The field has established that internal representations carry recoverable structure, that specific behaviours can sometimes be traced to specific mechanisms, that several intuitive explanation methods do not survive scrutiny, and that self-reported reasoning from a model is not a description of its internal process.
It has not established a general method for taking an arbitrary output and saying why. It cannot yet certify that a system will not do something. And it has not produced explanations that a non-specialist can act on, which is the version that regulation and public expectation both assume exists.
Whether the gap is temporary is disputed. Optimists point to steady progress and to the fact that the programme is young and small relative to the effort spent building the systems. Sceptics argue that a fitted high-dimensional function may simply not have a compact explanation, and that no amount of effort produces one. Both positions are held by serious people.
Common questions
Can a model be asked to explain itself?
It can produce an explanation, and that text is generated by the same process that generated the answer rather than read off any internal record. It may be accurate, and there is no mechanism guaranteeing it corresponds to what actually happened inside. Treat it as a hypothesis rather than a report.
Are simpler models always more interpretable?
Usually, though not automatically. A decision tree with thousands of branches or a linear model over ten thousand engineered features is not meaningfully readable either. Interpretability comes from a small number of components with meanings a person can hold in mind, which is a property of size and framing rather than of model type.
Does interpretability research make systems safer?
It contributes tools for auditing behaviour and diagnosing failures, which is a real contribution. It does not currently provide guarantees, and treating an interpretability result as assurance that a system will behave in some way goes further than the methods support.
Features writer, AI Worth Knowing
Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.





