Limits & Risks
Capability is a claim about something nobody can observe directly
Every statement that a system can or cannot do a thing is an inference from behaviour on a sample of cases, and the gap between the behaviour and the trait is where most disputes in the field actually sit.
By Samar Bhatia3 min read

We watch behaviour and then talk about traits
Nobody measures a capability. What gets measured is performance on a set of items, and the word capability is then applied to whatever is presumed to have produced that performance. The move from one to the other is an inference, and it is made so routinely that it usually goes unmarked.
This is familiar territory in fields that measure human attributes, where the distinction between an observable score and an underlying trait has been argued over for a century. The vocabulary developed there — whether an instrument measures what it claims to, whether it measures the same thing across different groups — applies here with very little modification.
Applying it is uncomfortable, because the answers are frequently that the instrument measures something narrower than its name suggests. A test called reasoning measures performance on those particular items, presented in that particular way, in that particular language.
The same trait can vanish when the wrapping changes
A system that answers a class of question correctly may fail on the same question rephrased, reordered, embedded in a longer passage or translated. When that happens, it is unclear whether the capability was absent or whether the test was measuring something incidental to it.
Sensitivity to presentation is a well-documented and awkward property. It means a single evaluation establishes much less than it appears to, and that differences between systems can reflect how well each happens to suit the format of the test rather than any underlying difference in ability.
The reasonable response is to test the same claim several ways and report the spread. This is more work, produces less quotable results, and is done inconsistently. The incentive structure around publishing scores does not reward it.
Being able to and being willing to are separate
A model may possess whatever internal machinery is needed for a task and still fail to use it, because the input did not evoke it, the format discouraged it, or training taught the system to decline. Failure is therefore weak evidence of absence, while success is reasonably strong evidence of presence.
This asymmetry has real consequences for anybody trying to assess risk. A system that refuses a request has demonstrated a policy, not an incapacity, and later changes to the surrounding software can alter what the same parameters will do without altering the parameters.
It cuts the other way too. Success achieved after many attempts, with the successful one selected afterwards, demonstrates something quite different from success on a first attempt, and evaluations vary in whether they distinguish the two.
The dispute about abilities appearing suddenly
A widely discussed claim is that certain abilities appear abruptly as models grow: absent, absent, then present. If true, this has serious implications, because it would mean that what a larger system can do is not predictable from smaller ones.
The counter-argument is that the abruptness is often produced by the scoring. A measure awarding credit only for an exactly correct answer will show nothing until performance crosses a threshold, even where the underlying improvement was smooth all along. Rescore the same runs with a measure that gives partial credit and some of the sharp jumps flatten out.
That analysis is taken seriously and it does not dispose of every case, and researchers continue to disagree about how much of the phenomenon survives it. The honest position is that the shape of improvement depends partly on how you chose to look, which is a real finding rather than a dodge.
The definition matters because rules are being written on it
Regulatory and voluntary frameworks increasingly express obligations in terms of what a system is capable of, with thresholds triggering additional scrutiny. That approach depends on capability being measurable in a way that different parties would agree on, and the preceding sections are reasons for doubting that it currently is.
The alternatives all have drawbacks. Defining thresholds by training compute is measurable and only loosely related to what a system can do. Defining them by application is more meaningful and much harder to police. Defining them by evaluation results makes the evaluation a target, with everything that follows from that.
None of this is an argument against measuring. It is an argument for reading capability claims as summaries of specific observed behaviour, and for treating the leap from observation to trait as the place where the interesting questions were quietly skipped.
Common questions
If a model solves a problem once, does it have the capability?
It depends on how the attempt was selected. One success out of many attempts, chosen after the fact, shows the behaviour is reachable rather than reliable. Consistent success across varied phrasings of the same task is much stronger evidence, and is what careful evaluations try to establish.
Why do two evaluations of the same model disagree so much?
Different items, different formats, different scoring rules and different settings for how the output was generated. Each of those can move a result substantially. Disagreement between evaluations is usually informative about the tests rather than evidence that one of them was run incorrectly.
Does this mean capability claims are meaningless?
No. It means they are claims about observed behaviour under stated conditions, and are useful in proportion to how clearly those conditions are described. A claim accompanied by its test items, its scoring rule and its variation across runs is worth a great deal more than a single number.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





