Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Zero-shot and few-shot describe the input, not an ability of the model

The terms count how many examples were supplied at the moment of asking, they say nothing about what the system was trained on, and the difference has quietly become impossible to verify.

By Samar Bhatia3 min read

Close-up of a ballpoint pen resting between the pages of an open book, perfect for education or reading themes.
Photograph by Talha Riaz via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The words came from a narrower setting

Both terms originate in image and text classification research, where they had a precise meaning. Zero-shot meant classifying into a category for which the model had seen no training examples at all, usually by relying on a description of the category rather than instances of it. Few-shot meant a handful of examples, sometimes a literal five.

The setting mattered because the training data was known. Researchers had assembled it, could list the categories in it, and could therefore state with confidence that a particular category had been excluded. The claim was checkable, which is what made it a claim.

That checkability is the part that has gone missing. The words survived and the conditions that gave them meaning did not, which is a common fate for technical vocabulary as it travels.

What they mean in current usage

Applied to a large language model, zero-shot now generally means the request contained no worked examples and few-shot means it contained some. That is a statement about the text supplied at the moment of asking, and nothing more.

This is a real distinction and worth having a word for, since supplying examples demonstrably changes behaviour, particularly on tasks with a specific output format. The phenomenon is usually called in-context learning, which is a better name for what happens even if the word learning is doing its familiar overtime.

What it is not is a measure of how much the model already knew. A system that has read the web has almost certainly encountered the task, or something close to it, many times. Withholding examples from the request does not withhold them from the model.

Nothing is learned in the sense the word implies

When examples are supplied in the request, no parameter changes. The model reads them the way it reads everything else and its output is conditioned on them, exactly as it would be conditioned on any other preceding text. Close the session and the effect is gone entirely.

The distinction from fine-tuning is therefore absolute rather than a matter of degree. One alters the artefact permanently and affects everybody who uses it; the other affects one response and vanishes. They are frequently discussed as though they were two strengths of the same intervention, which they are not.

Why supplying examples helps at all is itself an interesting question, and the explanations offered — that the examples identify which of many learned behaviours is wanted, that they specify a format, that something resembling an internal fitting procedure occurs — are not fully reconciled with one another.

The claim has become unverifiable at scale

A zero-shot result was once evidence that a system generalised to something genuinely new. Establishing that today would require knowing what was in a training corpus assembled from a large fraction of the public internet, and for most systems that composition is either undisclosed or too large to search exhaustively.

So a modern zero-shot figure establishes something much weaker: that the model performed without examples in the request. Whether the task was novel to it is unknown, and in most cases the safe assumption is that it was not.

This is not an accusation of bad faith. Researchers reporting these numbers generally say what they did, and the words carry an implication from their original setting that the current setting cannot support. The implication is what misleads, rather than any individual claim.

Reading the terms sensibly

Treat both as descriptions of experimental setup, comparable to reporting the temperature setting or the length of the input. They tell you how the test was run. They do not tell you what the system is capable of, and they certainly do not tell you it encountered the problem for the first time.

The comparison that remains meaningful is within a single model: how much better it does with examples than without. That difference is measurable, it varies enormously by task, and it is more informative than either figure on its own.

It is also worth separating the term from the related habit of counting examples as though they were a currency. Two examples chosen well can outperform ten chosen carelessly, so the number in the label is a poor summary of what was actually supplied and an even poorer basis for comparing one reported result against another.

One last note on the vocabulary itself. Shot came from an analogy with attempts, and the analogy has decayed to the point where the word contributes nothing. Saying with examples and without examples costs two extra syllables and confuses nobody.

Common questions

Does supplying examples always improve results?

No. It helps most with output format and with tasks where the instruction is hard to state precisely, and it can hurt when the examples are unrepresentative, since the model will generalise from them faithfully. It also consumes space in the context that something else might have used.

Is few-shot the same as fine-tuning on a few examples?

No, and the confusion is common. Fine-tuning changes the stored parameters and the change persists across every future use. Supplying examples in a request changes nothing permanent and applies only to that response. The two also fail in different ways and cost very different amounts.

Why do published zero-shot results vary so much between papers?

Because the setup differs in ways the label does not capture: the exact wording of the instruction, how the output was parsed, which scoring rule was applied, and how the answers were generated. Two groups reporting the same term can be running quite different experiments.

Jargonzero-shotfew-shotin-contextterminology
Samar Bhatia
Senior writer, AI Worth Knowing

Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.