Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Picking from a fixed list and composing from nothing are different problems

The two dominant shapes of machine learning task differ in how error can be measured at all, and treating the second as though it were the first sits behind a great deal of misplaced confidence.

By Daniel Okonkwo4 min read

A restaurant server operates a touchscreen POS system, enhancing efficient order management.
Photograph by SpotOn POS via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A closed set of answers changes what is knowable

Classification assigns an input to one of a fixed set of categories: spam or not, which of ten digits, which of a thousand object types. The set is decided in advance by a person, and everything the system can ever say is drawn from it. That constraint is more powerful than it looks.

Because the possible answers can be enumerated, the possible errors can be enumerated too. You can count how often each category was mistaken for each other one and lay the counts out in a grid. Nothing hides. Every mistake is a specific, nameable substitution that somebody can go and inspect.

Generation has no such set. A model asked to write a paragraph has an effectively unlimited space of outputs, of which a great many are acceptable and a great many are not, with no boundary that anybody has drawn. There is no grid to fill in and no obvious substitute for one.

Classification errors have a shape you can draw

Once the grid exists, the vocabulary follows. Of the items the system flagged, what share were right. Of the items that should have been flagged, what share it caught. The two move against each other, and where they balance is a decision about consequences rather than a property of the model.

That decision is a threshold. Most classifiers don’t output a category at all; they output a score, and a person chooses the cut-off above which it becomes a flag. Moving the cut-off trades one kind of error for the other, and the right position depends on which of the two mistakes costs more.

This is unglamorous, and it is one of the more genuinely useful things the field has. A system that can be tuned along that curve can be matched to an operational reality, including the option to abstain in the uncertain middle band and hand the case to a person instead.

Generation has no denominator

For generated text there is no set of wrong answers to count against. The standard workarounds compare the output with one or more reference answers written by people, measuring overlap of words or of meaning, which penalises any perfectly valid answer that the references happened not to phrase that way.

The alternative is judgement: ask people, or increasingly ask another model, to rate the output. Both routes carry their own problems. Human raters disagree with one another and drift over a long session; model raters carry the biases of their own training and tend to favour text resembling what they would have produced.

So the honest position is that generation quality is measured through proxies, all of which are known to be imperfect, and that comparisons between generative systems rest on those proxies. Everybody working in the field knows this. It is stated much less often outside it.

A score you can calibrate against one you cannot

A well-built classifier can be calibrated, meaning that among the cases it scored at seven in ten, roughly seven in ten really do belong to the category. Calibration is checkable after the fact, and it is what makes a score usable inside a decision procedure rather than merely suggestive.

Generative systems offer nothing comparable. The probability the model assigned to the fragments it produced is a statement about text, not about the truth of what that text asserts, and reading it as confidence is a category error. Fluent, high-probability output can be entirely fabricated.

This is the practical crux. Classification hands you a number you can reason about; generation hands you prose you have to verify. The second is more flexible and the first is more accountable, which is a real trade-off rather than a matter of one being the more advanced technology.

Generation took over anyway, and something was lost in the move

It has become common to solve classification problems with a generative model, simply by describing the categories and asking for one. The convenience is enormous: no labelled dataset, no training run, and the categories can be changed by editing a sentence. For many purposes that is a sensible engineering choice.

What it discards is the threshold. The output arrives as a word rather than a score, so the ability to tune where the system errs, to abstain, and to check calibration against real outcomes goes with it. Various techniques recover an approximation of a score, and how trustworthy those approximations are is disputed.

There is a reasonable counter-argument that flexibility is worth more than calibration for most applications, and that a purpose-built classifier is a false economy when the category list changes every quarter. Both positions are defensible. What is not defensible is failing to notice that the trade was made.

Common questions

Is one of these harder than the other?

They are hard in different places. Classification struggles when categories overlap, when examples are scarce for some of them, or when the labels themselves are inconsistent. Generation is hard because success is not defined precisely enough to optimise directly, so training aims at a proxy and hopes the gap is small.

Why does a generative model give a category confidently even when unsure?

Because it was trained to produce plausible continuations, and a hedged answer is often less plausible in context than a decisive one. Nothing in the output format carries uncertainty unless the model was specifically trained to express it, and expressed uncertainty is itself generated text rather than a measurement.

Can the two approaches be combined?

Frequently, and it is a common architecture. A cheap classifier decides whether a case needs the expensive generative system, or a generative system drafts something that a classifier then screens. Each component can then be evaluated on its own terms, which is far easier than evaluating an undifferentiated whole.

How It Worksclassificationgenerationevaluationcalibration
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.