Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A token is not a word, and several odd behaviours follow from that

Language models do not read letters or words; they read fragments from a fixed vocabulary, and the seams between those fragments explain failures that otherwise look inexplicable.

By Manish Trivedi3 min read

Detailed view of server racks with glowing lights in a data center environment.
Photograph by panumas nikhomkhai via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The unit of text is not the one you use

Before any text reaches a language model it is cut into pieces called tokens, and those pieces do not correspond to words. Common words usually survive intact as one token. Rarer ones get split into fragments, and a long technical term may be broken into four or five parts that mean nothing individually.

The set of available fragments is fixed before training begins. A procedure runs over a large sample of text, starts from individual characters, and repeatedly merges whichever adjacent pair occurs most frequently, building up a vocabulary of a chosen size. Frequent sequences earn their own entry; infrequent ones remain assembled from smaller parts.

That vocabulary is then permanent for the life of the model. It cannot be revised without retraining, and everything the system ever reads or writes passes through it. It is worth appreciating how thoroughly this shapes what follows.

Counting letters is genuinely hard from the inside

A model asked how many times a particular letter appears in a word is being asked to inspect something it cannot see directly. The word arrived as one or two opaque identifiers, not as a sequence of characters, so the internal representation of it need not contain any explicit record of its spelling.

Models often get these questions right anyway, because spelling is discussed constantly in written text and the relationship between fragments and letters can be inferred statistically from all that discussion. But it is inference rather than inspection, and inference degrades on unusual words, unusual languages and unusual questions.

The same explanation covers rhyming, reversing a string, counting syllables and puzzles that depend on the shape of a word rather than its meaning. None of these are hard problems in themselves. They are hard for this architecture because the architecture threw the relevant information away at the door.

Some languages cost several times more than others

Vocabularies are built from a sample of text, and if that sample is dominated by one language then that language gets the efficient fragments. Scripts and languages under-represented in the sample are reconstructed from smaller pieces, sometimes from individual characters, so the same sentence consumes far more tokens.

This is not merely an accounting curiosity. Context limits are measured in tokens, and where usage is billed, billing is measured in tokens too. A speaker of an under-represented language therefore fits less into a given window and may pay more for the same content, for reasons that have nothing to do with the language’s complexity.

It also affects quality in a subtler way. A word represented as one coherent unit gives the model a cleaner handle than the same word smeared across six fragments, though how much this matters relative to the sheer quantity of training text in that language is debated and probably varies by task.

Numbers are where the seams show most clearly

Digits are chopped up by exactly the same frequency-driven procedure that handles letters, and the results are irregular. Depending on the vocabulary, a long number may be split into groups that bear no relationship to place value — the boundaries fall where the training text happened to make certain digit sequences common.

Arithmetic performed on such a representation is doing something quite unlike the column-wise procedure taught in school. This is one reason numerical mistakes in these systems tend to look strange rather than merely careless, and why the same calculation can succeed at one length and fail at another.

Several projects now handle digits deliberately, splitting them individually or in fixed groups to make the representation regular. It helps. It does not turn a statistical text model into a calculator, which is why arithmetic is increasingly handed off to an actual calculator instead.

What tokenisation is not responsible for

It is tempting, once you have this explanation, to reach for it everywhere. Resist that. Tokenisation does not explain why a model states a false fact confidently, why it agrees with a mistaken user, or why it produces plausible references that do not exist. Those come from the training objective and the data, not from how text was chopped up.

Nor is tokenisation universally regarded as necessary. Work on models that operate directly on raw bytes or characters continues, and it removes these problems at a straightforward cost in efficiency, since sequences get much longer. Whether that trade eventually becomes worthwhile is an open question rather than a settled one.

The reasonable position is narrow. A specific family of failures — spelling, counting, character manipulation, some arithmetic, and unequal treatment across languages — traces back to the vocabulary. The rest of the system’s shortcomings need their own explanations.

Common questions

Roughly how much text is a token?

For ordinary English prose the common rule of thumb is that a token averages around three quarters of a word, but it is only a rule of thumb. Code, unusual names, other scripts and heavy punctuation all shift the ratio substantially, sometimes by several times.

Does a space count as part of a token?

Usually the leading space is attached to the word that follows it, so the same word can have two different representations depending on whether it starts a line or sits mid-sentence. This is one of the small irregularities that makes reasoning about token boundaries by eye unreliable.

Could a model just be told the spelling of a word?

It can, and writing a word out letter by letter in the input does reliably improve character-level tasks, because the letters then genuinely are separate tokens. That is a workaround rather than a fix, and it does not change what the model can perceive when nobody has done it for it.

How It Workstokenstokenisationtextvocabulary
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.