Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Attention is a weighted lookup and the name oversells it considerably

The mechanism behind modern language models lets every part of a sequence consult every other part, and its real advantage was never the biological metaphor.

By Manish Trivedi3 min read

Various tangled wires connected to system near black metal cases in server room
Photograph by Brett Sayles via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The problem it was built to solve

Before this design took over, sequences were processed in order, one element at a time, with the network carrying a summary of everything it had seen so far. That summary was a fixed-size bundle of numbers, and everything relevant from earlier in the sequence had to survive inside it.

The consequences were predictable. Information from far back got squeezed out, long-range relationships were unreliable, and because each step depended on the one before it, the work could not be spread across many processors. Training was slow in a way that no additional hardware fixed.

Attention addresses both problems at once, and it is the second one — the parallelism — that changed the field, though it is rarely the part that gets explained.

What the mechanism actually computes

Each position in the sequence produces three sets of numbers from its current representation: one describing what it is looking for, one describing what it offers to others, and one carrying the content it would pass along if selected. The names are query, key and value, borrowed loosely from databases.

Every query is then compared against every key, producing a relevance score for each pair. Those scores are normalised so they sum to one, and each position builds its new representation as a weighted blend of all the values, in proportion to those scores. A position that is highly relevant contributes most; the rest contribute a little.

That is the entire operation. It is a soft lookup — instead of retrieving one record, it retrieves a mixture of all of them, weighted by relevance. Nothing about it is conceptually deep, and its elegance lies in the fact that all these comparisons are independent of each other and can therefore be done simultaneously.

Parallelism was the real prize

Because every position attends to every other position in a single operation, an entire training example can be processed at once rather than stepped through. That turns training into a series of very large matrix multiplications, which is precisely the operation modern accelerators are built to perform.

This is why the architecture spread so quickly. It was not that it was obviously smarter; it was that it converted a sequential problem into a shape that available hardware could chew through at enormous rates, which meant far more data could be trained on in the same wall-clock time. Scale followed from that, and much of what came afterwards followed from scale.

A design that fits the hardware beats a more elegant design that does not. That lesson recurs throughout the history of computing and it is worth holding onto when the next architecture is announced.

The cost grows faster than the input

Comparing every position with every other position means the number of comparisons grows with the square of the sequence length. Double the input and the attention work roughly quadruples. This is the fundamental reason long contexts are expensive and why extending them has been an engineering campaign rather than a switch.

A great deal of work exists on cheaper approximations — restricting each position to a neighbourhood, selecting a subset of positions to attend to, or reformulating the maths so the quadratic term disappears. Many of these work well on some tasks and lose something on others.

Whether full attention is essential or merely convenient remains genuinely unresolved. Approximations keep proving adequate for particular purposes while the unapproximated version keeps winning at the frontier, and reasonable people read that pattern in opposite ways.

What attention maps do and do not tell you

Because the relevance scores are explicit numbers, they can be extracted and drawn as diagrams showing which positions attended to which. These pictures are appealing and they get used heavily to explain what a model is doing internally.

They should be read carefully. The scores describe where information was blended from at one layer of many, not what the model concluded or why, and there is a substantial published argument about whether such maps constitute an explanation at all. High attention to a word does not establish that the word determined the output.

The broader caution generalises. A mechanism that can be visualised is not thereby understood, and the visual appeal of a diagram is not evidence about the process underneath. Interpretability is an active research area precisely because the easy answers turned out to be insufficient.

Common questions

Why are there multiple attention heads?

Each head performs the same operation with its own learned queries, keys and values, so different heads can specialise in different kinds of relationship — one tracking grammatical agreement, another tracking longer-range references. The specialisation emerges rather than being assigned, and it is often messier than the tidy examples suggest.

Does the model know what order the words came in?

Not from attention itself, which treats the input as an unordered set. Position information is added separately, encoded into the representations before attention runs. How best to encode it is an ongoing design question, and different choices affect how well models handle inputs longer than they were trained on.

Is attention anything like human attention?

Only as a metaphor, and a loose one. The mechanism assigns weights to everything simultaneously rather than selecting and excluding, which is close to the opposite of what the psychological term describes. The name has caused a fair amount of confusion for people arriving from cognitive science.

How It Workstransformersattentionarchitecturesequences
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.