Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A context window is a budget, and everything competes for the same space

The number quoted for how much a model can take in covers the instructions, the conversation, any attached material and the reply, and treating it as a measure of comprehension is a mistake.

By Imran Sheikh3 min read

Close-up of an open Spanish Bible showing the introduction to Chronicles.
Photograph by Leslie Duarte Castro via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

What the number actually counts

A context window is the maximum length of the sequence a model can process at once, counted in the fragments the model reads rather than in words or characters. It is a hard architectural limit rather than a preference, and exceeding it produces an error or silent truncation depending on the system.

The important part is what shares that space. Provider instructions, any system-level configuration, the entire conversation so far, retrieved documents, tool descriptions, tool results and the response being generated all draw on the same budget. Some of it is spent before you have typed anything.

This is why an apparently enormous limit can feel smaller in use. A window that comfortably holds several documents holds far fewer once each turn of the conversation, and everything the system inserts behind the scenes, is counted too.

Cost does not grow gently with length

The mechanism at the heart of these models lets every position in the sequence attend to every other position, so the work involved grows with roughly the square of the length rather than in proportion to it. Doubling the input can therefore roughly quadruple the computation for that stage.

A great deal of engineering has gone into softening that curve, with approximations, caching of previous computation and architectural variants that restrict which positions attend to which. These help substantially and none of them makes length free.

So a long window is expensive in a way that is not always visible when the pricing is quoted per unit of input. Filling a very large window has real cost, and there is a reason providers price long inputs the way they do.

Capacity is not the same as attention

The most persistent misunderstanding is that a model uses a long input evenly. It does not. Material near the beginning and the end of a long input reliably exerts more influence than material in the middle, an effect robust enough across systems to be treated as a property of the approach rather than a defect of any one implementation.

The consequence is that fitting a document into the window does not guarantee the model has used it. A detail buried in the middle of a very long input may as well not be there for some queries, and nothing in the output indicates that it was skipped.

This is why evaluations that test whether a model can find a specific inserted fact are informative but incomplete. Finding one distinctive sentence is easier than integrating information spread across the whole input, and the two capabilities improve at different rates.

Bigger windows changed the architecture of applications

When windows were small, any application handling substantial material had to split it, retrieve the relevant pieces and assemble them, which meant building a retrieval pipeline before anything else worked. Larger windows removed that requirement for many use cases, and a great deal of machinery became optional.

The trade is now a real engineering decision rather than a constraint. Placing everything in the window is simpler, uses no separate infrastructure and costs more per request. Retrieving selectively is cheaper per request, more precise if the retrieval is good, and introduces a component that can fail on its own terms.

Which is preferable depends on how often the material changes, how large it is and how much precision matters. Anyone announcing that long windows have made retrieval obsolete is describing one set of circumstances as though it were all of them.

Reading the quoted number sensibly

Treat it as a ceiling on what fits rather than as a claim about what will be used well. Performance on tasks requiring genuine synthesis across a long input generally degrades before the limit is reached, and the point at which that begins is a property worth testing rather than assuming.

It is also worth remembering that the window is per request, not cumulative. A long conversation does not accumulate understanding; it accumulates text that must be re-read within the same budget, and once that budget is under pressure something gets dropped.

The number is real and it is an engineering specification, not a measure of comprehension. Those get conflated in nearly every announcement, and separating them explains most of the disappointment that follows. A larger window is a larger room, and nothing about a larger room guarantees that anybody looked in the corners.

Common questions

Why is the limit measured in tokens rather than words?

Because tokens are the units the model actually processes, and the relationship between tokens and words varies by language and by content. Ordinary English prose converts at a fairly steady ratio, while code, unusual names and many non-English scripts consume considerably more tokens for the same visible text.

Does a longer window mean better answers?

Only if the additional material is relevant and the model uses it. More input also means more opportunity for distraction, since irrelevant content in the window can pull a response off course. Adding context indiscriminately is not reliably an improvement.

What happens when a conversation exceeds the window?

It depends on the system. Some truncate the oldest material, some summarise it into a shorter form, and some simply refuse. Truncation and summarisation both lose information silently, which is why a long conversation can suddenly appear to forget something that was established early on.

Jargoncontext windowtokenslimitsterminology
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.