Jargon
Latency and throughput are different clocks and improving one usually costs the other
One measures how long a single person waits, the other how much work a system gets through, and almost every design decision about serving a model is a choice between them.
By Zoya Rahman3 min read

Two questions that sound like one
Ask how fast a system is and you have asked something ambiguous. Latency is the delay experienced by one request from the moment it is submitted to the moment it is answered. Throughput is the number of requests the system completes in a period. They are related, they are not the same, and a system can be excellent at one while being poor at the other.
A supermarket illustrates it well enough. Opening more tills raises throughput without shortening the queue at any individual till, and serving each customer faster shortens their wait without necessarily increasing the number served, since the constraint may be elsewhere entirely.
The distinction matters more here than in ordinary software because the trade between the two is unusually direct, and because the cost structure of running a model rewards the choice that annoys users.
Generation has two distinct waits
Text generation is sequential: each fragment is produced using everything already produced. So a response is not computed and then delivered. It is assembled piece by piece, and this gives the user two separate experiences of delay.
The first is the wait before anything appears, which covers reading the input, processing all of it, and producing the first fragment. This scales with how much input there is, which is why a long document pasted into a conversation produces a noticeable pause before the reply begins.
The second is the rate at which text then arrives. That is roughly constant per fragment and mostly a property of the hardware and the model’s size. Streaming output exists precisely because the second wait can be hidden behind reading speed, while the first cannot be hidden at all.
Batching is where the trade becomes explicit
The hardware these models run on is at its most efficient when performing the same operation across many sets of numbers simultaneously. Serving one request at a time leaves most of that capacity idle, which is why systems group requests together and process them as a batch.
Larger batches mean better use of expensive hardware and therefore lower cost per request, which is throughput. They also mean a request may wait for the batch to assemble, and shares the machine with others, which is latency. The operator sets that dial, and the setting is an economic decision rather than a technical inevitability.
This is why the same model can feel quick or sluggish depending on where you access it, and why cheaper access tiers are often slower. Nothing about the model changed. The batching policy did.
Memory movement, not arithmetic, sets the pace
A detail that surprises people: generating text is usually limited by how fast parameters can be moved from memory into the processor, rather than by the arithmetic itself. Producing a single fragment requires touching a great many parameters and doing relatively little with each.
This explains several otherwise puzzling facts. Compressing a model to use fewer bits per parameter speeds up generation, because less data has to be moved. Batching helps because the same parameters, once fetched, serve many requests at once. And processing a long input is comparatively efficient, since it is arithmetic-heavy in a way generation is not.
It also sets a floor. Once you are limited by memory bandwidth, faster arithmetic buys nothing, which is why hardware development in this area has concentrated so heavily on memory rather than on raw computational speed.
Why the vocabulary is worth getting right
Claims about speed are frequently made without saying which quantity is meant, and the two support very different conclusions. A system described as handling an impressive volume may still keep each user waiting; a system with an impressively quick response may be doing so by running well below capacity at considerable expense.
For anything interactive, latency dominates the experience in a way that is well established in interface research: delays past a fraction of a second are noticed, and past a few seconds attention drifts. For anything processed in bulk overnight, latency is nearly irrelevant and cost per unit of work is everything.
So the useful question is not how fast a system is, but which clock is being reported and which one your use actually depends on. Those are frequently not the same clock, and the gap between them explains a great many disappointed expectations.
Common questions
Why does a long input slow down the first response?
Because the whole input must be processed before any output can be produced, and that processing grows with input length. Once generation starts, the rate of arriving text is largely unaffected by how long the input was, which is why the delay appears as a pause at the beginning rather than as slower text.
Does a smaller model always respond faster?
Usually, since fewer parameters mean less data to move for each fragment produced, but the system around the model matters too. Queueing, batching policy and network overhead can dominate, and a small model on a congested service can easily feel slower than a large one on an uncongested one.
Is streaming output just a presentation trick?
It is a genuine improvement in perceived responsiveness rather than in total time. The complete response takes just as long, but the wait becomes readable rather than blank, which changes the experience substantially. Whether that counts as a trick depends on how strictly you define one.
Deputy editor, AI Worth Knowing
Zoya joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and would rather show the working than assert the conclusion.





