Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A batch is a group processed together, and its size is a real decision

The word turns up in training and in serving with different meanings, and in the first of those the number chosen affects both the speed of learning and, more contentiously, the quality of the result.

By Imran Sheikh3 min read

Close-up of an open braille book on a library table representing accessible reading.
Photograph by Thirdman via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Why examples are fed in groups at all

To take one step of training you need an estimate of how the error would change if each parameter moved. Computing that over the entire dataset gives the most accurate estimate and is prohibitively slow for anything large, since every step would require reading everything.

Computing it from a single example is fast and gives a very noisy estimate, pointing in roughly the right direction and wandering considerably. The compromise, which is what everybody actually does, is to compute it over a group of examples and average the result. That group is the batch.

The size of the group therefore trades estimate quality against steps per hour. It also happens to determine how well the hardware is used, since processors of this kind are efficient when given many similar operations at once and idle when given few.

The noise is not purely a nuisance

A smaller batch produces a noisier estimate of the direction to move, and that noise turns out to have a useful side effect. It stops the process settling into narrow configurations that fit the training data very precisely, because the jitter shakes it out of them, and configurations that survive the jitter tend to generalise better.

This is one explanation for a long-observed pattern in which very large batches train faster in wall-clock terms and sometimes produce models that perform slightly worse on unseen data. The observation is fairly robust. The explanation is not settled, and several accounts compete.

Some researchers argue the effect disappears once other settings are adjusted properly, and that the gap is a symptom of insufficient tuning rather than a property of batch size. Others hold that the noise is doing genuine work that cannot be recovered another way. Both positions have supporting results.

It interacts with everything else

Change the batch size and the appropriate step size changes with it, because averaging over more examples reduces the variability of the estimate and permits a larger, more confident move. Rules of thumb exist for scaling the two together, they work over a useful range, and they are not exact.

Very large batches also run into a limit where increasing the group stops improving the estimate meaningfully, because the direction is already well determined and additional examples mostly confirm it. Past that point the extra computation buys very little, and identifying where it lies is empirical rather than principled.

None of these interactions can be reasoned about independently, which is why training recipes are reported as complete configurations rather than as isolated choices. A number that works in one recipe may be poor in another that differs elsewhere.

The other batch, at the other end of the pipeline

The same word describes something different in serving. Requests from many users arriving at roughly the same time are grouped and pushed through the model together, because processing twenty inputs at once costs far less than twenty times processing one.

That grouping has nothing to do with learning and no bearing on the parameters. It is a scheduling technique, and its trade-off is between how efficiently the hardware is used and how long an individual request waits for the group to fill.

The homonym causes real confusion, particularly since both appear in discussions about cost. A larger training batch is a decision about how the model was built. A larger serving batch is a decision about how requests are queued today, and it can be changed this afternoon.

Why the term is worth pinning down

Batch size appears in nearly every published training configuration and is one of the numbers people copy without much thought. It is also one of the numbers that cannot be copied safely on its own, since its effect depends on the step size, the model, the data and the number of processors.

And it is a good illustration of a broader pattern in this field: a choice made for a practical reason, here fitting work into memory, turning out to have consequences for the quality of the result that nobody intended and that are still being argued about.

That pattern recurs often enough to be worth expecting. Several of the settings now treated as meaningful began as accommodations to hardware, and the theory explaining why they matter, where there is one, arrived long afterwards.

Common questions

Does a bigger batch always train faster?

Faster per step of progress through the data, up to the point where the hardware saturates or the estimate stops improving. Whether it reaches a given quality sooner in wall-clock terms depends on how well the other settings were adjusted, and on whether more processors are available to use the larger group.

Is batch size worth adjusting for a small project?

It is one of the settings most constrained by the hardware available, so in practice it is often chosen by what fits in memory rather than by preference. Where there is room to choose, moving it requires adjusting the step size alongside, or the comparison measures the wrong thing.

Why do serving costs mention batching so often?

Because it is one of the main levers an operator has over cost per request. Grouping requests improves the use of expensive hardware substantially, and it does so by making individual users wait slightly longer, which is why cheaper service tiers often carry noticeably slower responses.

Jargonbatchtraininggradientsterminology
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.