How It Works
Averaging several models beats polishing any one of them, up to a point
Combining independently trained predictors reliably improves accuracy, the reason is a piece of arithmetic rather than anything clever, and the practice quietly dominated applied machine learning for a decade.
By Daniel Okonkwo4 min read

Errors that disagree can cancel; errors that agree cannot
Take several models trained on the same problem and average their predictions. If they make the same mistakes, the average makes those mistakes too and nothing improves. If they make different mistakes, the errors pull in different directions and partially cancel, while the correct signal, being common to all of them, survives.
That is the whole idea, and it is arithmetic rather than insight. The gain depends on how uncorrelated the errors are, which is why effort goes into making the members differ: different random starting points, different subsets of the data, different subsets of the input features, sometimes entirely different kinds of model.
It also sets the ceiling. Members that agree perfectly give no benefit whatsoever, and members chosen to be so different that most of them are bad will drag the average down. The useful region is a set of individually decent models that happen to be wrong in different places.
Two families, built on opposite instincts
One approach grows the members independently. Each is trained on a random sample of the data drawn with replacement, so every member sees a slightly different world, and the results are averaged or voted on. Decision trees combined this way became one of the most widely used tools in applied statistics, largely because it works without much tuning.
The other grows them in sequence, each new member trained to correct what its predecessors got wrong. Members are deliberately weak on their own and specialised towards the residual errors of the group. This tends to produce stronger results and is fussier to configure, since it will happily chase noise if nothing stops it.
The two families answer different questions. The first reduces the variance that comes from having only a finite sample of data. The second reduces bias by adding capacity exactly where the current combination is failing. Practitioners argue about which to reach for first, and the honest answer is that it depends on the data.
This is what actually won, for years
Through the period when public modelling competitions were a visible part of the field, winning entries were overwhelmingly ensembles, often large and inelegant ones. Combining a dozen models with different preprocessing routinely beat any single carefully constructed model, and the margin was small but stubbornly consistent.
The lesson taken from that was mixed. Competition scoring rewards the last fraction of a percentage point, and a fifty-model blend that captures that fraction is usually unusable in production. So the technique earned a reputation for being effective and impractical at the same time, which was fair enough.
On structured tabular data — rows and columns, the kind of problem most organisations actually have — ensembles of trees remain very hard to beat. Neural approaches have made repeated attempts, and the results are genuinely contested, with published comparisons landing on both sides depending on which datasets were chosen.
The cost is paid on every single use
An ensemble of ten models costs ten times as much to run as one of them, and occupies ten times the memory. For a batch job scored overnight that may be irrelevant. For a system answering interactive requests at volume it is often decisive, and the accuracy gain has to justify the multiplied bill.
This is where the technique meets its usual fate: somebody trains an ensemble to find out how well the problem can be solved, then trains a single model to imitate it, and deploys that instead. The imitation captures much of the benefit at a fraction of the cost, which is a reasonable trade rather than an equivalent one.
Where the cost per request is already large, as with big generative models, running several of them and combining the results is rarely attractive. The economics that made ensembling routine on small models make it exceptional on large ones.
The boundary between one model and several is blurrier than it looks
Some training techniques have been described as ensembling in disguise. Randomly switching off units during training, for instance, has been interpreted as training an enormous collection of overlapping sub-networks and averaging them at the end. The interpretation is intuitive and it isn’t universally accepted.
Averaging the parameters of several checkpoints, or of several separately trained models, sometimes improves results in ways ordinary intuition doesn’t predict, since the parameters of a neural network are not obviously the sort of thing that can be meaningfully averaged. It works often enough to be a standard trick, and it doesn’t always work.
One thing worth separating out: an architecture that routes each input to one of several internal sub-networks is not an ensemble, despite the vocabulary suggesting otherwise. It selects rather than averages, and it exists to save computation. Different mechanism, different purpose, confusingly similar name.
Common questions
Does combining several answers from one model count as an ensemble?
Loosely, yes — sampling several outputs and taking the most common one exploits the same cancellation of uncorrelated errors. The members are far less independent than separately trained models, since they share every parameter, so the benefit is usually smaller than a true ensemble of distinct models would give.
Why not build one bigger model instead?
Often that is the better use of the budget, and on problems with abundant data it usually is. Ensembles earn their place when data is limited, when the individual models are cheap to train and run, or when different model families make errors that are genuinely different in character.
Do ensembles make a system more trustworthy?
They tend to improve accuracy, and the disagreement between members gives weak evidence about uncertainty. They don’t correct a bias present in the training data, because every member learned from the same flawed record and will agree confidently on precisely the same mistake.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





