Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

A hyperparameter is a setting somebody chose, and the choosing is where results come from

The division between values a model learns and values a person picks is the most consequential distinction in the vocabulary of training, and it is the one least often explained.

By Imran Sheikh3 min read

A detailed close-up of a Bible page showing scripture text and verses.
Photograph by Brett Jordan via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Two kinds of number, and only one of them is learned

Inside any trained system there are numbers adjusted automatically by the training process, and numbers fixed before training began that the process cannot touch. The first are parameters. The second are hyperparameters, and the prefix is doing exactly the work it appears to: they sit above the learning rather than inside it.

The list is long and mundane. How large a step to take when adjusting parameters, how many examples to process together, how many layers, how wide each one is, how much randomness to inject to discourage memorisation, how long to train. None of these is discovered by the model; each is a decision.

The distinction is sharp and the boundary is negotiable. Methods exist that learn things ordinarily chosen by hand, and every such method introduces new settings of its own. The category is defined by which loop a value sits outside, not by the nature of the value.

The settings matter more than the vocabulary suggests

It’s tempting to treat these as engineering details beneath the interesting questions. They are not. The same architecture trained with a poor choice of step size may fail to converge at all, and a good choice may make an ordinary design outperform a clever one. Comparisons between methods are frequently comparisons between tuning efforts.

This has produced a recognised problem in the literature. A new method is proposed, carefully tuned by its authors, and compared against a baseline tuned considerably less carefully. Several published re-examinations have found that established baselines, given equivalent attention, close much of the reported gap or all of it.

The effect is not usually deliberate. Researchers spend their time on the thing they are proposing, which is entirely natural and produces a systematic bias in favour of whatever is new. Where reviewers now ask about tuning budgets, they are addressing this directly.

Finding good values is a search with no gradient

The obvious method is to try combinations and keep the best. Testing every combination on a grid is thorough and wasteful, because most settings do not matter much and the grid spends equal effort on all of them. Random sampling within sensible ranges was shown to be more efficient for exactly that reason and became standard practice.

More elaborate approaches build a model of how settings relate to results and choose the next trial to be informative, which is a sensible idea that pays off when each trial is expensive. Others start many runs and stop the unpromising ones early, spending the budget on what looks likely to work.

What none of them avoids is that every trial requires training a model. For small models this is routine. For very large ones a full search is impossible, and choices are made by extrapolating from smaller runs and by accumulated craft, which is a polite word for informed guessing.

Tuning is where a test set gets contaminated quietly

The purpose of holding data back is to obtain an honest estimate of performance on cases the model has not shaped itself around. Choosing settings by looking at that held-back data violates the arrangement, because the settings then carry information from it and the estimate is no longer independent.

This is why a third split exists. Training data fits the parameters, a validation set guides the choice of settings, and a test set is looked at once at the end. The discipline is well understood and frequently eroded, since running a few more experiments after seeing the test result is an easy thing to do and hard for anybody outside to detect.

Repeated over a community sharing a public benchmark, the effect accumulates. Everybody tunes against the same test set through the papers they read, and the benchmark gradually stops measuring what it measured at the start.

Reading the term in the wild

When a result is reported without mention of how settings were chosen, the useful assumption is that a search happened and its budget is unstated. That does not make the result wrong; it makes the comparison less informative than it appears, particularly where the margin is small.

It is also worth distinguishing this from the settings a user adjusts at the point of use, which are sometimes loosely called by the same name. Those change the behaviour of a finished model and change nothing about what it learned, which is a completely different kind of decision.

The general point is that a trained model embodies a great many human choices that are invisible in the finished artefact. The word names them, which is most of its value.

Common questions

Can hyperparameters be learned automatically?

To a degree, and the methods that do so are widely used. They shift the difficulty rather than removing it, since each such method has settings of its own and each still requires training runs to evaluate candidates.

Why is the step size singled out as important?

Because it controls how far parameters move on each update, and both extremes fail. Too small and training is impractically slow or stalls in a poor region; too large and the process oscillates or diverges. The workable range is often narrow and problem-specific.

Do very large models require the same tuning?

They require the same decisions and cannot afford the same search, so the values are typically transferred from smaller experiments or from published practice. This is one of the less examined sources of uncertainty in large training runs.

Jargonterminologytrainingtuning
Imran Sheikh
Editor, AI Worth Knowing

Imran has written about how it works, in the world, limits & risks for most of the last decade and thinks most subjects are more interesting once you know how they work.