Jargon
Inductive bias is what a model assumes before it has seen anything
Every learning system carries built-in preferences about which explanations to favour, learning is impossible without them, and the choice of which ones to build in is most of what model design consists of.
By Samar Bhatia3 min read

Learning from examples requires an assumption to be possible at all
Given a set of observations, infinitely many rules fit them and disagree about everything else. A system with no preference between those rules cannot generalise, because it has no basis for choosing one over another. This is not a practical difficulty; it is a logical one, and it was established formally long ago.
What resolves it is a preference built in beforehand: simpler explanations over complicated ones, smooth relationships over jagged ones, patterns that hold across positions over patterns that do not. Any such preference is an inductive bias, and it is the thing doing the generalising.
The word bias is unfortunate, because in ordinary use and in fairness research it means something undesirable. Here it means a prior commitment, and it is not merely acceptable but required. A system without one cannot learn anything from finite data.
The assumptions live mostly in the architecture
Some are explicit. A model applying the same detector across every position of an image assumes that identity does not depend on where something appears. A model processing a sequence in order assumes that recent items matter more than distant ones. A model treating an input as a set assumes order carries no information.
Others are subtler and live in the training procedure. The habit of preferring small parameter values, the choice of how to initialise, the tendency of the optimiser to settle in certain kinds of solution rather than others — all express preferences about what a good explanation looks like without anybody stating them.
This is why architecture matters even though every architecture is ultimately arithmetic. Two models with the same number of parameters, trained on the same data, will generalise differently because they favour different explanations of it.
Stronger assumptions need less data and constrain more
The trade is direct. A model that assumes a great deal about the problem can learn from few examples, because most of the answer was supplied by the assumption. It will also fail on any problem where the assumption does not hold, and will fail confidently, since it has no representation of the alternative.
A model that assumes very little needs correspondingly more data, since it must discover from examples what the other model was given. In exchange it is not restricted to problems fitting a particular mould, which matters when the mould is wrong or unknown.
Neither position is superior. Which is right depends on how much data exists and how well the assumption matches reality, and both of those are empirical questions rather than matters of principle.
The argument about whether to build assumptions in
A widely discussed position in the field holds that historically, methods relying on general learning and large computation have eventually overtaken methods encoding human knowledge about the problem, and that effort spent on the latter is usually wasted. The argument is supported by several decades of examples across vision, speech and games.
The counter-argument is that the successful general methods still contain substantial built-in assumptions, that those assumptions were what made them work, and that the position is therefore an argument about which assumptions rather than about whether. It also notes that the general approach requires data and computation that are not always available.
This is a live disagreement among serious people, and it is not settled by pointing at recent results. Where data is abundant the general approach has clearly done better; where it is scarce, structure still earns its keep, and most of the world’s problems have scarce data.
Why the term is worth knowing
It names the reason two models behave differently on data neither has seen, which is otherwise mysterious. It also explains why a model can be excellent on typical cases and wrong in a specific, consistent way at the edges: the assumption that carried it through the middle stops being true.
It clarifies what a training set does and does not determine. Data constrains a model; it does not select among the explanations consistent with it. Something else does that, and the something else was chosen by whoever designed the system, usually without stating it as a choice.
And it supplies a better question than most about any system: not how much data it saw, but what it was built to assume about the problem before it saw any.
Common questions
Is a model with fewer assumptions more objective?
No. It has different assumptions, generally weaker and more general ones, and it still has them. The idea of learning without prior commitments is not achievable, which is the substance of the formal results on this.
How does this relate to bias in the fairness sense?
They are different concepts sharing a word. Inductive bias is a prior preference among explanations and is necessary. Bias in the fairness sense concerns systematically unequal outcomes, and while a poorly chosen inductive bias can contribute to it, the terms are not interchangeable.
Do very large general models have weak inductive biases?
Weaker than purpose-built architectures, and far from absent. The way they process sequences, the way they are trained and the composition of what they are trained on all express strong preferences about what patterns to find.
Senior writer, AI Worth Knowing
Samar has been reporting on how it works, in the world, limits & risks since long before it was fashionable and would rather show the working than assert the conclusion.





