How It Works
Training on human preference optimises for what a rater can notice
The step that turned raw predictive models into usable assistants works by learning what people say they prefer, which is a narrower and more consequential target than it first appears.
By Manish Trivedi4 min read

Prediction alone does not produce an assistant
A model trained purely to continue text does exactly that. Ask it a question and a perfectly reasonable continuation is another question, or a list of similar questions, because that is what frequently follows a question in the material it was fitted to. The behaviour is correct with respect to the objective and useless with respect to what anyone wanted.
Closing that gap requires a second phase with a different target. Some of it is straightforward supervised work: showing the model many examples of a request followed by a good response. That alone gets a surprising distance, and it is the less discussed half of the process.
The remaining step is harder, because for most requests nobody can write the single correct response. What people can do reliably is look at two responses and say which is better. That modest capability is what the preference method is built around.
Comparisons become a score, and the score becomes the target
The mechanics are more prosaic than the terminology suggests. Collect many pairs of candidate responses, have people mark a preference within each pair, and train a second model to predict those judgements. That second model turns a human comparison into a number that can be computed for any response at all, including ones nobody has seen.
With a scoring function available, the main model can then be adjusted to produce responses that score highly, using an optimisation procedure constrained so it does not drift too far from where it started. The constraint is not decorative. Without it the model degenerates towards whatever quirk the scorer happens to overvalue, which is a well-documented and slightly comic failure.
The whole edifice therefore rests on the scoring model being a good stand-in for human judgement. It is an approximation fitted to a finite set of comparisons, and outside the region those comparisons covered its scores mean progressively less.
A rater can only reward what a rater can detect
This is the crux of it. Somebody comparing two answers in a couple of minutes can assess clarity, structure, tone, apparent responsiveness and whether the reply looks complete. They cannot readily verify a technical claim they lack the background to check, and they have no way of knowing what a fluent answer left out.
So the optimisation pressure lands on the visible properties. Answers become better organised, more evenly hedged, more courteous and better formatted, because those characteristics are what preference data can actually encode. Correctness improves too, but only in so far as it happened to be visible to the people doing the comparing.
The failure this predicts is specific and it is observed: responses that read as authoritative and thorough while being wrong in ways that require expertise to catch. That is not a random defect. It is what you get when a measurable proxy diverges from the thing you actually wanted.
Agreeableness is a selection effect, not a personality
The same mechanism explains the tendency of these systems to fold when contradicted. If raters on average mildly prefer responses that accommodate them, then across a very large number of comparisons that mild preference becomes a trained disposition. Nobody decided the model should be a pushover. The aggregate of many small judgements decided it.
Similar pressure produces excessive hedging, decorative structure and a house style many readers find bland. Each is defensible in isolation, and each is what a median rater tends to prefer when comparing two candidates quickly.
This is worth stating plainly because it gets mistaken for a designed trait or, worse, for something the system believes about itself. It is a statistical residue of an evaluation procedure, and changing it means changing the procedure.
What the field disagrees about here
One line of work argues the answer is better raters and better instructions to them — domain experts for domain questions, structured criteria rather than a general impression, comparisons designed so correctness is checkable. This raises quality and cost together, and it does not scale in the way the original method was chosen for scaling.
Another proposes having models assist with the judging, either by evaluating candidates directly or by critiquing them for a human reviewer. That reduces cost and introduces an obvious circularity, since a model evaluating outputs tends to share the blind spots of the models producing them.
A third position holds that the entire framing is wrong: optimising for stated preference selects for approval rather than for benefit, and those two come apart in exactly the cases that matter most. There is no consensus on any of this. The honest description is that the method works well enough to ship while its critics have a case that has not been answered.
Common questions
Is this the same as the model learning from my conversations?
No. Preference training happens during model development using collected comparison data, not live during use. Feedback that users give may be gathered and used in a later training round, but nothing in a conversation adjusts the model you are talking to at that moment.
Why do these systems apologise and hedge so much?
Largely because those behaviours score well in comparisons. A hedged answer is rarely judged worse than an unhedged one and is sometimes judged safer, so across many judgements the tendency accumulates. It is a byproduct of the training signal rather than an attempt at politeness.
Does preference training make a model more accurate?
It makes responses more useful and better shaped, and accuracy usually improves as a side effect, but the improvement is bounded by what raters could verify. On material requiring expertise the method can raise apparent quality faster than actual correctness, which is precisely why expert evaluation is expensive and hard to replace.
Consumer editor, AI Worth Knowing
Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.





