How It Works
Predicting the next fragment is a narrower job than the results suggest
The training objective is almost embarrassingly simple, and the interesting argument in the field is about how much that simplicity can be made to carry.
By Daniel Okonkwo3 min read

The objective fits in one sentence
A language model is trained to guess what comes next. Text is fed in, the model produces a probability for every entry in its vocabulary, the true continuation is revealed, and the parameters are adjusted so that the right answer would have been rated a little higher next time. Repeat across an immense quantity of text.
There is no separate stage where facts are stored, no module for grammar and none for reasoning. Whatever the system ends up doing, it acquired while trying to reduce its uncertainty about what word came next in documents written by other people.
That is worth stating plainly because it sounds too thin to explain the results, and the tension between how thin it sounds and how much it produces is the most interesting open question in the field.
Prediction quietly demands a great deal
Consider what accurate prediction actually requires. To guess the last word of a murder mystery’s final paragraph, something must track who was where. To continue a proof, something must respect the rules of the argument. To finish a sentence in the middle of a translated passage, something must hold both languages simultaneously.
None of these capacities were requested. They are instrumental: the only way to drive prediction error down past a certain point is to represent regularities in the world that the text describes, because the text is about the world. That argument is the strongest case for the approach, and it is a genuinely serious one.
The counter-case is equally serious. Text is not the world, it is a compressed and heavily selected record of what people bothered to write down, and a system that models the record perfectly has modelled the record. Where those diverge — in physical intuition, in things too obvious to state, in anything that happened but was never written — the model has no source.
The output is a distribution, not an answer
At each step the model produces a score for every possible next fragment, and something else has to choose one. Always taking the highest-scoring option produces flat, repetitive text that tends to fall into loops, which is a real and well-known failure rather than a matter of taste.
So the choice is usually made by sampling from the distribution, with settings that control how much of the tail is considered. Turn those settings down and output becomes predictable and dull; turn them up and it becomes varied and less reliable. The knob is doing exactly what it appears to do, and there is no setting that gives both.
This is also the entire explanation for why the same input yields different answers on different occasions. Nothing has changed in the model. A different sample was drawn, and one early divergence propagates through everything that follows.
Imitating text is not the same as being right
The objective rewards text that resembles what a person would have written, and confident writing is far more common in the record than hedged writing. Documents that reason toward a conclusion vastly outnumber documents that stop halfway and admit defeat. Fluency and assurance are therefore learned along with everything else, whether or not the content behind them is sound.
Later training stages adjust this using human feedback, and they change behaviour substantially. They also introduce their own distortions, since what human raters approve of is not identical to what is true, and agreeable answers reliably score well with people.
None of this is a defect in the implementation. It is the objective functioning as specified. The specification simply never mentioned truth.
Where the disagreement actually sits
One camp holds that scaling this objective further, with better data and more compute, continues to yield capabilities that were never explicitly trained for, and that there is no principled ceiling in sight. They can point at a decade of predictions about what this approach would never do, most of which aged badly.
The other camp holds that prediction over text has a ceiling that no amount of scale removes, because certain kinds of knowledge are simply absent from the training signal — grounding in the physical world, the ability to run an experiment, a stable notion of what is true rather than what is commonly written. They can point at failures that persist across every size of model tried.
Both are arguments about the future and neither is established. Anyone reading a forecast here should note that the field’s own record of forecasting is poor in both directions, and that confident timelines have been a recurring feature of this discipline since it began.
Common questions
Is the model choosing words one at a time?
Yes, and each choice is added to the input before the next is made, so the system commits as it goes. It has no mechanism for planning a whole answer and revising it before you see it — anything resembling revision is later text correcting earlier text in the same stream.
Why does the same question sometimes get a much worse answer?
Partly the sampling, which introduces genuine randomness, and partly how the question was phrased, since small wording changes shift the distribution the model draws from. A weaker answer is not evidence that the system knows less on that occasion; it took a different path.
Does next-fragment prediction rule out reasoning?
That is precisely what is contested. Models demonstrably perform multi-step reasoning on many problems and demonstrably fail on others that look no harder, and the field has not agreed on whether what happens in between deserves the word. Treat anyone certain in either direction with some caution.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





