Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Training data has an end date and the model cannot feel the edge of it

Everything a system knows was fixed at some point in the past, the boundary is blurrier than a single date suggests, and the system has no sense of standing at it.

By Manish Trivedi3 min read

Individual viewing a laptop displaying a cracked and colorful digital screen indoors.
Photograph by Beyzanur K. via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

A cutoff is a hard boundary experienced as nothing at all

Training happens on a collection of text assembled up to some point, and after that point the parameters are frozen. Anything that happened later is simply absent — not marked as unknown, not represented as a gap, just not there. A person who has been out of contact knows they have been out of contact. A model has no equivalent awareness.

The consequence is that questions about recent events get answered from whatever the statistics support, which usually means the state of affairs at the end of training presented in the present tense. Prices, office holders, product versions, legal rules and best practice in fast-moving fields all get reported as current when they are not.

This is a different failure from fabrication, though it looks similar from outside. The information was true. It has expired, and nothing in the system tracks expiry.

The boundary is fuzzier than a single date

A cutoff date implies a clean line, and the reality is more gradual. Data collection happens over an extended period, different sources are gathered at different times, and material about an event keeps being written long after the event, so coverage of anything recent is thin relative to how it will eventually be covered.

That produces a decay rather than a cliff. Knowledge of things well before the boundary is dense and reliable; knowledge of things shortly before it is patchy, because the internet had not finished writing about them yet. A model may know that something happened without knowing how it turned out.

Additional training stages after the main run complicate it further. A system may be adjusted with more recent material for specific purposes, so its behaviour reflects several different vintages at once, and a single stated date describes the arrangement only roughly.

Models are unreliable narrators of their own vintage

Asking a system when its training data ends produces an answer, and that answer is generated the same way every other answer is. It may come from an instruction supplied at the start of the conversation, in which case it is as accurate as that instruction. It may be inferred from the training data itself, in which case it tends to be wrong.

The characteristic error is underestimation. Because recent material is thinly represented, the model’s internal sense of the present is dragged backwards towards the period that is densely covered, and it may report a date noticeably earlier than the actual boundary.

The practical implication is that this is a fact to look up in documentation rather than to ask about. It is one of a small number of questions where the system is systematically the wrong source about itself.

Retrieval moves the problem rather than removing it

The standard remedy is to fetch current material and supply it alongside the question, so the answer is grounded in something dated rather than in frozen parameters. This works well and it is why so many deployed systems now search before answering.

It introduces a new set of questions, all of them about the retrieval rather than the model. Is the source current? Is it authoritative? Did the search find the relevant document or a superficially similar one? When retrieved material conflicts with what the parameters encode, which wins — and the answer is not consistent, which is uncomfortable.

There is also a subtler issue. Retrieved text competes for attention with everything else in the input, and a strongly held prior from training can override a document that says otherwise, particularly when the document is terse and the prior is dense. This is an active area of engineering rather than a solved one.

A slower question about what gets trained on next

Since these systems train on publicly available text, and a growing share of publicly available text is machine-generated, later systems will train partly on the output of earlier ones. What follows from that is genuinely disputed and worth marking clearly as unsettled.

One line of argument holds that repeatedly training on generated output degrades a model, narrowing the range of what it produces as rare patterns drop out of successive generations. Studies demonstrating this effect exist, generally under conditions where generated data replaces original data wholesale.

The counter-argument is that nobody trains that way in practice. Curation, filtering, mixing with fresh human material and deliberately generated training data are all standard, and carefully constructed synthetic data has improved systems rather than degrading them. Which dynamic dominates over years is not something anyone can currently claim to know, and treating either scenario as established goes beyond the evidence.

Common questions

Why does a system sometimes know about a recent event and sometimes not?

Because coverage near the boundary is uneven rather than absent, and because many deployed systems search the web for some queries and not others. From the outside these two explanations are hard to tell apart, which is why an interface that indicates when it has searched is more useful than it looks.

Does a new version of a model relearn everything?

Usually a new major version is trained afresh rather than patched, which is why capability and knowledge tend to move together in steps. Smaller updates adjust behaviour through additional training on top of existing parameters, and these can change tone and habits considerably without much changing what the system knows.

Is stale information more dangerous than no information?

Frequently, yes, because it arrives with the same confident presentation as current information and offers no cue to check. This is most serious in domains where rules change on a schedule — tax, benefits, immigration, medical guidance and safety standards — and it is a good reason to treat any date-sensitive answer as a starting point for verification.

Limits & Risksknowledge cutoffstalenessretrievaldata
Manish Trivedi
Consumer editor, AI Worth Knowing

Manish has written about how it works, in the world, limits & risks for most of the last decade and prefers a plain explanation to a clever one.