Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

Training on your data means several different things, and only one is training

The phrase covers a contractual question about permission and a technical question about whether any parameter changed, and collapsing the two makes both considerably harder to reason about.

By Naina Sethi3 min read

A bearded man in suit playing chess with robotic arm, showcasing AI strategy.
Photograph by Pavel Danilyuk via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

Three arrangements share one sentence

When a service says it uses your data, at least three distinct things may be meant. The text you sent may simply have been included in the request and then discarded. It may have been stored in a searchable index that the system consults later. Or it may have been added to a corpus used to train a future model.

These differ in nearly every respect that matters: how long the data persists, who can see it, whether it can be removed, and whether it influences anybody else’s results. Only the third alters a model’s parameters, and only the third is training in any technical sense.

The vocabulary in public discussion does not distinguish them, which is unfortunate, because most anxiety attaches to the third while most actual processing is the first two. A person worried about the wrong one may take precautions that address nothing.

Passing through and being absorbed are different events

Text supplied in a request is read by a frozen model, influences the output, and then has no further effect on the system. Nothing is retained inside the parameters, because inference does not write anything back. The model that answers the next person is byte-for-byte the model that answered you.

A retrieval index is storage, plainly. The material sits in a database, it can be searched, it can usually be deleted, and it is available to whoever the access controls permit. This is an ordinary data-protection question of the sort organisations have handled for decades, and it should be reasoned about as one.

Training is the only case where the material is absorbed into a model in a form that cannot be located afterwards. Once fitted, contributions from any individual document are distributed across parameters with no record of origin, which is why removing them later is a research problem rather than an administrative task.

Retention is a separate axis from training

Services commonly keep request logs for a period regardless of whether they train on them, for debugging, billing, abuse investigation and legal obligation. Human review of a sample is also standard practice in operating any large system, and it is disclosed with varying prominence.

So a product can honestly say it does not train on customer content while retaining that content for months and permitting staff to read some of it. Both statements are true. They answer different questions, and the one people usually care about is not always the one being answered.

The reverse also occurs: material may be used to train while being retained only briefly in raw form. Retention period, access, and training use are three separate settings, and knowing one tells you very little about the other two.

Whether these arrangements require permission, what form the permission takes, and whether an opt-out must be offered are legal questions whose answers differ by jurisdiction, by the category of data and by whether the account is personal or organisational. They also change. Anyone with a real obligation here needs advice about their own situation rather than a general article.

What can be said structurally is that defaults do most of the work. Where training use is enabled unless the user changes a setting, the vast majority of data ends up in scope, because most people never open the settings. Where it is disabled by default, the reverse holds. This is a design decision with larger consequences than the policy wording.

Business agreements often differ from consumer terms on precisely this point, which is why the same underlying system can carry two quite different sets of commitments depending on how it was purchased.

Deletion does not run backwards through a training run

If material was used to train a model that has already been built, deleting the source record does not remove its influence from that model. The usual remedy is to exclude it from future training runs, which means the effect persists for as long as the existing model is deployed.

This is uncomfortable against the way data protection is normally framed, where deletion is supposed to be effective and verifiable. Techniques for removing specific influence after the fact exist and are an active research area; their reliability is contested, and demonstrating that removal succeeded is harder than performing it.

The honest summary is that the three arrangements have three different reversibility properties, and the least reversible one is the one hardest to observe from outside. That asymmetry, more than any particular company practice, is what makes the topic difficult.

Common questions

Does a chatbot remember what I told it earlier?

Within a session it is being re-sent the earlier text, not remembering it. Across sessions, anything that persists is stored in an ordinary database associated with your account, which is storage rather than learning. Neither case changes the model, and both are governed by the service’s retention settings.

If a service says it does not train on my data, is that the whole picture?

It answers one of three questions. It does not tell you how long content is retained, whether staff or contractors may review samples, or whether the content is indexed for later retrieval. Those are separate commitments and worth looking for separately in whatever documentation exists.

Can data used in training ever be traced back to a person?

Not by inspecting the model, which contains no index of sources. Fragments of memorised text can sometimes be extracted, particularly if a passage appeared many times in the corpus, so the possibility is not zero. Records of what was included generally live in the training pipeline rather than in the model itself.

In The Worlddataprivacyretentionpolicy
Naina Sethi
Features writer, AI Worth Knowing

Naina joined to cover how it works, in the world, limits & risks and stayed for the awkward questions and is unreasonably interested in the detail nobody else checks.