Limits & Risks
A model cannot be made to forget, and that breaks how deletion is supposed to work
Information from training is distributed across billions of parameters rather than stored in a record, so removing it after the fact is a research problem rather than an administrative one.
By Daniel Okonkwo3 min read

Deletion assumes a place where the thing is kept
The right to have personal data removed, which exists in various forms in several legal systems, rests on an implicit picture of how data is held: there is a record, it sits in a database, and it can be found and deleted. For conventional systems this picture is accurate and the operation is routine.
A trained model does not work that way. During training, text influences the adjustment of parameters, and the influence of any one document is spread thinly across an enormous number of values that also encode everything else. There is no row to delete and no location to inspect.
So the question of whether a model can comply with a deletion request is not evasion. It is a genuine technical difficulty that the law was not written with in mind, and the various proposed answers are all partial.
Memorisation is real, uneven and hard to predict
Models generally do not reproduce their training data, because the objective encourages generalisation rather than storage. But they sometimes do, and the pattern of when has been studied enough to state some regularities: repetition in the corpus matters a great deal, unusual or highly structured strings are more prone to it, and larger models memorise more readily than smaller ones.
The uneven part is what makes this awkward. A person whose details appear once in an obscure document is unlikely to be reproducible. The same details appearing across many scraped pages become far more so, which means exposure correlates with how widely the information already circulated rather than with how sensitive it is.
Establishing whether a particular piece of information is recoverable from a model is itself difficult. Absence of evidence from a few thousand attempts at extraction is weak evidence of absence, given the size of the space of possible prompts.
Retraining is the reliable answer and an impractical one
The one method that certainly removes a document’s influence is training again without it. That works and it costs what training costs, which places it out of reach as a response to individual requests arriving continuously.
This is not merely a matter of expense. Retraining produces a different model with slightly different behaviour everywhere, which then needs re-evaluating and re-deploying. Treating it as a routine operation misunderstands what a training run is.
Various cheaper approaches go under the heading of machine unlearning: methods that attempt to remove the influence of specific data by targeted adjustment. Some show real progress on well-defined cases. None currently offers the guarantee that a deletion regime implies, and evaluating whether an unlearning method worked is an unsolved problem in itself.
Filtering the output is what usually happens instead
The practical response in deployed systems is to intervene at the edges: block the model from producing certain material, filter outputs containing personal details, or refuse categories of request. This addresses the visible symptom and it is not the same as removal.
A filter can be bypassed, can fail on phrasings it did not anticipate, and does nothing about the information still latent in the parameters. If the weights are ever released, or the model accessed through a different route, the filter is not there.
It is worth being clear that this is a mitigation presented, sometimes, as a remedy. The distinction matters for anyone reasoning about risk rather than about compliance paperwork, because the two lead to different decisions about what a system should be allowed to do with sensitive material in the first place. A filter is a control on one route of access. It is not a statement about what the model contains.
Prevention is the only stage where control is straightforward
Everything becomes easier before training. Filtering personal information out of a corpus, deduplicating aggressively to reduce memorisation, and adding formal noise to the training process are all things that work at the point where the data is still identifiable and separable.
Each has costs. Corpus filtering is imperfect at scale and removes legitimate material along with the rest; noise-based methods trade measurable privacy protection against model quality, and the strength of protection is a dial rather than a switch. These are real engineering trades with no free option.
The structural lesson generalises beyond privacy law. With this technology, decisions made before training are cheap and decisions made afterwards are expensive or impossible, and a great deal of what people expect to be adjustable later simply is not.
Common questions
If my information is in a model, can it be removed?
Not with certainty by any method short of retraining without the source data. Providers may filter outputs, and targeted unlearning methods exist and are improving, but neither currently provides the assurance that deleting a database record does. This gap between legal expectation and technical capability is unresolved.
Does a model store copies of the pages it trained on?
Not as copies. The parameters hold a compressed statistical residue, which is why models usually paraphrase rather than reproduce. Material that appeared very frequently in the corpus can nonetheless be reproduced closely, and that is the exception the whole debate turns on.
Do conversations become part of a model?
Not automatically and not immediately. Whether conversation data is retained and used in future training depends on the provider’s terms and any settings offered, and these vary considerably between services and change over time. It is one of the few things here where reading the policy actually answers the question.
Contributing editor, AI Worth Knowing
Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.





