Skip to content
What the technology actually is
AI Worth KnowingWhat the technology actually is

These systems have no memory, and every workaround is something else wearing the name

A conversation that appears to build on itself is being reassembled from scratch each time, and the various features marketed as memory are storage bolted to a system that cannot store anything.

By Daniel Okonkwo3 min read

Detailed view of HTML and JavaScript code displayed on a computer monitor.
Photograph by Digital Buggu via Pexels
Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The model is the same before and after every exchange

Once training ends the parameters are fixed. Nothing that happens during use alters them, which means the system that answers your tenth message is numerically identical to the one that answered your first. Whatever continuity you experience is not stored inside the model, because there is nowhere inside the model for it to go.

What actually happens is more mundane. The entire conversation so far is fed back in as input on every turn, so the model reads the whole history afresh each time and produces the next response from it. The illusion of a thread being held is really a transcript being re-read at speed.

This is worth internalising because it explains a family of behaviours that otherwise look like carelessness — losing track of something established earlier, contradicting a stated constraint, forgetting a preference. There was never a thread to lose.

Recall degrades within a single conversation for a structural reason

Because the whole history is re-read, everything must fit in a bounded window. When it does not, something has to go, and the usual approaches are truncation of the oldest material or summarisation of it into something shorter. Both discard information, and neither has any way of knowing which detail will matter later.

Even inside the window, attention over a long input is not uniform. Material at the beginning and end tends to exert more influence than material in the middle, which is a robust and much-studied effect and not a defect of any particular system. Something stated once, in passing, forty exchanges ago is competing with everything since.

So a request to remember something is not stored as an instruction with special status. It is text among other text, weighted statistically, and its influence decays as the conversation grows.

Persistent memory features are databases with a retrieval step

Systems that appear to remember you across sessions typically write facts to a store and retrieve relevant ones later, inserting them into the input before the model runs. This is genuinely useful and it is architecturally ordinary — the same retrieval machinery used for documents, pointed at notes about a user.

Its failure modes are those of retrieval rather than those of memory. A fact that was never written cannot be recalled. A fact written imprecisely is recalled imprecisely. And relevance is judged by similarity, so material that matters for a non-obvious reason may not surface when it should.

There is also an accumulation problem that nobody has solved elegantly. Human memory forgets, revises and reprioritises constantly. A store of extracted facts does not, so stale preferences and superseded details persist and continue to be retrieved unless something explicitly removes them.

Why it cannot simply be fixed by updating the model

The obvious remedy — adjust the parameters as new information arrives — runs into a well-documented obstacle. Training a network on new material tends to degrade what it previously did well, because the same parameters encode both, and the update has no way to protect the old behaviour while changing the new.

A related risk is that continuous updating from ordinary use would let whatever people said flow directly into the model’s behaviour, with no filtering stage and no way to reverse a bad change. The history of deployed systems that learned from live interaction is not encouraging on this point.

Research into updating models incrementally without this damage is genuinely active and has produced partial results. Anyone claiming the problem is solved is overstating it, and anyone claiming it is fundamentally unsolvable is also going beyond the evidence.

What the absence of memory means in practice

It means the system cannot learn your situation over time in the way a colleague does. It can be told, repeatedly, and it can be given a store of facts to consult, and the results can be quite good. But the improvement lives in the plumbing around the model rather than in the model itself.

It also means claims about a system that knows you should be read carefully. The knowledge is a file, held by whoever operates the service, retrievable and deletable, with all the privacy consequences a file has. That is a different arrangement from a person remembering you, and it deserves a different word.

The honest framing is that these systems are stateless engines with storage attached. Everything that looks like continuity is the storage doing its job, and understanding that predicts the failures rather than being surprised by them.

Common questions

Does the model learn from my conversation for other users?

Not directly. Conversations do not alter the running model. Providers may retain data and use it in a later training round depending on their terms, which is a separate process with its own settings, and it is worth reading those terms rather than assuming either answer.

Why does it forget an instruction I gave earlier?

Because the instruction is ordinary text in a long input rather than a stored rule, and its influence competes with everything else present. As the conversation lengthens, older material may be summarised, truncated or simply outweighed. Restating a constraint is not a workaround so much as an accurate response to how the system works.

Is a longer context window the same as better memory?

It helps and it is not the same thing. A larger window lets more material be re-read each turn, but attention across a long input remains uneven and cost rises with length. More room to put things is not the same as reliable recall of what is in the room.

Limits & Risksmemorystatecontextarchitecture
Daniel Okonkwo
Contributing editor, AI Worth Knowing

Daniel covers how it works, in the world, limits & risks and the questions readers actually send in and prefers a plain explanation to a clever one.