Language Models Lack True Memory Across Conversations
Most people watching AI systems fail ask the wrong question. They ask whether the model is smart enough. The more useful question is whether the model is doing what it appears to be doing. Those are not the same thing, and the gap between them is where the trouble lives.
A large language model in a long conversation looks like it is building on what came before. It sees the earlier exchange, it responds to the current prompt, and the surface reads as continuity. But appearance and mechanism have quietly parted ways. What the system retains from an earlier problem does not automatically help it solve the next one. Sometimes it helps. Sometimes it hurts. And the system itself has no reliable way to tell you which is happening.
That is the finding at the center of a new paper on how retained computation from earlier problems affects later ones. The work is not about whether models are powerful. It is about a specific deception baked into how they process a conversation — one that users, and increasingly the institutions deploying these systems, are not equipped to see.
The History You Think You Have
When you ask a model a question after it has already answered three others, you assume the earlier work is available to it. In one sense it is: the previous turns sit in the context, and the model reads them again as it processes your new prompt. In another sense, it is not available at all. The internal representations that produced the earlier answers are gone. What remains is text — a transcript, not a memory.
The paper’s authors ran preliminary experiments to test what that transcript actually does. [1] Their result undercuts the intuition: retained history can raise or lower later-turn accuracy, even within the same domain. Retained history can raise or lower later-turn accuracy, and it does so even when every problem comes from the same domain. The same conversation that sharpens one answer degrades another. There is no clean rule the user can apply from the outside.
This is the first layer of the deception. A model that has “seen” an earlier solution looks like a model that has learned from it. The transcript creates the impression of cumulative reasoning. But the effect of that transcript is not reliably positive, and the system gives no signal about which way it is pulling. The user sees a coherent conversation. The mechanism underneath is doing something else.
To get underneath the surface, the researchers used controlled replay. They isolated the internal state changes that occur for each specific pairing of a problem with a particular history. What they found is stranger than simple interference. Across different histories, those internal changes preserve similar relationships among the current problems. In other words, the history reshapes how the model represents the problem in front of it, and it does so in a structured way — not as noise, but as a systematic shift.
That distinction matters. Random corruption would be easier to dismiss. A structured shift means the model is not merely forgetting or confusing. It is re-encoding the present through the lens of the past, and the lens is not always ground glass.
The Fix That Admits the Problem
Faced with this, the authors did not try to make history uniformly helpful. They built a mechanism that treats the earlier computation as a resource to be redirected rather than a memory to be trusted. They call it STAIR — Stale-Token Attention for Inter-query Reuse.
The design is worth understanding because it is an implicit admission of the gap. STAIR captures keys and values from earlier response generation and stores them in a fixed bank. When the current query reads from that bank during prompt processing, the system has learned to redirect those reads. The base model stays frozen. Only 12,288 parameters are trained — a rounding error against the scale of the model itself.

The framing is telling. “Stale-token attention” names the problem directly: the tokens from earlier turns are stale. They are not a living memory. They are residue, and the question is whether you can teach the model to use that residue without being misled by it. The answer, across three Qwen models and four benchmarks, is that average later-turn accuracy improves by up to 11.67 percentage points over the unmodified model with history. [1] Read that number carefully. It is not a claim that history helps. It is a claim that history, left alone, was costing the model something — and that a small, targeted intervention recovers part of that loss. The improvement is measured against a baseline that already had the history. The history was there. The model was just not using it well.
This is the second layer of the deception. The system appears to incorporate prior context. It does not, at least not in the way the appearance suggests. The fix does not add memory. It adds a learned filter on how the existing pseudo-memory gets read. The model still does not remember. It has simply been taught to glance at the transcript in a less damaging way.
What the Transformer Was Never Built to Do
To understand why this happens, you have to look at the machinery. The transformer architecture, described in the 2017 paper “Attention Is All You Need,” processes text through a mechanism called dense attention. [1] Every token in a block of text is compared with every other token through a form of multiplication. That is how the model encodes meaning — by weighing each word against all the others.
Dense attention is remarkably good at capturing meaning within a block. It is also, by construction, not a memory system. It has no persistent state between calls. When a conversation continues, the earlier turns are re-fed as input, not recalled from storage. The model re-reads the transcript every time. There is no accumulation, only re-processing.
The transformer architecture was not designed as a memory system. It has no persistent state between calls, and earlier turns are re-fed as input rather than recalled from storage. The STAIR paper is a precise instance of the workaround pattern: the flaw is that retained history is not retained in any meaningful sense, and the fix is a small trained module that manages the damage.
The STAIR paper is a precise instance of that pattern. The flaw is that retained history is not retained in any meaningful sense. The workaround is a small trained module that manages the damage. The base model is untouched. The architecture is not fixed. A patch is applied.
This is where the deception becomes structural rather than incidental. A user interacting with a long-context model experiences continuity. The model references earlier points, maintains tone, follows the thread. Everything about the interaction says: this system is tracking our conversation. But the system is not tracking anything. It is re-reading a document it wrote and guessing at coherence. The continuity is an artifact of re-processing, not evidence of memory.
The Cost of Looking Like You Remember
The consequences extend beyond a single conversation. When a model appears to build on earlier work but is actually re-deriving its understanding from a transcript, the errors compound in ways that are hard to attribute. A wrong turn in turn three does not just sit there. It gets re-read in turn four, re-interpreted in turn five, and by turn ten the model is confidently reasoning from a premise it never actually held.
The paper’s controlled replay experiments show that this is not random. Across different histories, the internal state changes preserve similar relationships among current problems. That means the distortion is consistent. The model is not flailing. It is systematically re-encoding present problems through a past that may or may not be relevant. And because the re-encoding is structured, it produces confident, coherent output that is nonetheless built on a shifted foundation.
This is the third layer of the deception, and the one with the longest tail. The system does not just fail to remember. It fails to remember in a way that looks like remembering. The output is fluent. The reasoning is internally consistent. The user has no reason to suspect that the ground has moved. The gap between what the system claims — through its behavior — and what it does is invisible from the outside.

The STAIR results make this concrete. The unmodified model with history was not broken. It was producing answers. Some were better than they would have been without history. Some were worse. The average was dragged down by the worse ones, and no one could tell which was which. The improvement came from teaching the model to read its own stale tokens more carefully — not from giving it better memory, but from reducing the damage of the pseudo-memory it already had.
The Institutional Blind Spot
Here is the part no one names openly. The institutions deploying these systems are building workflows, products, and decisions on top of a capability that does not exist in the form they believe it does.
A customer service agent that “remembers” a prior complaint is not remembering. It is re-reading a transcript and hoping the re-reading produces the right answer. A coding assistant that “builds on” earlier code in a session is not building. It is re-processing the file and generating a continuation that may or may not respect what came before.
None of this is hidden in a malicious sense. The systems do not lie. They simply behave in a way that invites a false interpretation, and the false interpretation is the one that makes them most useful to sell and easiest to deploy. The gap between appearance and mechanism is not a bug the vendors are covering up. It is a property of the architecture that the surrounding narrative has learned to ignore.
The STAIR paper is a small, technical correction to one slice of that gap. It shows that the problem is real, measurable, and partially fixable with a tiny number of trained parameters. It does not show that the problem is solved. The base model still has no memory. The transcript is still stale. The filter is better, but the underlying deception — the appearance of continuity without the substance — remains.
What the work really demonstrates is that the field is now in the business of managing the gap rather than closing it. The transformer was never built to remember. The industry built products that assume it does. The research is now patching the difference, one mechanism at a time, while the products continue to present continuity as a feature.
The question worth asking is not whether the next model will remember better. It is whether the institutions relying on these systems understand that the memory was never there.
