Context window amnesia has a signature moment. Forty turns into a build, the assistant proposes work that contradicts the objective you set on day one and never changed. You feel ignored. The mechanism that produces this lives in the pillar piece on agent memory architecture, and this post deliberately stays with the symptom, because having the name changes what you ask for.
Here is the correction that makes the term make sense. A chat model has no memory that could fail in the first place. Nothing was learned inside the model, so nothing inside the model can be forgotten. Amnesia is not forgetting; it is no longer being shown.
Every turn, the software around the model assembles a transcript and ships a fresh copy of it. The model reads the copy, writes the next response, and keeps nothing. The next turn gets a new assembly, and parts of what was shown before are quietly missing from it. When the assistant forgets, the failure is upstream of the model, in the assembly.
The Problem: A Conversation That Stops Being Shown
The symptom has a shape. A session starts sharp. The assistant knows your constraints, your naming habits, the thing you are building and the definition of done you argued about for an hour. Deep into the session it starts proposing work that contradicts an early decision, asking for an input it was already given, or drifting toward a version of the goal nobody agreed to. People search for why does my AI forget, and the answers on offer are usually vibes: it gets confused, the context gets long, try restating things clearly.
Nobody is confused. The transcript is being trimmed, and the trim falls where the cheapest cut falls.
- The model is re-shown its context every turn and stores nothing between calls.
- When the transcript outgrows the window, the oldest parts stop being shown.
- The oldest parts are where the objective and the constraints live.
- A bigger window delays the cut. It does not end it, and nothing crosses the gap between sessions.
- A fix is a system that selects, retains, compresses, and reconstructs context. That is what the pillar covers.
1. Nothing Was Ever Learned, So Nothing Is Being Forgotten
A model at inference time is a function: it maps what it is shown in this moment to what it says next. What it appears to know about you is either baked into its training weights or sitting on the page in front of it. Between the two, only the page changes, and the page is written by whoever assembles the prompt.
So the assistant has exactly one memory, and it is the window.
This reframing matters because the everyday language does not. We say the agent forgot, as though there were a store inside that dropped a record. There is no store. There is a page, and the page got shorter, or the part that mattered moved somewhere the model does not look very hard.
When your assistant forgets, check the page, not the brain. That single sentence separates every usable memory product from the ones that only look clever in demos.
2. What Gets Dropped First, And Why That Order Is The Worst
The transcript grows every turn and the window does not. Something must stop being shown. Almost every assistant cuts from the oldest end, because the current exchange feels like the only thing that matters right now, and a cut at the front is the cheapest code to write.
The order is the worst possible one because of what lives at the front. A session’s oldest content is where the objective statement, the constraints, the rejected alternatives, and the definition of done were written down. The newest content is the current state, which is what the trim is trying to protect. Oldest-first therefore eats the frame and keeps the motion. The assistant gets progressively better at local work and worse at knowing what the work was for.
A study by Noah Liu and colleagues at Stanford measured what happens even before anything is cut. Answer accuracy across a long context follows a U-shape: high when the relevant fact sits at the beginning or the end, degraded in the middle. A constraint in the middle of a long prompt is present and still underweighted. Position is a tax you pay even when nothing is dropped.
This is not a hypothetical about other products. When the Ocai transcript sheds old messages, it leaves a line where they were: Older messages hidden to keep context bounded. You can watch the amnesia happen and read its own description of what it did, which is more honesty than most assistants offer. How the cuts are decided, and what a real compactor preserves instead of dropping, is the pillar’s territory, and it is covered there.
3. Where Amnesia Bites: Three Places
The symptom is not one failure. It is three, at three different scopes, and users solve each one differently and badly.
| Scope | What goes missing | What it feels like |
|---|---|---|
| Mid-session | The early objective, constraints, decisions | Repeating yourself inside one conversation |
| Between sessions | The entire previous conversation | Explaining the project from scratch every morning |
| Between projects | Conventions learned in another workspace | Every new repo feels like the first one |
Mid-session is the one people complain about, because it happens while you are watching. Between sessions is the one that quietly costs the most: the cold start every Monday morning where you re-brief the tool on a project it spent all of last week inside. Between projects is the one that makes power users cynical, because the assistant that learned your conventions in one workspace arrives at the next one as a stranger.
Note that only the first of the three is a window problem. The other two are continuity problems, which is why window size and memory are constantly confused in product conversations and why the two claims lead to different designs.
4. The Tax You Pay For It
Users do not sit idle while their assistant forgets. They build a workaround culture, and the workarounds are all labor.
One: pasting summaries of your own summaries into new chats, so the assistant starts at page twenty instead of page one. Two: keeping a notes file that describes the work the assistant is supposed to help with, which is ironic the second you say it out loud. Three: restating preferences every session. Use four-space indents. We chose Postgres. Do not invent numbers. You have said these things forty times to forty separate chats, and each chat received them like new information.
Add it up on a normal week and the re-briefing is a real line item: an hour here, two there, plus the slow tax of work thrown away because the assistant optimized for a goal you corrected on day one.
The inversion is what deserves attention. A tool sold as a thinking partner quietly turns its user into its context manager. The human holds the state; the machine performs on it. That is the opposite of the bargain, and it is the actual product experience of context window amnesia.
5. Why A Bigger Window Does Not Fix It
Start with the assumption most builders make: I have a million-token window, so I do not need a memory system. The data says otherwise, for at least four reasons.
First, a bigger window still has a middle. The U-shape is not a size problem; long contexts simply give content more middle to disappear into. Second, re-shown is not used. Every token in a fat transcript is read at a cost, and cost and latency per turn scale with window size, which is why vendors put a working ceiling below the advertised one anyway. Ocai’s reference implementation sets its working ceiling at 180,000 tokens even where the preset model window goes to 204,800, and every product you use picks some number like this. Your transcript does not care about the number. It keeps growing.
Third, windows bound every session and nothing bounds the work. A large research chat, a multi-week build, a migration: all of them eventually outgrow any finite ceiling, so the question is not whether content stops being shown but whether anything survives the cutoff. A bigger window changes the day the bill comes due. Fourth, none of it crosses the gap between sessions. The window lives inside one call. Closing a tab and reopening is zero shown, at any window size.
6. How The Symptom Shows Up In The Assistant You Use Today
You may be reading this inside a competing assistant, and the diagnosis holds there too, because the mechanism is universal. Watch for these in your daily tool. They are behaviors, not accusations, and every product ships some version of them.
The chat that goes bad. A long session drifts. The assistant stops honoring constraints it saw early, you start a fresh chat, and the problem vanishes. That was oldest-first trimming plus positional fade, and the fresh chat was a cold start, not a fix.
The memory feature that stores the wrong layer. It remembers that you prefer concise answers, and it forgets the working state of the task you did last Tuesday: which files moved, which decision was reversed, which test still fails. Facts are easy to store; working state is the hard part, and most memory features do not touch it.
The declared memory that goes stale. Custom instructions and rules files are memory you write once. They never learn from what happened in later sessions, so they drift out of date and start fighting your current intent.
The running-brief habit. You keep a summary of the project and paste it into each new chat. That is you doing reconstruction by hand, the last job of the fix below, unpaid.
None of these are user errors. They are rational coping with a product that has amnesia. And notice what they have in common: every one of them makes you the memory subsystem from the section above, just with better tooling.
7. What An Actual Fix Requires
A fix is a system that does four jobs the model cannot do for itself. It selects, which means deciding on the way in what deserves to persist. It retains, storing those decisions and facts somewhere outside the window. It compresses, shrinking long exchanges without destroying what made them matter. And it reconstructs, assembling a fresh context every turn from what survived, so the model is shown the frame and not just the last motion.
That is architecture, not prompt size. The layer taxonomy, the write path, and the compaction mechanics are the pillar’s subject, and reading this section of the cluster without the other is like taking a tour of the engine through the windshield.
You can evaluate any assistant against those four jobs in a week. Ask it three questions. Can it show you what it remembers about you, right now? Does it get sharper over weeks on the same project, or does it stay flat? Do your decisions survive a restart? A product whose intelligence is flat across weeks has no accumulation, and no accumulation is the clinical sign of amnesia wearing a nice interface.
Conclusion: Amnesia Is A Systems Failure, Not A Model Failure
Take a position: context window amnesia is a failure of the systems shipping the models, not a limitation of the models themselves. The model is the reader, never the memory. Everything worth doing about continuity is a decision somebody made about the page, and a vendor who has not made those decisions is selling you a reader and calling it a partner. Until an assistant can show you its memory, you are the memory subsystem, unpaid and unthanked. That is the honest state of the market. It is also why the unglamorous work of selection and retention will beat the next doubling of the window, because doubling the desk does nothing for the librarian.
FAQ
Does a longer context window reduce context window amnesia?
It delays the first cut of content. Positional decay, per-turn cost, and the gap between sessions do not scale down with window size. A longer window is a bigger desk, not a better memory.
Why does the assistant drop what I said early instead of recently?
Because oldest-first is the cheapest trimming policy to implement, and the current exchange feels like the only thing that is needed immediately. The objective happens to live exactly where the cheap cut lands.
Is amnesia the same thing as hallucination?
They are different failures that often pair up. Hallucination is producing content that was never there. Amnesia is failing to be shown content that was. And amnesia causes hallucination: with the objective gone, the model fills the gap with something plausible.
Does starting a new chat help?
It resets the symptom, not the cause. You get a clean window and a cold start. Everything that lived only in the transcript is gone with it.
My assistant has a memory feature. Why does it still lose the plot?
Because a memory store fixes continuity, not the mid-session problem, and some features store only facts and not the working state. The window and the store are two different failures wearing one symptom.
See the assembly decide what survives: Start a Session Now
- Noah Y. Liu et al., Lost in the Middle: How Language Models Use Long Contexts, Transactions of the Association for Computational Linguistics, 2024. https://arxiv.org/abs/2307.03172
- The cluster pillar: Agent memory architecture: how the agent remembers, compacts and forgets, the Ocai blog. /blog/agent-memory-architecture
- Ocai, Context compaction reference, the agent-framework repository, 2026. https://example.com