The capability
What a memory is
There are mainly two types of memory when it comes to LLMs. The first is short-term memory, which is the list of messages the model gets on every call during the conversation you're in right now. Everything said so far in that thread is in the list, and when the thread ends the list goes with it. M5 was all about managing that list.
The second is long-term memory, which survives outside the current conversation. If you've used ChatGPT, you've probably already encountered that kind of memory. You can open a new chat, ask about something you talked through weeks ago in a different chat, and it still knows. Or you tell it once that you write in British English and don't like bullet points, and every new chat follows that without you saying it again. That second kind is what this module is about.
According to OpenAI, ChatGPT does this in two ways. Saved memories are things you ask it to remember, like "remember I'm vegetarian", and they get written down during the chat. Since April 2025 it also learns from your chat history, through a background process OpenAI calls dreaming, which reads across your past conversations and keeps rewriting what ChatGPT knows about you so new chats start with that context (OpenAI, June 2026). OpenAI doesn't publish how ChatGPT finds the relevant parts of old chats, so whether it also embeds them and searches them the way RAG does (M19 covers that) isn't public. What OpenAI does describe is the background job that keeps ChatGPT's notes about you up to date.
The same post explains why they rebuilt it. OpenAI says saved memories "tend to go stale over time and eventually become incorrect or irrelevant". Elsewhere in the post it lists staying current as one of the goals, with the example of a memory that says you're planning a birthday party for next Saturday, which stops being true once Sunday arrives. Keeping memory current when the facts behind it change is most of this module.
When you build your own agent, none of that comes for free. The model keeps nothing between calls, which was M3's point, so something has to decide what's worth keeping and write it down, and something has to load it back later.
The writing happens in one of two ways, and they're the same two ChatGPT uses. The first is during the conversation. You give the model a tool called something like remember_fact. The model decides by itself when something the user said is worth keeping, and calls it. The tool call is only a request, so your code is what checks the value and writes the row, the same as with any other tool the agent has. That's how ChatGPT's saved memories work, and OpenAI's post says they "were only written during the conversation and relied on strong cues to decide when to trigger memory", like someone saying "remember I'm traveling to Singapore in July".
The second is after the conversation. Once a session ends, a separate model call reads the whole transcript and pulls out anything worth keeping, whether or not anyone asked for it to be saved. That's closer to what ChatGPT's dreaming does, and it catches the things the model didn't think to save while it was busy answering.
Loading back is simpler. At the start of the next conversation your code looks up what's stored for that user or project and puts it in the list the model sees. A fact that gets saved and never loaded doesn't do anything.
So you have a write policy, which is whatever decides what gets saved, in either of those two ways, and a read policy, which decides which saved facts get loaded on a later turn and where they go in the context. They're separate pieces of code and they break separately. M6 built a table for exactly this, keyed by customer id. The same shape works keyed by a user or a project. What belongs in that table was left to this module.
The running example for this module is a coding agent. A small team uses it on their web app's repo, and it remembers facts about the project between sessions, like which command runs the tests and where staging deploys go. The team and its repo are made up, and so are the numbers. In the figure, a developer tells the agent in February where staging deploys go, and the agent reads that fact back in September.

The model itself still doesn't remember anything. It only reads what's in the list on this call, so when an agent seems to remember you, that's your code loading the right fact. When it forgets you, the fact was never saved, or the query didn't fetch it, or it was fetched and never put in the list. M5 didn't cover a fourth way it goes wrong, which is that the fact made it into the list and was wrong, and that's the one this module spends most of its time on.
Why this is a capability of its own
After M5 and M6 you might think there isn't much left to do. The store lives outside the model and reading from it is a lookup by id. Writing a row and reading it back is a morning's work.
It stays that simple while you're testing it. Once real people use the agent for months, memory gets complicated quickly, and if you don't design it properly it starts working against you. Something has to decide what's worth saving and when, what to do when a new fact contradicts an old one, and which of everything saved gets loaded into this conversation. Get those wrong and the agent confidently acts on things that are no longer true, or fills its context with notes that have nothing to do with the task.
Facts going out of date are the clearest example. A developer tells the agent one thing in February and something different in June, or the agent writes down that a form field is required and a month later someone changes the form. Every write was right when it happened, and the rows stop being true anyway. Nothing in the store notices, because an old row looks exactly like a current one.
Where it shows up
Long-term memory shows up anywhere an agent deals with the same person or the same system more than once.
- A coding agent's progress file, read at the start of a session so the next one carries on from the last one.
- A tutoring app's note of which topics a student has already covered and which ones they struggled with.
- An account record a sales assistant carries from one call into the next.
- An agent's own notes about a tool, like which field on an internal admin form fails silently when it's left blank.
All of them have the same write and read code behind them. What's different is how much it costs when the agent acts on something that's gone out of date.
Where this starts
The deploy target that was true in February
The team's coding agent has a remember_fact tool with a short list of fact types it's allowed to save, like test_command, package_manager and deploy_target. The model calls it when a developer tells it something about the project that should stick, your code checks the value, and the developer confirms before the row is written. Facts are stored per repo, so everyone on the team gets the same ones.
In February a developer sets one up. She tells the agent that staging deploys go to the uat branch, and that when she says "ship this to staging" it should just push there. The model reads that as a deploy_target and calls remember_fact. She confirms, and a row goes into the project's facts. Everything about that exchange is correct.
In June the team repurposes uat. The sales team needs a stable environment for customer demos, so uat becomes the demo box and staging moves to a new preview branch. The same developer is in a session asking the agent to fix a flaky checkout test, and she mentions it on the way, writing "also, uat is the sales demo box now, staging moved to preview". The agent fixes the test. The model never calls remember_fact, though, because that tool is for facts a developer asks it to save and confirms, and she did neither. She was asking about a test and happened to mention the thing that had changed, so nothing in the turn looked like a request to remember something.
In September she's finishing a checkout redesign and says "ship this to staging". The agent loads the repo's facts and finds the February deploy_target. It pushes to uat. That afternoon the sales team demos the product to a prospect on uat, and the checkout page is half-built.

Pull that apart and there are two separate failures. The June turn should have produced a write and didn't, because the agent's only way to write was the model calling a tool during the conversation, and nothing in that turn prompted it to. A pass over the transcript after the session would have seen the sentence about uat. The September turn should have found a reason to distrust a seven-month-old deploy target and had none available, because the row carried a value and a timestamp and nothing about how long that kind of fact stays true.
Neither failure raises an error anywhere. Every component behaved as written. The tool was called correctly in February, its conditions weren't met in June so it didn't run, and the store was queried correctly in September. M2's traces show a clean session end to end. The agent was confidently wrong about the team's own setup on the strength of a row it had every reason to trust.
What to keep
What deserves to become a memory
Before deciding where memories live or how they get written, it helps to be clear about what should become a memory at all, because most of what people say to an agent shouldn't. Things a developer says need different handling depending on what kind of statement they are.
A standing instruction or preference is something said as an ongoing rule, like "we use pnpm in this repo" or "Sam reviews anything in payments". This is what memory is for. It was said directly, it's meant to apply next time too, and saving it saves the developer from repeating it every session.
A one-off request applies to this time only, like "just run the checkout tests, I'm in a hurry" or "push this hotfix straight to hotfix-42". It often reads exactly like a standing instruction, and saving it as one is how a single shortcut quietly becomes the default. Words like "this time" or "for now" give it away, and when they aren't there, the safer assumption is that a request is one-off until it's repeated or confirmed.
An observed event is something that happened, like "the uat server was down this morning" or "the migration failed at 3pm". It can be worth keeping as a record of the past, which is the episodic kind of memory the next section covers, but it isn't a rule. "The uat server was down" tells you nothing about where staging deploys.
An inference is something the agent worked out that nobody said. If a developer mentions she's been reviewing a lot of Sam's PRs, it's tempting to conclude that Sam is the reviewer for her area, and nothing she said supports that. Inferences can be right and they can be useful, but they're the agent's guess, and a guess stored next to stated facts looks exactly like one of them.
One more trap cuts across all of these, which is who or what the statement is about. "I'm also helping Priya on her mobile-app repo, which uses yarn" is a stated fact, and it's about a different repo. Saving it as "this repo uses yarn" is the same mistake as turning "I'm buying a gift for someone who likes roses" into "this customer likes roses". The statement was accurate, and the memory is wrong because it's filed under the wrong owner.
So a candidate memory has to carry more than a value. It needs to say what kind of statement it came from, who or what it's about, how long it's meant to apply, and the exact words it came from, so the pipeline can decide what to do with it.
| Said in a session | Kind | What happens to it |
|---|---|---|
| "We use pnpm in this repo." | Standing, stated | Saved as a fact for this repo |
| "Just run the checkout tests this time." | One-off | Used now, nothing saved |
| "Until Friday, don't touch the billing folder." | Standing, with an end date | Saved with its end date |
| "The uat server was down this morning." | Observed event | Kept in the session notes, never as a fact |
| "I've been reviewing a lot of Sam's PRs." | Invites an inference | Nothing saved, or the agent asks if it's relevant |
| "Priya's mobile-app repo uses yarn." | About another repo | Never filed under this repo |
Section 13 runs every one of these kinds through a small pipeline, with and without these checks, so you can see how often skipping them goes wrong.
The options
Where long-term memory can live
Before deciding what the agent should save, it helps to know where it could go. The options run from a text file to a graph database, and each one makes different things easy, so the choice shapes everything that comes after it.
The simplest option is to have no long-term memory at all. If every conversation stands on its own, like a tool that reviews one pull request and then it's done, the message list from M5 is all you need. Adding a memory layer to something like that only adds things that can go wrong, so it's worth asking whether the agent will ever see the same person or project twice before you build anything.
A notes file loaded every time
The next step up is a plain text file that gets loaded into the context at the start of every session. Coding agents use this a lot. Claude Code reads a CLAUDE.md file from the repo, and many other coding tools read an AGENTS.md file, which is the same idea under a shared name that its site says over 60,000 open-source projects use. The file holds things like how to run the tests, which folders not to touch, and how the team likes code written, and the agent gets all of it every time.
Claude Code also keeps notes of its own that it writes itself, which it calls auto memory. That one is a folder with an index file called MEMORY.md and one file per topic. As of September 2026 it loads the first 200 lines or 25KB of the index at the start of every conversation, and it only opens a topic file when it needs what's in it (Anthropic, Claude Code memory docs). Anthropic's API memory tool works the same way from the model's side. The model asks to read or change files in a memory folder, and your code carries out each request against storage you control.
A notes file is easy to read and easy to fix by hand, and if it lives in git you can see who changed what and when. What it can't do is stay small or tell you when a line has gone out of date. Everything in it loads every time, and the same docs say to keep a CLAUDE.md under 200 lines because "Longer files consume more context and reduce adherence." And if a line goes out of date, like one saying staging deploys to uat, nothing in the file tells you.
A table of facts
The third option is the one M6 built, a database table where each row is one fact, with columns like type, value, written_at and repo_id. The coding agent's project_facts table is this. Your code loads rows by key, so there's no searching involved, and you can change one fact without touching the others or expire a fact once it's old.
The cost is that somebody has to decide the fact types up front. deploy_target and test_command have a home, and a loose observation like "the checkout tests are flaky on Mondays when the fixtures reset" doesn't fit any type, so it either gets forced into one or it isn't saved.
A store you search
The fourth option is to save notes or chunks of past conversations along with an embedding of each one, and at the start of a turn search for the ones closest in meaning to what's being asked. That's retrieval, which M19 covers in full. This can hold far more than you could ever load, and it's the natural place for the loose things a table has no room for, like what happened the last time someone changed the payment code.
The downside is that search isn't exact. It can miss the note you needed or bring back one that only sounds related, and updating or deleting one specific memory is awkward because you have to find it first.
A graph with dates on it
The last option is a knowledge graph, where the things the agent knows about, like people and services, are nodes, and the relationships between them are edges. Zep is a published example built for agent memory. Each edge carries the time range when it was true, and when a new fact contradicts an old one, an LLM compares them and Zep "invalidates the affected edges" by setting an end date on the old one. The old fact stays in the graph with its dates, so the agent can answer both "where does staging deploy now" and "where did it deploy in March" (Rasmussen et al., 2025).
That handles facts that change and facts that connect to each other better than any of the other options. It also costs the most, since every write needs model calls to pull out entities and check them against what's already there, and you're running a graph database on top of everything else.
| Option | How the agent reads it | Good for | Where it goes wrong |
|---|---|---|---|
| No long-term memory | Only the current message list | Tools where each conversation stands alone | Anything that needs to carry over |
| Notes file | Whole file loaded every session | Instructions a person writes and reviews | Grows, and can't tell you a line is stale |
| Table of facts | Rows loaded by user or project id | Specific facts you update and expire | Things that don't fit a fact type |
| Searchable store | Search by meaning each turn | Lots of loose notes and past sessions | Misses, near-misses, hard to edit one item |
| Graph with dates | Query entities and their current edges | Facts that change and connect | Model calls on every write, more to run |
Most agents use more than one
The systems in this section each combine more than one of these, and the reason is clearer once you split memory by what kind of thing is being remembered. Researchers who study agents borrow three terms from psychology for this. Procedural memory is how to do things, like "run pnpm test before committing". Semantic memory is facts, like which branch staging deploys to. Episodic memory is past events, like what happened in last Tuesday's session (Sumers et al., 2023).
Each kind has a natural home. Procedural memory belongs in the notes file, since it's instructions a person should be able to read and approve, and that's what a CLAUDE.md is. Semantic memory belongs in the table, where each fact can be updated and dated on its own. Episodic memory belongs in the searchable store, because there's a lot of it and you only ever want the few pieces that bear on the current task.
The other thing most designs share is a split between a small part that's always loaded and a big part that gets fetched when needed. MemGPT, an early design from 2023, gives the model a fixed-size block of notes about the user that sits in every prompt, and the model edits that block by calling functions. Everything else goes to storage outside the context. Old messages that get pushed out of the window are "stored indefinitely in recall storage and readable via MemGPT function calls", and there's a separate archive for longer text the model chooses to keep (Packer et al., 2023). Claude Code's always-loaded index with topic files on the side is the same shape, and so is ChatGPT's summary of you with your chat history behind it.
The coding agent in this module ends up using three of the options. The team's CLAUDE.md holds how they work, project_facts holds specific facts like the deploy target, and past session logs are searchable for the times the agent needs to know what happened before.
Procedures the agent learns for itself
Procedural memory doesn't have to be written by a person. Over many sessions an agent keeps running into the same things, like the checkout test that only passes after the fixtures are reset, or the deploy script that fails unless the VPN is on. A note describing how to handle one of those is a procedure the agent learned from experience. LangChain's docs say it's "fairly uncommon for agents to modify their model weights or rewrite their code" and "more common for agents to modify their own prompts", and that's what a learned procedure usually is in practice, a line added to the instructions the agent loads every session. Claude Code's auto memory is built for this, and its docs describe it as letting "Claude learn from your corrections without manual effort."
Because a learned procedure changes how the agent behaves on every future task, it deserves more care than a single fact. The coding agent's after-session pass can propose one when the same fix shows up in two or more sessions, and it goes into a pending section of the notes file that a developer reviews before it becomes an instruction. It loads into every session once it's accepted, so a wrong one affects every task the agent does from then on, where a wrong fact only affects the tasks that load it.
The write side
Deciding what gets saved
Wherever the memory lives, something has to decide what goes into it. Section 1 described the two moments that decision can happen, during the conversation when the model calls a tool, or afterwards when a separate pass reads the transcript. They trade off against each other in ways worth knowing before you pick.
Saving during the conversation means a new memory is there straight away, and the user can be told it happened. LangChain's memory docs describe this as making memories "immediately available for use in subsequent interactions". The same docs point out the cost, which is that "the process of reasoning about what to save to memory can impact agent latency", and that the model is doing two jobs at once, answering the user and deciding what to remember, so it does each a little worse (LangChain, memory concepts). That's the June failure from section 2. The model was busy fixing a test and didn't treat a passing remark as something to save.
Saving afterwards takes that pressure off. A separate model call reads the whole transcript with nothing else to do, adds no latency to the conversation, and can catch things nobody asked it to remember. The catch is timing. The docs note that "infrequent updates may leave other threads without new context", so if the pass runs once a night and the developer opens a second session an hour later, the new fact isn't there yet. A sensible default is to run it when a session ends, with a nightly job as a backstop for sessions that never close cleanly.
Asking the user
The most reliable filter is to ask. The coding agent does this already, repeating the fact back and waiting for the developer to confirm before anything is written. Asking catches the extractor's mistakes before they reach the store, and it's the most dependable way to tell a standing preference from a one-off, since "just run the checkout tests" could mean either.
The cost is that people stop reading prompts they see too often. If the agent asks after every session, the developer starts clicking yes without reading, and the confirmation stops meaning anything. So save the question for facts the agent will act on without checking again, like where to deploy or which files not to touch, and let low-stakes things like how she likes commit messages worded go in without asking.
Scoring what's worth saving at all
Not everything that comes up in a session deserves a row. The Generative Agents paper, from researchers at Stanford and Google in 2023, asked the model to rate each new memory for importance on a scale of 1 to 10, with a prompt that anchors 1 at "purely mundane (e.g., brushing teeth, making bed)" and 10 at "extremely poignant (e.g., a break up, college acceptance)". Their example returned 2 for "cleaning up the room" and 8 for "asking your crush out on a date" (Park et al., 2023). They used that score to rank memories later, and you can use the same idea as a gate. For the coding agent, "the deploy target changed" is a 9 and "she renamed a local variable" is a 1, and a threshold between them keeps the store from filling up with noise.
The score comes from a model, so it has an error rate like any M13 classifier and it needs checking against a person's judgement on a sample. It also sees one memory at a time, so it can't tell that the fifth note about flaky tests this week counts for more than the first.
What a saved memory holds
A candidate that gets through needs to be stored with enough around it that anyone, including the agent months later, can tell where it came from and whether it still applies. The coding agent's rows look like this.
| Field | Example | Why it's there |
|---|---|---|
repo_id | web-app | Who the fact belongs to, taken from the session and never from the model |
type | deploy_target | One of a short allowed list |
value | preview | The fact itself, checked in code |
kind | standing | Standing, or standing with an end date |
evidence | "staging moved to preview" | The developer's exact words |
source_turn | s_14027 / 1 | Where to find the full conversation |
confirmed | false | Whether the developer confirmed it |
valid_until | null | The end date, for temporary rules |
status | active | Active, superseded or deleted |
written_at | 2026-06-30 | When it was saved |
The evidence field is cheap and it catches a lot. The validator checks that the quoted words appear in what the developer said, which catches the extractor inventing a fact, and it also catches the extractor quoting the agent's own reply back as if the developer had said it. If the quote isn't in the developer's message, the candidate is rejected.
It also makes the default outcome for a session obvious, which is to save nothing. Most sessions are about renaming a function or fixing some CSS, and a pipeline that finds three memories in every session is saving noise. Most sessions should end with no new rows at all.
How big a piece to save
A common design, and the one this team built, pulls facts out of the conversation and stores them as short statements, because short statements are cheap to load and easy to read. That design has a measured cost, and knowing what the cost is changes where you draw the line.
LongMemEval is a benchmark of 500 hand-written questions about a user's own chat history with an assistant, published at ICLR 2025. Its authors tested storage granularity directly, comparing whole sessions against individual rounds of the conversation and against compressed user facts. Storing a round beat storing a session when GPT-4o was doing the reading, and made little difference with Llama 3.1 8B Instruct. Compressing further into individual user facts "harms overall performance due to information loss", while improving accuracy on the questions that need several sessions combined (Wu et al., ICLR 2025). Those experiments used GPT-4o and two sizes of Llama 3.1 Instruct, which were the models of 2024 and 2025, so treat the ordering as the finding and the margins as dated.
Compression hurts for a reason that applies directly to a store you build yourself. A fact like deploy_target: push to uat is a summary of a longer exchange in which the developer explained that uat auto-deploys to the staging server and that nobody else uses it, which is why pushing half-finished work there was fine. The summary is what a later session can act on. The explanation is what a later session would need to work out whether the instruction still applies once uat has a different job. Extraction throws away the second thing to make the first thing small, and you find out which one you needed months later.
So the practical shape of a write policy is two records and not one. The extracted fact is what the read policy loads by default, and a pointer back to the turn it came from is what an agent follows when the fact looks doubtful. Storing the pointer costs a foreign key. Storing the turn costs whatever your conversation table already costs, since M6 stored the transcript anyway.
The two ways a write policy misses
The same LongMemEval paper ran a study on shipping consumer assistants that names both failure directions in one experiment. Human annotators worked through 97 questions with 3 to 6 session histories, talking to the products turn by turn through their web interfaces, all of it during the first two weeks of August 2024. Reading the same histories straight into GPT-4o's context scored 0.9184. ChatGPT scored 0.5773 on the same questions with the same underlying model, and Coze scored 0.3299.
The authors describe what went wrong in one sentence, and the two halves of it are the two write-policy failures anyone building this will meet. ChatGPT "tended to overwrite crucial information as the chat continues", and Coze "often failed to record indirectly provided user information".
| Failure | What the policy did | What the user sees | Where it shows up for the coding agent |
|---|---|---|---|
| Overwrote | Replaced a record that was still current | An answer built on something the agent decided later and got wrong | The standard test command replaced by a one-off pytest -k checkout |
| Never recorded | Took no action on a fact stated in passing | The agent asks again, or acts on the old value | The June sentence about uat |
Those numbers are two years old and both products have changed since, so don't read much into the scores. The two failures are still the ones you'll hit. A policy that writes eagerly overwrites things that were still true, and a policy that writes only on an explicit confirmed request misses everything a user mentions while talking about something else. This team sits at the conservative end on purpose, since a memory the model can write on its own is also a way for text it reads, like a README or an issue, to plant instructions in it, which M10 covers. The agent only writes when the model calls remember_fact during the conversation and the developer confirms, and the June turn is what that decision cost.
The fix is to use both ways of writing. The model keeps calling remember_fact when a developer asks for something to be saved, and an after-session pass reads the transcript for everything else.
That pass needs two different standards inside it. Deciding that a sentence contains a candidate fact is cheap and reversible, so it can afford to be generous. Deciding that a candidate fact should replace a stored one is expensive and hard to undo, so it has to be strict. A policy that treats those as one decision has to pick a single level of caution for both, and whichever it picks, one of the two failures above is what it buys.
def on_session_end(session, repo_id):
"""Two decisions, run at different thresholds."""
# Generous: anything that looks like a durable fact becomes a candidate.
candidates = propose_facts(session) # a model call over the transcript
for c in candidates:
if c.type not in ALLOWED_TYPES or not value_is_valid(c):
continue # allowed types and value checks
current = store.current_fact(repo_id, c.type)
if current is None:
store.write(repo_id, c, source_turn=c.turn, state="active")
elif c.value != current.value:
# Strict. A replacement is a decision, so it waits for review.
store.queue_supersession(current, c, source_turn=c.turn)The propose_facts call is a model reading a transcript and returning candidate records, which is the M13 classifier shape with the label space being your fact types. It has a confusion matrix and thresholds, and you measure it the way M13 measured the ticket sort. The difference is that the two error directions land in different places, and section 10 prices them.
Add a candidate-fact extractor to our coding agent. After a session ends, send the transcript to the model with our six allowed fact types and ask it to return zero or more candidates, each with a type, a value, the turn index it came from, and a confidence between 0 and 1. Validate every value in Python against the same rules the remember_fact tool uses. When no active fact of that type exists, write it. When one exists and the value differs, write a row to fact_supersessions holding the old id, the new value, the source turn and the confidence, and leave the active fact alone. Log every candidate, including rejected ones, so we can score the extractor later.
The balance
Too little memory, and too much
Once the agent is saving things, the next problem is how much of it to keep and load. Get it wrong in one direction and the agent forgets things it needed. Get it wrong in the other and it drowns in notes that have nothing to do with the task. Both have been measured.
Too little is what section 7 is about in detail, but the short version is that squeezing memory down loses facts. LongMemEval found that compressing sessions into individual facts lost information, and a June 2026 paper called Supersede, which section 7 goes through, found that a short memory the agent kept itself kept dropping the latest value of things that changed.
Too much is just as bad, and it's the more tempting mistake, because loading everything feels safe. LongMemEval tested it by giving models a user's whole chat history, about 50 sessions and 115,000 tokens, and asking questions whose answers were somewhere in it. GPT-4o got 87.0% right when it was given only the sessions that held the answer and 60.6% when it was given the whole history, and the open models they tried lost between 36% and 55% of their accuracy the same way (Wu et al., ICLR 2025). Those are 2024 models and newer ones handle long contexts better, but the point M5 made about long contexts still holds. Every extra note is something the model has to read past, and it costs money on every call. The Mem0 team reports that their memory layer cut token costs by more than 90% compared with sending the full history, on the benchmark they tested, which is their own number on their own evaluation, so read it as a direction and not a guarantee (Chhikara et al., 2025).
So the job is to keep the store full enough that nothing you need gets lost, and to keep what reaches the model small enough that the useful part isn't buried. Several mechanisms help with that, and a good memory layer uses most of them.
Rank what you load
When there's more in the store than you can load, you need a way to pick. Generative Agents picked with three scores added together. Recency went down exponentially with every hour of the simulation since a memory was last retrieved, with a decay factor of 0.995 per hour. Importance was the 1 to 10 rating from when the memory was saved. Relevance was how close the memory's embedding was to the current situation. Each score was scaled to between 0 and 1, the three were added with equal weight, and "The top-ranked memories that fit within the language model's context window are included in the prompt" (Park et al., 2023).
You don't have to copy those numbers, and for a coding agent with a handful of typed facts per repo you probably don't need scoring at all, since the lookup by fact type in section 8 does the job. Ranking starts to pay off once the agent keeps episodic memory, like past session notes, where there are thousands of items and only a handful bear on today's task.
Update old notes before adding new ones
The easiest way to write memory is to add a new note every time, and it's also the fastest way to end up with ten slightly different notes about the same thing. Mem0 handles this with an update step. Each new candidate fact gets compared against the 10 most similar memories already stored, and a model picks one of four operations, "ADD for creation of new memories when no semantically equivalent memory exists; UPDATE for augmentation of existing memories with complementary information; DELETE for removal of memories contradicted by new information; and NOOP when the candidate fact requires no modification to the knowledge base" (Chhikara et al., 2025).
That step keeps the store from filling up with duplicates, and it's also where things go wrong, because a model is deciding whether two notes mean the same thing. LangChain's docs warn that "some models may default to over-inserting and others may default to over-updating", so it needs measuring like any other classifier. Notice too that Mem0's DELETE throws the old memory away. Section 6 argues for keeping it with an end date, the way Zep does, so you can see what the agent used to believe.
Summarise old history into fewer notes
Raw history piles up fast, so the designs above all boil it down every so often. Generative Agents did this with what they called reflection. Once the importance scores of recent memories added up past 150, the agent read its latest 100 memories and wrote a few higher-level conclusions, which then went into memory alongside everything else. In their simulation that happened two or three times a day. MemGPT does something similar when the context fills up, pushing old messages out and writing "a new recursive summary using the existing recursive summary and evicted messages". ChatGPT's dreaming, from section 1, is a background version of the same idea.
For the coding agent this could be a weekly job that reads the week's session notes and writes one short summary per area of the codebase, like "checkout tests are flaky when fixtures reset, three sessions this week", and then the raw notes can be searched less often. What you lose is detail, which is the same trade LongMemEval measured, so keep the raw notes around and let the summary point back to them.
Put a cap on what's always loaded
A hard limit on the always-loaded part forces all of the above to happen. MemGPT's in-prompt block of user notes is a fixed size, so the model has to decide what stays in it. Claude Code does the same with its memory index. When the index gets near 200 lines or 25KB, Claude Code tells the model to shorten it and to "keep one line per entry, move detail into topic files, and merge or drop stale entries" (Anthropic, Claude Code memory docs).
A cap on its own doesn't make the memory better, and the Supersede results in section 7 show that giving the memory more room didn't help either. What a cap does is make pruning happen regularly, and it means the always-loaded part can't slowly grow until it's the noisy full history you were trying to avoid.
Keep a small core and a big archive
Put together, those mechanisms point to one shape. A few things go into every prompt, like the notes file and the handful of facts this task needs. Everything else lives outside the prompt, in a table or a searchable store, and gets pulled in when a task calls for it. Capping and pruning keep the core small, and the archive can keep growing without costing you anything until you load something from it.
For the coding agent, the core is the team's CLAUDE.md plus whichever project_facts rows this kind of task needs, loaded at the start of every session. The archive is every past session note, searchable by meaning, which the agent reaches for when a task touches code it has worked on before.
Supersession
Facts change and the stored row doesn't
Everything up to here treats a stored fact as correct or missing. The September failure was neither. The row was right when it was written, and then the team changed what uat was for.
The paper this section leans on is Supersede, a single-author preprint from June 2026 that we couldn't find anyone repeating yet, so treat what follows as one careful measurement. It calls a value that has been replaced by a newer one superseded, and answering from the current value while ignoring the replaced ones is the ability it measures. LongMemEval calls the same ability knowledge updates. It's a narrow ability, and the measurement below describes it as a distinct and unsolved failure.
It took the knowledge-update questions from LongMemEval, the ones where a user states a value and later changes it, and compared two ways of answering them. Full context puts every session into the model's context and asks the question, which is an upper bound where memory isn't the bottleneck. Bounded memory gives the agent a notes field of 300 characters, rewritten one session at a time, with the raw sessions never fed back. The gap between those two columns is the cost of maintaining a memory.
| Model | Full context | Bounded memory | Paired McNemar |
|---|---|---|---|
| gpt-4.1-mini | 82% | 63% | p = 0.0035 |
| gpt-4.1 | 91% | 64% | not computed |
| gpt-5.4 | 92% | 77% | p = 0.0033 |
Full-context accuracy climbs with model strength and saturates around 92%, which says the questions are answerable and the models can read. Bounded-memory accuracy climbs far more slowly, from 63% to 77%, which says the loss is happening at maintenance time and a stronger model recovers only part of it. On gpt-5.4 the paired test counted 13 questions that full context answered and the memory agent didn't, against 1 in reverse (Patel, 2026).
One objection to that is that 300 characters is too small a budget and the whole result is an artifact of an undersized memory. The paper tested that directly, and giving the memory more room didn't help.
| History | Memory budget | Accuracy |
|---|---|---|
| ~2 sessions | 300 characters | 68% |
| ~48 sessions | 300 characters | 28% |
| ~48 sessions | ~7,150 characters | 28% |
Lengthening the conversation 24 times over at a fixed memory took accuracy from 68% to 28%. Giving the agent 24 times more memory recovered none of it. The larger budget was used, since every one of the 25 answers changed between the two long-history rows, and it helped and hurt in equal measure, with the paired test counting 4 flips each way. The failure tracks how much history the memory has had to survive, and the compression ratio makes no difference to it.
That cell rests on 25 questions on gpt-4.1-mini, which is small. What it establishes is a direction and a rough size, and the useful part of the direction is that more memory doesn't fix this, and neither does a better model.

The failures behind those numbers are ordinary. The paper reports that errors are dominated by the relevant fact being compressed away or not overwritten, and gives examples of both. Asked how many Korean restaurants the user had tried, where the answer was four, the agent said the user had not mentioned any. Asked where a person named Rachel had moved to, where the answer was the suburbs, gpt-5.4 said there was no information about Rachel. The model read the question fine, but the updated fact was no longer in its memory.
What to do about it in a store you control
The experiment gives an agent one notes field and asks it to keep that field current, which is close to MemGPT's fixed-size block of user notes from section 4, and it's part of why the numbers look the way they do. A row in Postgres can do something a notes field can't, which is keep the old value.
Write a superseded fact as a new row and mark the old one superseded, with the id of the row that replaced it and the timestamp of the turn that did it. The store then holds a history per fact type with one row marked active, and the read policy takes that row. It's the same idea as the end dates Zep puts on graph edges, from section 4, and the opposite of Mem0's DELETE from section 6, which removes the contradicted memory entirely.
An agent that writes a bad replacement can be undone, because the previous row is still there with its source turn attached. A fact whose value has changed three times in a month is visibly volatile, and volatility is a signal the read policy can use, since a deploy target that has moved twice this year is worth confirming out loud. And a question about what the setup used to be is answerable, which helps when someone is working out why the agent did what it did back in February.
The cost is that your read policy now has to pick, and picking the newest active row is right until two rows arrive from sessions that overlapped. M6's advice about handling one turn per session at a time covers the common version of that. The uncommon version, where two developers run the agent on the same repo at the same time, needs the write to be conditional on the row it's superseding still being active, which is an ordinary compare-and-set in SQL.
UPDATE project_facts
SET state = 'superseded',
superseded_by = :new_fact_id,
superseded_at = now()
WHERE id = :current_fact_id
AND state = 'active';
-- 0 rows updated means another session moved first. Re-read and decide again.Not every change is the same kind of change
The deploy target moving from uat to preview is one kind of change, where the world moved and the old value was right at the time. There are others, and they need different handling.
A correction is when the memory was wrong from the start, like the extractor saving Priya as the payments reviewer from a request that was only for one PR. Superseding it would leave a history saying Priya was the reviewer until June, which never happened. A correction should mark the old row as wrong, and anything built from it, like a weekly summary that mentions Priya, needs rewriting, or the mistake comes back through the summary.
A temporary rule has its end built in, like "until March 9th staging goes to preview while uat is rebuilt". It should be saved with its end date at write time so it stops applying on its own. The coding agent keeps it beside the default value, so when the date passes the default takes over again without anyone having to remember to switch it back.
A request to forget is different again. When a developer says "forget what I said about the staging branch", the value has to go, and that includes copies you might not think of, like summary notes built from it and its embedding in the search index. Keep a tombstone row saying a fact of that type was deleted and when, without the value, so the history still makes sense. A deletion policy should also say plainly what's kept afterwards. OpenAI's help center, as of September 2026, says it "may retain logs of deleted saved memories for up to 30 days for safety and debugging purposes."
The read policy
Deciding at read time what to load
A read policy that loads every fact stored for a repo is fine for a while. Most of this team's repos have four or five active facts, and five short rows cost almost nothing. It's the simplest policy to ship, and it stops working once a repo has been around long enough to pile up history.
Facts accumulate and the context doesn't grow to match. A repo that's been worked on for three years has perhaps thirty rows. A package-manager migration and a couple of reorganised deploy setups account for most of them, and two of the thirty bear on the task she's asking about now. M5 measured what the other twenty-eight do, which is compete for the model's attention and cost money on every turn. And the more of the store you load, the more likely it is that a stale row is sitting in the context next to the current one, which is the failure section 7 measured.
Which fact types could bear on this turn at all is the first decision, and it's cheaper than it looks. A request to rename a function doesn't get any better for knowing the deploy target, so a small mapping from the kind of task to the fact types worth loading removes most of the store before any model call happens. The kind of task is usually obvious from the request, since "ship this" is a deploy, and a cheap classifier from M13 handles the rest. That is a lookup table someone maintains by hand, and it's the right amount of machinery for six fact types.
Which rows of those types to take comes next, and for most types it's the newest active row, which the supersession scheme from section 7 makes a one-line query. Where a type can legitimately hold several current rows, like a list of paths the agent shouldn't edit, all of them get loaded and the extra tokens are the price.
Where in the context the rows go is the last decision, and M5 already answered it. Facts go in the stable part of the context near the system prompt, which keeps the cached prefix intact across turns, and they go in as structured lines with their timestamps attached. LongMemEval found that the reading strategy alone moved question-answering accuracy by as much as 10 absolute points across three models, when the retrieved items were given in a structured format and the model was asked to note which ones it used (Wu et al., ICLR 2025). A row that reaches the context in a format the model reads past has cost you the same as not loading it.
Picking five out of hundreds
Typed facts are easy to load, since there are only a few per repo. Session notes are harder, because after a year the coding agent has hundreds of them and a task can be loosely related to dozens.
Do the steps in this order. First filter by owner, with the repo id from the session written into the query itself, so notes from another repo can't come back however similar they sound. That's a hard filter in code, and the model never gets a say in it. Then search what's left by meaning and keep the top 20 or so. Then re-rank those using the signals the search can't see, like how recent each note is and its importance score from when it was saved, which is the Generative Agents combination from section 6. Finally cut to a small budget, say five notes or 1,500 tokens, whichever comes first.
Say the developer asks the agent to fix a failing checkout test and the repo has 340 session notes. Search by meaning brings back the 20 closest, and most of them mention checkout or tests. Re-ranking pushes up three notes from the last month about the flaky fixture, and a note with a high importance score about the payment provider's test mode, and it pushes down a two-year-old note about a checkout layout change. The top five go into the context, each with its date.
Those numbers are invented, like the team's other numbers, and the budget is the thing to tune. If it's too small the useful note gets cut, and if it's too big you're back to the noise problem from section 6.
Load the age along with the value
A row's timestamp is the cheapest signal in the store, and it's easy to drop it on the way into the context without noticing.
Put it in. A line reading deploy_target: push to uat (confirmed 2026-02-11, 223 days ago) gives the model something the bare value doesn't, and it gives your own code something too, because a rule that holds back any deploy_target older than 60 days for confirmation is four lines and doesn't involve a model. The September request would have come back as a question about whether uat was still staging, and she'd have said no.
How long a fact type stays true is a property of the type and somebody has to write it down. Nothing about a row tells you that a code style rule tends to last for years and a deploy target tends to move, so the table of fact types carries a maximum age and what to do when a row exceeds it.
| Fact type | Trusted for | Past that |
|---|---|---|
code_style | No limit | Loaded as-is |
package_manager | 12 months | Loaded, flagged for confirmation |
test_command | 6 months | Loaded, flagged for confirmation |
reviewer | 6 months | Loaded, flagged for confirmation |
do_not_edit | 90 days | Loaded, flagged for confirmation |
deploy_target | 60 days | Held back, agent asks |
The check that applies the table runs on every loaded row before the model sees anything. It measures age from the last time a developer confirmed the row, and an unconfirmed row goes out as a question whatever its age, which is the rule from section 9.
TRUST_DAYS = {"package_manager": 365, "test_command": 180, "reviewer": 180,
"do_not_edit": 90, "deploy_target": 60} # code_style: no limit
HOLD_BACK = {"deploy_target"}
def placement(row, today):
"""Returns "load", "flag" or "ask" for one active row."""
if not row.confirmed:
return "ask"
limit = TRUST_DAYS.get(row.type)
age = (today - (row.confirmed_at or row.written_at)).days
if limit is None or age <= limit:
return "load"
return "ask" if row.type in HOLD_BACK else "flag"The September row was confirmed on 2026-02-11 and is 223 days old against a 60-day window, so placement returns "ask" and the push starts as a question.
Those windows are invented, like the team's other numbers, and what carries over to a real system is the shape of the table. A team that has run the agent for a year replaces the invented windows with the measured rate at which each type changes, which falls out of the supersession rows once you have been writing them for a while.
Memory and decisions
How memory should affect what the agent does
Loading the right memory still leaves the question of what the agent does when memory disagrees with something else it knows. That needs a rule written down in advance, because if it's left to the model, the answer depends on how the prompt happens to be worded that day.
The coding agent has three sources that can disagree. The first is the current request, what the developer is asking for right now. The second is the repo itself, which is the system of record for a lot of facts, since a lockfile says which package manager is in use and the CI config says what the test command is. The third is memory, which is what the agent was told before.
The rule is that memory is a default and never an override. The current request wins, because the developer is right there saying what they want. The repo wins over memory for anything the repo can answer, because it's the source of truth and memory is a copy of something someone once said about it. Memory fills in whatever neither of those settles, and a memory that's inferred or unconfirmed only gets to ask a question, never to act.
A few examples make it concrete. Memory says the tests run with pnpm test, and the developer asks for pnpm test:fast this time. The agent runs pnpm test:fast and leaves the stored fact alone, since a one-off request doesn't change the standing rule. Memory says the repo uses pnpm, but the repo has a package-lock.json and no pnpm lockfile. The agent goes with npm, tells the developer the memory looks out of date, and proposes a correction. Memory says staging deploys to uat, nothing in the repo says otherwise, and the fact is 223 days old. The agent asks before pushing, because a deploy is hard to undo and the fact is past its trust window.
That last example shows the other half of the rule, which is that how much the agent trusts a memory should depend on what it's about to do with it. Formatting a commit message from a remembered style preference is fine to do without asking. Pushing to a branch or skipping a reviewer on the strength of a memory should come with a confirmation, however recent the memory is.
def resolve(fact_type, request, repo, memory):
"""Returns (value, source, ask_first)."""
if request.states(fact_type): # 1. what they asked for now
return request.value(fact_type), "request", False
from_repo = repo.lookup(fact_type) # 2. lockfile, CI config, git
if from_repo is not None:
remembered = memory.get(fact_type)
if remembered and remembered.value != from_repo:
memory.flag_possible_correction(fact_type, from_repo)
return from_repo, "repo", False
remembered = memory.get(fact_type) # 3. confirmed and current
if remembered and remembered.confirmed and not remembered.expired:
return remembered.value, "memory", fact_type in HIGH_RISK
return None, "none", True # 4. nothing settles it, askThe two errors
What a wrong memory costs in each direction
A memory component can fail in two directions and the design turns on how far apart their costs are, which is the accounting M13 set up for classifiers and M14 applied to routes.
A missing memory is a fact the agent should have had and didn't. The developer gets asked something she already told it. The cost is a few seconds of her time and some annoyance, and it's bounded, because she can always answer again. A coding agent that forgets the deploy target asks which branch staging is on, which is an ordinary thing for a coding agent to ask.
A wrong memory is a fact the agent had and acted on, where the value no longer holds. The cost is whatever the action cost. For this team that's a half-built checkout page in front of a prospect, and then an afternoon of two people reverting it and cleaning up. There's no turn in which the developer gets to correct it, because nobody asked.
| Missing memory | Wrong memory | |
|---|---|---|
| What happens | The agent asks a question it has already had answered | The agent acts on a value that no longer holds |
| Who notices | The developer, immediately | Nobody, until the outcome arrives |
| Cost | A few seconds and some annoyance | The cost of the action, plus the cleanup |
| Recoverable in the session | Yes | No |
| What it looked like here | "Which branch is staging on?" | A broken checkout page in a sales demo |
The agent also asks questions on purpose, and those are a different thing. When the read policy holds back the 223-day-old deploy target, the agent asks "Staging was uat when you set this up on February 11th. Is that still right?" The row was found and loaded, and the question shows the developer the stored value with its date so she can confirm or correct it. A missing-memory question carries no stored value, like "Which branch is staging on?", because nothing was found. Both cost her a few seconds. A confirmation means the read policy did what the trust window told it to, and her answer is new evidence, which either refreshes the row's confirmation date or starts a supersession. A missing-memory question means the write side or the read side dropped something.
The gap between those two columns is what sets every threshold in the component, and here it's large. The extractor's confidence bar goes up. Supersession gets queued for review before anything is applied. The trust windows in the previous section get short. And the read policy holds old facts back and asks the developer, because asking costs her a few seconds and guessing wrong cost a demo.
That's the same trade M13 made on the ticket sort and M14 made on routes, and it keeps recurring because every one of these components produces an output your system acts on without a person looking. The design question is never how accurate it is on average. It's which of the two errors you would rather have, and how much of the cheap one you'll buy to avoid the expensive one.
Where the costs are close together the answer changes. A consumer assistant remembering that a user prefers metric units has a wrong-memory cost of nearly nothing. Writing eagerly and correcting on complaint is the right design there, and it's the one those products use. Work out the ratio for your own system before copying a threshold out of a library's defaults.
Failures and security
Memory failures and how to guard against them
Memory has failure modes of its own, and they tend to build up quietly over time. Each one needs a safeguard when the memory is written and another when it's read, because the check on the way in will sometimes miss.
A guess becomes a permanent fact
The extractor infers something, it gets saved with the same status as a stated fact, and from then on every session treats it as true, since nothing downstream can tell it was a guess. On the write side, the kind field and the evidence check from sections 3 and 5 stop most of these, so an inference either isn't saved or is saved as unconfirmed. On the read side, unconfirmed memories load as questions, like "I think Sam reviews payments, is that right?", and never as instructions.
One user's or one repo's memory reaches another
The coding agent works on several repos, and a consumer assistant might have millions of users. If the query that loads memory is built from anything the model produced, a prompt that mentions another repo can pull that repo's facts in. On the write side, the owner id comes from the authenticated session and never from the model's output. On the read side, the owner filter is part of the database query, the way M10 set up per-customer isolation, so there's no code path that loads memory without it.
Someone plants an instruction in memory
This is the attack M10 covers in full. Text the agent reads, like a README or an issue comment, contains something that looks like an instruction, and the agent saves it, so every later session loads it and the instruction outlives the conversation that planted it. On the write side, save only typed facts with values your code checks, never free-text instructions, and take evidence only from the developer's own messages, never from tool output or files the agent read. On the read side, load memories as data inside clearly marked delimiters, below the system prompt and separate from it.
The delimiters help the model see where each stored fact starts and ends and keep it apart from your instructions. They aren't a security boundary, because the model reads the text inside them the same way it reads everything else, and it can still follow an instruction it finds there. M10 went through this with spotlighting, which marks text with delimiters. It held up against a fixed benchmark of known attacks, and attacks tuned against it over many tries got through almost every time (Nasr et al., 2025). So the protection comes from checks in code. A deploy_target value has to pass value_is_valid from section 5, which for this type means matching a branch name, so a sentence of instructions can't be stored as one. And a push or anything else hard to undo goes through the confirmation from section 9 whatever the row says.
Old mistakes reinforce themselves
This one is easy to miss. Some ranking schemes boost a memory every time it's used, and Generative Agents measured recency as hours "since the memory was last retrieved". A wrong memory that keeps getting retrieved stays fresh, so it keeps winning the ranking, so it keeps getting retrieved. Summaries do something similar, since a weekly summary written from a wrong note copies the mistake into a new note, and now two notes agree with each other. And if the agent's own replies are ever saved as evidence, it can end up citing itself. On the write side, never accept the agent's own output as evidence, and keep a pointer from every summary back to the notes it came from so a correction can find the copies. On the read side, measure recency from when a memory was written or last confirmed by a person, and re-confirm high-stakes facts on a schedule.
| Failure | When writing | When reading |
|---|---|---|
| A guess becomes a fact | Kind field and evidence check, inferences unconfirmed | Unconfirmed memories load as questions |
| Memory crosses owners | Owner id from the session, never the model | Owner filter inside the query |
| A planted instruction | Typed facts only, evidence from the user's own words | Delimited as data, and risky actions confirmed in code |
| A mistake reinforces itself | No self-citation, summaries point to their sources | Recency from writing or confirming, scheduled re-checks |
Memory also changes the security review M10 ran, since a write persists past the conversation that made it. The narrowing M10 used against planted instructions, saving only facts the developer asked for and confirmed, is the same caution that meant the June sentence was never recorded. So security and coverage pull against each other here. The way through is the split from section 5, where the extractor can suggest facts freely because replacing a stored one still goes through review.
Measuring it
Measuring whether the memory helped
Memory is harder to measure than the components before it because a single turn doesn't contain the evidence. A classifier's output can be judged against a label on the same ticket. Whether a memory was right requires knowing what the developer said three months ago and what has happened since, which means a test case in this module is a sequence of sessions and not a message.
The public benchmarks are built in that shape, and each one scores a different task, which decides whether its number is a ceiling for your system or background.
| Benchmark | The task it scores | Size | Date |
|---|---|---|---|
| LongMemEval | Questions about a user's own chat history with an assistant, across extraction, multi-session reasoning, knowledge updates, temporal reasoning and abstention | 500 questions, about 115k tokens of history each | ICLR 2025 |
| LongMemEval-V2 | Whether a web agent's memory holds what it learned about one environment, including workflows and the gotchas that only show up in use | 451 questions, up to 500 trajectories and 115M tokens | May 2026 |
| The knowledge-update subset | Answering from the current value of a fact the user changed partway through | 78 questions, oracle split | Used in Patel, 2026 |
A coding agent that remembers facts about a repo is closer to the second row than the first, since its memory is about an environment and not about a person. LongMemEval-V2's headline, where a memory method built on a coding agent reached 72.5% average accuracy against 48.5% for the strongest retrieval baseline, comes from web agents learning a customised web application (Wu et al., 2026). That's the nearer comparison for this team, and it still isn't their repo.
None of these replaces a dataset from your own traffic, for the reason M1 gave about every other component. What the public numbers give you is the shape of the test, which is worth copying even where the numbers are about somebody else's task.
A test that follows a memory through its whole life
A good memory test case walks a fact through the same life it would have in production. Introduce it in one session, add a few sessions about other things so it isn't sitting in recent context, then change or delete it in a later session, and then ask something in a final session that only has the right answer if every step in between worked. LongMemEval's knowledge-update questions are built this way, and the coding agent's test cases copy that shape.
Each test case should exercise one kind of statement from section 3, so a failing test case points at one step. A test case where the developer asks the agent to forget the deploy target tests deletion. One where she mentions another repo's package manager tests ownership. One where she asks for pnpm test:fast once tests whether a one-off request stays one-off, and one where the final request contradicts memory tests the precedence rule from section 9.
Score more than the final answer. Look at the store after every session, which tells you whether the right thing was saved and whether anything was saved that shouldn't have been. Look at what was loaded for the final question, which tells you whether the memory was found when it was needed and whether deleted or expired memories stayed out. Then look at the answer. That way a failure points at the stage that caused it. A wrong answer with the right row loaded is a reading problem, and a wrong answer with nothing loaded is a saving or retrieval problem, which the store after each session will sort out.
Finally, run the same test cases with memory turned off. Some questions can be answered without any memory, like the ones where the final request states the value itself, and if the memory system doesn't beat the no-memory run on the rest, it's adding cost and risk without helping.
Building the dataset from traffic you already have
The agent's session logs hold every session on every repo, which means the raw material is already on disk. A test case is one repo's sessions in order, a question to ask at the end and the answer that is correct given everything in the sequence.
The cheap way to find good ones is to look for repos whose stored facts have changed. Every supersession row from section 7 marks a point where something the team told the agent stopped being true, and a question asked after that point has a known correct answer and a known tempting wrong one. Fifty of those, drawn from real traffic across the six fact types, is a dataset that catches the failure this module is about. Where the store has no supersession history yet, the same rows can be found by asking a model to scan pairs of sessions for a value that changed, and then having a person check what it returns.
Score two things separately, because a pooled number hides which half is broken. Retrieval asks whether the row the answer needed was in the context at all, which your traces answer without a model call. Answer correctness asks whether the reply used it. A 2026 study of memory pipelines found that separating evidence identification from the step that applies the answer policy was where most of its gains came from, and that swapping only the component which applies the policy moved accuracy by 2.0 percentage points on average (Reddy and Challaram, 2026). The same paper found no significant advantage when it checked its pipeline on LongMemEval, so it's a result about one kind of question. If you measure only the final answer you can't tell those apart, and they have different fixes.
Count how often the agent asked for a value from scratch when the store held an active row for that fact type, confirmed and inside its trust window. That is your missing-memory rate on live traffic. Leave out the confirmation questions the read policy asked on purpose, the ones that quote a stored value back with its date, because counting them would make every shortened trust window look like a memory failure. Count those separately as a confirmation rate, and track how often the developer's answer changes the value. A fact type whose confirmations almost never change anything has a window that's too short. The confirmation count comes straight from the read policy's log, since placement records every row it held back. The missing-memory count needs a cheap classifier over the agent's questions to tag which fact type each one asks about.
Build it and test it
A small memory lifecycle, run and measured
Everything in this module fits into one pipeline small enough to run on a laptop, and running it is how you find out whether yours works. The script research/M16-memory-lifecycle.py does this for the coding agent, and before looking at any numbers it's worth being clear about what it is. It's 36 hand-written test cases, four in each of the nine categories listed in the second table below, run on one local model, qwen3:8b through Ollama at temperature 0. The repo and the sessions are invented, and so is every value in them. That makes the results a sighting of how the pieces behave, and nowhere near a benchmark.
Each test case is a few sessions followed by a final question in a later session. After every session, one model call reads the transcript and proposes candidate memories, each with a type, a value, a kind, the repo it's about, an end date if there is one, and a quote as evidence. The same candidates then go through three different pipelines.
- No memory keeps nothing, so the final question sees only itself.
- Naive memory saves every candidate the extractor returns and loads everything of the right type, and it never checks, replaces or deletes anything.
- Lifecycle runs the checks from this module. It saves only standing facts about this repo whose evidence appears in the developer's own words, replaces old values with the history kept, deletes on request, drops dated rules once they expire, and loads the newest active row per type with its date.
The answering prompt is the same in all three and includes the rule from section 9, so a value stated in the request wins over memory, and the agent answers UNKNOWN if neither settles it. Any difference between the rows comes from the pipeline.
def validate(c, s):
"""The lifecycle's checks. Returns (ok, reason)."""
if c.get("type") not in TYPES:
return False, "type not allowed"
ev = norm(re.sub(r"^\s*(developer|dev|user)\s*:\s*", "",
c.get("evidence") or "", flags=re.I))
if not ev or ev not in norm(s["lines"][0]): # the developer's words
return False, "evidence not found in what the developer said"
if (c.get("about_repo") or REPO) != REPO:
return False, f"about another repo ({c.get('about_repo')})"
if c.get("kind") == "one_off":
return False, "one-off request"
if c.get("kind") in ("inferred", "observed"):
return False, f"{c.get('kind')}, not stated as a rule"
return True, "ok"| Pipeline | Answered correctly | Acted on a stale, deleted or wrong value |
|---|---|---|
| No memory | 15 of 36 | 2 of 36 |
| Naive memory | 13 of 36 | 14 of 36 |
| Lifecycle | 28 of 36 | 4 of 36 |
The naive pipeline did worse than having no memory at all, 13 correct against 15, and it acted on a wrong value 14 times. The reason shows up in what it saved. It kept 12 rows that should never have been saved, like one-off requests and facts about another repo, and it kept things the extractor made up. After a session about fixing a date picker's CSS, the extractor proposed deploy_target: production, with its own reasoning given as the evidence, "The developer is working on a production-ready feature (date picker), implying deployment to production is standard." The naive pipeline stored it, so when the recall test asked where staging deploys, it loaded both uat and production and the model answered UNKNOWN.

The no-memory run got 15 right because some questions don't need memory. All four precedence test cases state the value in the request itself, and in every inference and forget case the right answer is UNKNOWN, which an agent with no memory gives for free. That's the reason to always run the baseline. Without it, the naive pipeline's 13 would look like memory helping a bit.
| Category | No memory | Naive | Lifecycle |
|---|---|---|---|
| Recall after unrelated sessions | 0 | 2 | 4 |
| A deliberate update | 0 | 1 | 3 |
| A change mentioned in passing | 0 | 0 | 3 |
| A one-off request | 0 | 2 | 3 |
| A fact about another repo | 2 | 0 | 2 |
| Something only inferable | 4 | 2 | 4 |
| A request to forget | 4 | 1 | 2 |
| A rule with an end date | 1 | 1 | 3 |
| The request contradicts memory | 4 | 4 | 4 |
The lifecycle rejected 25 candidates. Seven were one-off requests, three were about another repo, one was an observed event and one had a type that isn't allowed. The other 13 were rejected because the quoted evidence wasn't in what the developer said, and reading all 13 shows every one was a correct rejection. Besides the invented production deploy target, they included code_style: mobile-first and test_command: npm test, both guessed from the same CSS fix, plus quotes taken from the agent's own replies. Some of the 13 repeat, since the distractor sessions are shared between the recall test cases. The evidence check is a single string comparison, and in this run it caught more bad candidates than every other check put together.
Where the lifecycle still failed
Eight test cases failed with the full lifecycle, and the traces split them into two groups.
Three were reading failures. The right row was loaded and the model answered with something else, "staging" when the loaded fact said preview, or "main" when a dated rule said preview until March 9th. A small local model is part of that, and a stronger model on the answering step would likely fix some of it, but that wasn't tested here.
Five were writing failures, and each one points at a check to add.
- A one-off saved as standing. "Sam's out today, so get Priya to review this one payments PR" came back labelled standing, so Priya became the payments reviewer. A cheap check for words like "today" and "this one" would catch it, on top of the model's label.
- The wrong owner. "My old team at my last job had Dana review every payments change" was filed as a fact about this repo. Reviewer facts are the kind that should need the developer's confirmation.
- The wrong type. "Tests here run with
pnpm test" was saved as a package manager, so the test command was never stored. - Two forgets that deleted nothing. Once, the extractor named the thing to forget as "staging branch", which isn't a fact type, so the delete matched no rows. The other time it read "you can forget the rule about the legacy folder" as the rule itself and kept it.
The first forget failure is a good example of reading a trace, since everything in it looks like it worked. The delete ran, and the only sign of trouble is the row count.
{
"session_2_forget": ["staging branch"],
"trace": ["ADD deploy_target=uat", "DELETE staging branch (0 rows)"],
"loaded": ["deploy_target: uat (saved 2026-02-02)"],
"answer": {"answer": "uat"}
}The fix is to treat a forget request like any other candidate, checking its target against the allowed types, and to treat a delete that touches zero rows as an error that gets logged and surfaced. Without that the developer is told the memory was dropped, and the next session deploys to uat anyway.
Two things the traces caught in the test itself
The first run of the script rejected correct facts. The extractor quoted evidence with the speaker's label attached, "Developer: We use pnpm here.", and the label isn't part of the message, so the exact-match check failed. The trace showed "REJECT package_manager=pnpm: evidence not found in what the developer said" next to a quote that was plainly right, and the fix was to strip the label before comparing. The trace from that run is saved as research/M16-run1-evidence-trace.txt.
The second run asked the deploy questions as "Ship this to staging. Which branch?", and the model read that as the branch the code was on. It answered "main" in five lifecycle test cases, three of them with the right fact loaded. Rewording to "Which branch do you push it to?" fixed two, and the other three are the reading failures above. That run scored 26 for the lifecycle, 11 for naive and 12 for no memory, and the numbers above are from the reworded run. Both of these were bugs in the test and neither was the memory's fault, and both were only visible because every step wrote down what it did.
Build a memory test harness for our agent. Each test case is a list of sessions (developer message, agent reply, date) and a final question with the expected answer and any wrong answers that count as using a stale or deleted memory. After each session, run our extractor and pass its candidates through our memory pipeline. Run every test case three ways, with no memory, with every candidate saved, and with our full pipeline, and use the same answering prompt for all three. Log every step as a trace line (ADD, NOOP, SUPERSEDE, REJECT with a reason, DELETE with the row count), save the loaded memories and the answer, and print accuracy per category for each pipeline. Treat a DELETE that touches zero rows as a failure in the report.
In production
Expire the old rows and review the new ones
Run expiry as a scheduled job that removes or archives old rows. A rule in the read path only hides them, and they stay in the store where an audit or a second code path will find them. Anthropic's own guidance for its memory tool says to periodically delete memory files that haven't been accessed in a long time, which is the same reasoning applied to files (Anthropic, memory tool docs). Rows past their trust window get archived on a schedule, and the archive stays reachable for a dispute while the answering path stops seeing them.
Review is what the queued supersessions from section 5 are for. A weekly pass over the queue, which for this team is a few dozen rows, tells you whether the extractor is proposing sensible replacements, which is the M13 measurement you owe it. It also surfaces the fact types where the team's setup changes often, which is what turns the invented trust windows in section 8 into measured ones.
Verification is where M15 already named the problem. A verifier checks an answer against the material it was supposed to rest on, and for a memory-backed answer that material is a row your own system wrote months ago. Checking the reply against the row tells you the reply used the row faithfully. It tells you nothing about whether the row is still true, which no check inside the conversation can settle. What you can do is make the agent say which stored fact it's about to act on, so the developer is the check, which is why the September push should have started as a question.
Let people see and fix what the agent remembers. ChatGPT has a memory summary page where you can type a correction or delete something it knows about you, and Claude Code's docs say of its memory folder, "Everything is plain markdown you can read, edit, or delete." For the coding agent that means a command that lists the repo's active facts with their dates and lets a developer edit or retire one. It's also the cheapest way to fix a wrong memory, since the person who knows it's wrong can change it themselves.
Putting it together
Putting it together
Long-term memory is anything your agent keeps from one conversation to use in a later one. Saving a note and loading it back is easy. What makes it a module of its own is everything around that, which comes down to three questions you should be able to answer for your own agent. How does it decide what to remember, how does it keep what it remembers accurate, and how does it use a memory when it's relevant and ignore it when it isn't?
The coding agent's memory layer ends up looking like this, and it's a reasonable starting point for any agent that works with the same people or projects over time.

It keeps three kinds of memory in three places. Instructions the team writes by hand go in a CLAUDE.md in the repo, which loads every session and which a person reviews like any other file. Specific facts about the project go in the project_facts table, one typed row each, with a timestamp and a pointer back to the turn they came from. Past sessions are saved as notes with embeddings, so the agent can search them when a task touches code it has worked on before.
Writing happens in two places. During the session, the model calls remember_fact when a developer asks it to save something, and the developer confirms facts the agent will act on later. After each session, a separate pass reads the transcript and proposes candidate facts, each scored for how much it's worth keeping. A candidate is only saved if it's a standing fact about this repo, with evidence that appears in the developer's own words. Those checks depend on the extractor's labels, so they stop most one-off requests and facts about other repos, and some still get through. In section 13 a one-off request about one payments PR came back labelled standing, and a reviewer from a developer's last job was filed under this repo. So the finished layer also checks for words like "today" and "this one" and asks the developer to confirm reviewer facts, which section 13 suggested and didn't test. New facts that pass are written straight away. A candidate that would replace an existing fact goes to a review queue, and when it's accepted the old row is marked superseded with an end date and kept.
Loading happens at the start of each session. The CLAUDE.md goes in whole, and it stays short because it's capped and trimmed. The agent works out what kind of task this is, loads the fact types that kind of task needs, takes the newest active row of each, and puts them near the system prompt with their age attached. Facts past their trust window are held back, and the agent asks about them. Past session notes are only searched when the task calls for them, and only the top few come back.
When a memory disagrees with something else, the current request wins, then the repo, and memory only fills the gaps, with a confirmation first for anything hard to undo like a deploy. The owner id comes from the session and the owner filter lives in the query, only typed facts are ever saved, and memories load as delimited data next to the system prompt, never inside it. The delimiters help the model read the facts, and what stops a planted line from doing damage is the value check on the way in and the confirmation before a risky action.
Keeping it healthy is a few scheduled jobs. A weekly summary turns the week's session notes into short notes per area of the code. An expiry job archives rows past their trust window. A weekly review of the supersession queue doubles as the check on how well the extractor is doing. The team measures all of it with test cases that follow a fact through its whole life, run with and without memory and scored step by step. On the 36 test cases in section 13 that pipeline answered 28 correctly, against 15 with no memory and 13 with naive memory, and its traces pointed at every one of the failures that were left.
If your agent is simpler than this, start smaller. A single notes file and a hot-path tool with confirmation is a fine first version, and you add the table when facts start changing, the search when there's too much to load, and the background pass once you notice things being said in passing and lost.
M24, State Understanding, takes the agent's picture of what's true right now and extends it past facts about a person to the state of a system it's working in. M19 covers retrieval, which is how the searchable part of the store works. M30 runs all of this inside agents whose sessions are long enough that facts change within a single run.
Checkpoint · recall · 6 questions
What the module said
- 01
What are the two halves of a memory component?
- 02
What does "superseded" mean when applied to a stored fact?
- 03
In the LongMemEval study of shipping consumer assistants, what were the two failure modes the authors named?
- 04
The coding agent needs to remember "run pnpm test before committing". Which kind of memory is that, and where does it belong?
- 05
Why does a test case for memory have to be a sequence of sessions?
- 06
A developer says "just run the checkout tests, I'm in a hurry." What should the memory layer do with it?
0 / 6 answered
Checkpoint · understanding · 8 questions
Reason it through
- 01
A team doubles their agent's memory budget after seeing stale-fact errors, and accuracy doesn't improve. What does the Supersede measurement suggest is happening?
- 02
Why does the coding agent run an after-session pass as well as letting the model call
remember_factduring the conversation? - 03
A coding agent's wrong-memory cost is a bad deploy and its missing-memory cost is one extra question. Which way should the extractor's confidence threshold move?
- 04
A team loads a user's entire history into every call so nothing gets forgotten. What did LongMemEval find about that approach?
- 05
Why should retrieval and answer correctness be scored separately when measuring memory?
- 06
What does loading a fact's age into the context buy you that the value alone doesn't?
- 07
Memory says the repo uses pnpm, but the repo has a
package-lock.jsonand no pnpm lockfile. What should the agent do? - 08
A reviewer suggests wrapping loaded memories in
<memory>tags and calling the planted-instruction risk handled. Why isn't that enough?
0 / 8 answered
Checkpoint · debugging · 6 questions
Debug it
- 01
A developer asks the agent to ship to staging and it pushes to a branch the team stopped using for staging in June. The traces show a clean session, the read policy loaded one
deploy_targetrow, and the row's value matches what she set up in February. Where is the bug? - 02
Your supersession dashboard shows
test_commandrows being replaced on the same repo several times a week, and the agent keeps running only a handful of tests before saying a change is done. What's the likeliest cause? - 03
The agent keeps asking which package manager the repo uses, even though the traces show an active
package_managerrow for it. The row is active and inside its trust window. What do you check next? - 04
After a deploy, memory-related errors rise on long-lived repos and stay flat on new ones. Fact extraction accuracy is unchanged on your dataset. What changed?
- 05
On the web-app repo the agent suggests
make test, which is the api repo's test command. The trace shows the memory query was built from repo names the model mentioned in its plan. What's the fix? - 06
The team shortens the
deploy_targettrust window from 60 days to 30, and the missing-memory rate on the dashboard doubles that week. Every new question in the traces quotes the stored branch and its date back to the developer. What's going on?
0 / 6 answered
Go deeper
OpenAI (June 2026): Dreaming, better memory for a more helpful ChatGPT · Wu et al. (ICLR 2025): LongMemEval, Benchmarking Chat Assistants on Long-Term Interactive Memory · Patel (2026): Supersede, Diagnosing and Training the Memory-Update Gap in LLM Agents · Wu et al. (2026): LongMemEval-V2, Evaluating Long-Term Agent Memory Toward Experienced Colleagues · Reddy and Challaram (2026): Reliable Post-Retrieval Assembly for Agent Memory · Park et al. (2023): Generative Agents, Interactive Simulacra of Human Behavior · Packer et al. (2023): MemGPT, Towards LLMs as Operating Systems · Sumers et al. (2023): Cognitive Architectures for Language Agents · Chhikara et al. (2025): Mem0, Building Production-Ready AI Agents with Scalable Long-Term Memory · Rasmussen et al. (2025): Zep, A Temporal Knowledge Graph Architecture for Agent Memory · Anthropic: how Claude Code remembers your project · LangChain: memory concepts, writing in the hot path or in the background · Anthropic: the memory tool, client-side file operations
That's the last one written so far
Pick your next module from the board.
