APIs, integration & security — in depth

Memory Decay and Forgetting Policies for Always-On Agents

Agents need two separate fixes for memory failure, applied in a strict order.

Contributing Editor · · 9 min read
Cover illustration for “Memory Decay and Forgetting Policies for Always-On Agents”
Agent State & Memory · October 8, 2026 · 9 min read · 2,118 words

The signal always arrives the same way. The model didn't change. The weights are identical. The memory store underneath it grew, and the geometry of retrieval began surfacing stale facts ahead of current ones.

Operators describe this as the agent "forgetting," but that word hides two separate mechanical failures with two separate fixes, and the order you apply those fixes in is not optional.

The first failure is session amnesia. Large language models are stateless: every inference call sees only what's sitting in the context window at that moment. Conversation history, user preferences, task state in progress, none of it persists unless something explicitly carries it forward. This produces a behavior any operator can spot immediately: users re-briefing the agent at the start of every session, repeating context they already gave it an hour ago.

The second failure is organizational ignorance, and it's far harder to catch because it doesn't look like a failure at all. An agent that never had access to governed metric definitions, data lineage, or business policy produces answers that sound just as confident as correct ones. Jitender Aswani of Starburst Data documented the scale of this gap: the same foundation model scores roughly half accuracy on enterprise data questions when it lacks structured business context, and climbs past 90% once that context is supplied. Same model, same weights, same prompt style. The only variable is whether the organization's own definitions were ever made available to it.

These two failures have to be fixed in sequence, not in parallel. An agent that cannot hold context across turns cannot reliably apply organizational context even when you hand it that context on a platter. Session memory is the prerequisite. Skipping it dilutes every investment in metric governance and lineage with an agent that forgets which metric it was even discussing three turns ago. The decay that follows is a predictable consequence of memory architecture, not model quality, because it follows a fixed order.

The two forgetting failures that existing memory benchmarks miss entirely

MEMPALACE scored competitively, though not at the top, on the Memora recall benchmark. Nothing in between. That single result says something important about how the field has been measuring memory quality: recall accuracy and the ability to forget on command are not two ends of the same scale. They are different capabilities, and a system can max out one while failing the other completely.

Most memory benchmarks test exactly one thing: can the agent retrieve a fact it was given earlier? MTEB and BEIR, the dominant retrieval benchmarks, cover tasks like classification, clustering, and semantic similarity, but neither probes deletion, forgetting, or whether stored facts stay current over time. That's a reasonable thing to measure. It's also not where production systems break.

Production failures are almost entirely on the forgetting axis, not the recall axis. A password a user rotated three months ago still gets suggested by the agent. In every one of these cases, retrieval worked perfectly. The system found the fact. The application needed the system to find nothing, and it had no mechanism to make that happen.

FORGETEVAL, a benchmark built by Yang, is one of the few evaluation suites that actually targets this axis. Scoring runs on deterministic substring match, with no LLM judge anywhere in the loop, which keeps the results reproducible in a way that LLM-graded benchmarks often aren't.

The practical consequence for operators is direct. Evaluating a memory system purely on retrieval benchmarks selects for systems that will pass every test you ran and still fail in production, because the tests never asked the one question that matters most: can this system be commanded to forget?

What "forgetting" decomposes into: four levers, not one dial

A forgetting policy is four distinct levers, each operating at a different point in the memory lifecycle, each doing something different to the underlying data. Treating them as interchangeable is how systems end up applying the wrong operation at the wrong moment.

The four levers, drawn from consolidation research with roots in older cognitive-science models of memory decay, are importance, merge, decay, and eviction. Eviction removes facts from the store permanently.

Importance and merge are write-time decisions. Decay and eviction happen later, at read time or during lifecycle management. That distinction matters because write-time filtering is the cheapest place to control quality: everything that makes it into the index has to be retrieved, reranked, and judged for as long as it exists, so the bar for entry should be high.

Two patterns dominate write-time importance filtering. The other extracts atomic facts from a conversation and stores only the facts that survive extraction, so conversational filler, repeated greetings, and procedural chatter never become memories at all.

Merge solves a specific and common problem. The agent ends up picking whichever fact its retrieval system happened to score higher on a given turn. The same question can get two different answers depending on timing.

Decay and eviction cause the most confusion, because they sound like variations on the same idea and do the opposite thing to the data. Decay doesn't delete anything. Eviction actually removes the fact, reclaims the space, and guarantees it's gone. That's the tool for data that must never resurface: GDPR erasure requests, rotated credentials, expired authorizations.

Decay is the lever most often skipped, and the most dangerous one to skip in a long-running agent. Decay only dampens memories that go unused. A memory retrieved constantly never decays, because every retrieval reinforces it, as in a documented case at Northwind, where a customer success agent greeted a returning customer by referencing a job the customer had left months earlier. Nothing was wrong with retrieval. The lever that should have suppressed an outdated fact was simply never applied.

A workable retention ladder follows from these four levers. Semantic memory, the durable facts about a user or entity, stays indefinitely until something supersedes it.

How architectural placement of the control plane determines which failures a memory system can recover from

The question that matters for any memory system is where the LLM sits relative to the operations that mutate memory: supersede, release, purge. That placement determines which categories of forgetting failure the system can recover from, and no single placement covers every category.

Yang's study tested thirteen system configurations across a shared 385-case adversarial surface and found three placement regimes, each with a different and only partly overlapping set of strengths. Move the LLM to a mutation-time hook, so it sits at the point where memories get superseded, released, or purged, and intent-aware deletion recovers to 78 to 85%, with gains showing up across nearly every other category at the same time.

The cost appears only on the mutation side, not on every read, because the recall hot path itself stayed unchanged.

This has a direct implication for a common design pattern. Teams that add LLM intelligence only at inscribe time, which is a frequent choice because it's simpler to build, gain strong canonicalization and lose nothing on that front. They remain blind to intent-aware deletion, the failure mode behind GDPR violations and credentials that keep surfacing after rotation. The architecture looks sophisticated and still can't do the one thing regulators and security teams actually need it to do.

FORGETEVAL's five structural families, supersession, decay, amnesia, purge, and drift, each carry their own placement requirement. A system built and tuned around one family will carry systematic blind spots in the others, by construction, not by oversight. This is also where constraint-decay failures in agents occur: when an agent violates a rule it enforced correctly a few turns earlier, the model hasn't gotten worse. The attention weight on that constraint simply dropped below the threshold needed to enforce it, because nothing in the control plane reinforced it. That's a memory architecture gap, not a reasoning failure.

These three questions matter for any team evaluating a memory system: where does the LLM sit relative to supersede, release, and purge operations? Can the system actually be commanded to forget something, not just asked to recall it? Can deletion be proven?

Why memory staleness is an invalidation problem, not a retrieval problem

Staleness is the complaint operators hear most often from production memory systems, and the instinct is to treat it as a retrieval tuning problem: better embeddings, better reranking, a smarter top-k cutoff. Staleness isn't a search quality issue but a provenance issue, so better embeddings, reranking, or top-k tuning don't fix it. A memory system needs to carry evidence of its own currency, not just evidence of its own content.

Three forces push staleness to the top of the list of consolidation challenges in any long-running agent. Context economics is the first: when retrieved context contains contradictory facts about the same entity, even a strong model can't produce a coherent answer from it, because nothing in the context tells it which fact is current. Entity drift is the second: users change jobs, codebases change dependencies, and the old fact and the new fact sit in memory with equal weight until something explicitly supersedes one of them. Index precision degradation is the third: as a memory store grows, more documents compete as near-duplicates for the same top-k retrieval slots. More documents competing means more opportunities for a stale fact to outscore a current one purely on retrieval mechanics.

Their evaluation found that memory usage improved code review precision and recall even when the memory pool included entries from abandoned branches, which is a strong argument for designing verification into the system rather than treating staleness as disqualifying.

That verification pattern generalizes well beyond code review. Every claim an agent writes to memory can be persisted alongside versioned evidence, the source artifact it was derived from at the time. When that underlying evidence changes, the claim gets flagged stale and stays in that state, durably uncertain, until something re-verifies it. Staleness becomes a first-class state a record can be in, inspectable and queryable, rather than a side effect that only shows up when a user notices a wrong answer.

Systems built on vector embeddings or opaque database records cannot prove that a deletion actually happened. Those systems can report that a deletion happened, but they often can't verify it actually did: the vector may be gone from one index while a copy lingers in a cache, a backup, or a derived table. A more rigorous approach re-counts residual rows across the full storage surface after every deletion and tracks deletion and provable deletion as two separate, explicit flags. That distinction is the operative difference for satisfying GDPR Article 17, for proper credential hygiene after a rotation, and for any user who needs to trust that a removal request did what it claimed to do. Applying the four levers from earlier, specifically eviction, matters only if the system can demonstrate the eviction took, not just assert it.

GDPR's right to erasure, codified in Article 17, gives every data subject the right to obtain erasure of their personal data from a controller without undue delay. For traditional databases, that's a solved problem: find the row, delete the row. For agent memory, it's a problem most enterprise teams haven't fully mapped yet, because agent memory doesn't live in one place or one format.

Agents retain and build on historical data across sessions by design, which makes them useful over time, but it also means they accumulate detailed profiles of users that can outlast the task those profiles were built for. Unlike a database record, agent memory might exist as a vector embedding in a semantic store, as weights adjusted through fine-tuning, or as entries in a shared enterprise context repository accessed by multiple agents at once. Each of those forms requires a different deletion mechanism, and a system that only knows how to evict from one of them will quietly fail erasure requests for the others.

The EU AI Act adds a second compliance layer on top of GDPR, and it runs on its own separate timeline. The Act's general application begins in August 2026, with obligations for high-risk systems arriving in December 2027. None of these phases replaces GDPR's existing erasure requirement. They sit alongside it: a memory system built to satisfy Article 17 today still has to account for a second, staggered set of obligations arriving over the following two years.

The architectural argument running through this entire piece converges here. A memory system that cannot prove eviction, that conflates decay with deletion, or that places its only LLM at inscribe time rather than at the mutation layer, carries a compliance gap, with a regulator on one side and a deadline on the other.

Sources

  1. AI Agent Memory Loss: Fix Session Amnesia and Context [2026]
  2. Control-Plane Placement Shapes Forgetting: An Architectural Study of Agent Memory Across Thirteen System Configurations
  3. Governing Evolving Memory in LLM Agents: Risks, Mechanisms, and the Stability and Safety Governed Memory (SSGM) Framework
  4. Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents

More in Agent State & Memory