Vector Memory Retrieval Tradeoffs in Long-Running Agents
Agents running continuously hit hard limits in flat-file memory systems that weren't built for them.

Picture the memory of a production AI agent as two files sitting on a disk: a 20 KB text file called MEMORY.md, and a folder of daily log files written in plain Markdown, one entry appended after another. That is the actual architecture behind OpenClaw, the leading open-source agent runtime, running in more than 250,000 deployments, and it works, not a simplified description. It also has a hard ceiling, and that ceiling is the subject of this piece.
Flat files like these were built for a world where a session started, ran for a few minutes, and ended. Short conversations don't need much memory, and a capped file plus an append-only log is a reasonable way to hold what little needs holding. The trouble starts when the job changes shape: agents now run continuously, execute tool calls across hours or days, and are expected to keep acting coherently on everything they've learned since the session began. A 20 KB cap and a daily log were never built to hold weeks of accumulated knowledge, and they don't bend to fit. They break, in specific and measurable ways, because the shape of the storage no longer matches the shape of the job now being asked of it.
The four failure modes that compound over time in flat-file systems
Longitudinal measurement of OpenClaw's memory behavior turns up four distinct failure modes, and the important thing to understand is that they don't sit side by side as independent problems. They feed each other.
The first is context collapse. When the MEMORY.md file hits its 20 KB cap, something has to give, and what gives is whatever gets truncated first, with no regard for importance. There's no prioritization step that asks whether the discarded text mattered. It's just a hard cutoff, the same way a glass overflows the instant it's full, regardless of what's in it.
The second is compaction discontinuity. As the session grows, older context gets summarized to make room, and this summarization step produces a measurable break in the agent's behavior. A 2026 analysis of OpenClaw found that 62% of context-compaction events produce a measurable behavioral break: the agent acts differently right at the moment its older context gets compressed, not gradually, but as a step change.
The third is structural blindness. Flat text has no concept of relationship. If two names appear in the same paragraph, a flat-file retrieval system has no way to know whether they're connected as employer and employee, or whether they just happened to land in the same sentence by coincidence. The system is reading characters, not structure, and it pays for that difference whenever a question depends on knowing how two facts relate rather than whether they both exist somewhere in the log.
The fourth is the absence of an attribution loop. When an agent executes a tool call, no one links the outcome (success or failure) back to whichever memory entry informed that decision. Without that link, the system has no way to learn which memories were useful and which led it astray. It can't get better at retrieval over time, because it has no record of what retrieval produced.
These four don't stay in their lanes. A compaction event that causes a behavioral break also destroys the provenance trail that an attribution loop would need to function, so the system can't learn to avoid the same compaction cliff next time. Context collapse deletes information before structural blindness even gets the chance to misread it. Each failure mode erodes the conditions the others need to be fixed, which is what makes the overall decline compound rather than plateau.
What the compaction cliff specifically does to safety-critical knowledge
Compaction discontinuity deserves its own closer look, because new research shows it is a correctness problem, and the distinction matters most wherever a mistake carries real consequences.
Compaction is often treated as a neutral step, just the system tidying up old text to save space. Researchers Zerhoudi, Mitrović, and Granitzer at the University of Passau, presenting at CIKM '26, show that treatment is wrong. Their paper names the effect the Compaction Cliff. Testing Claude Code's /compact prompt on Sonnet 4.6 across 20 production agent configurations, they found that safety rules survive a single round of compaction only 53% of the time, and after five rounds, survival drops to a small fraction of the original rules. Structurally, a compactor applies one uniform summarization policy to everything in front of it, whether that's a hard safety constraint or a throwaway note about what happened three turns ago. A safety rule needs its exact wording preserved to stay enforceable. An episodic log entry doesn't. The compactor treats both the same way, and the rule loses.
The paper's example makes the stakes concrete rather than abstract. A medical agent's knowledge base contains one line: "Patient is allergic to penicillin." After compaction, that line gets paraphrased, or dropped outright, and three turns later the agent recommends amoxicillin, a penicillin-class antibiotic, to the same patient. Nothing exotic happened to cause that failure. Ordinary summarization, running exactly as designed, erased the one fact that existed to prevent that exact outcome.
The researchers' answer is called Knowledge Triage. Instead of one compaction policy for everything, each line in the knowledge base gets classified into one of five types: Constraint, Procedural, Belief, Preference, or Episodic. Each type then follows its own retention rule, enforced by three operators (TypeCompact, TypeDecompose, TypeRetrieve) built specifically to preserve exact wording for anything load-bearing. A verifier checks the result before it becomes the next working set and flags the output as unsafe if a constraint has gone missing.
How pure vector retrieval trades one problem for another
The obvious response to flat-file breakage is to abandon flat files and move to a vector database, and that move does solve context collapse: nothing gets truncated by a hard byte cap anymore, because retrieval works by similarity search instead of file size limits. But solving one problem surfaces the next: retrieval difficulty persists in a different form.
Vector search works by comparing meaning, not matching words, and that gives it both its strength and its failure point. A query about "the client's budget constraint" can miss a memory stored as "the project cap is fixed," even though the two phrases describe the same fact, because the embeddings for those two phrases don't land close enough together in vector space. Semantic similarity and factual relevance are not the same measurement, and a retrieval system built only on the first one will miss real answers that happen to use different words.
A 2026 guide from Vectorize names the tradeoff directly: combining semantic search with keyword matching, graph traversal, and temporal signals produces retrieval that holds up better than any single strategy on its own. But each strategy added to the stack adds its own complexity and its own latency cost. A system running four retrieval strategies in parallel has four places where something can slow down, four sets of results to merge, and four failure modes layered on top of whatever the original architecture already had. That's a real engineering cost, but it's a solvable one, closer to a tuning problem than a correctness problem.
What the research systems reveal about what works
Three research systems, each built independently, attack the retrieval bottleneck from different directions, and reading them side by side shows which architectural choices actually move performance rather than just adding complexity.
MEMTIER, from Ben-Gurion University, builds a three-part memory architecture for OpenClaw: a structured episodic store in JSONL format, a retrieval engine that weighs five separate signals, a feedback loop that updates those weights based on whether retrieved memories actually helped, and a background process that promotes episodic facts up into a longer-term semantic tier. On the full 500-question LongMemEval-S benchmark, MEMTIER reaches an accuracy of 0.382 and an F1 score of 0.412 using a lightweight model running on a single consumer GPU, a substantial jump over a no-retrieval baseline. With structured facts pre-populated ahead of time, single-session recall climbs to between 0.686 and 0.714, above the 0.560 baseline reported for RAG with BM25 and GPT-4o in the original LongMemEval paper, but the evaluation protocols differ enough that you should read this comparison as directional rather than a precise head-to-head.
MEMTIER's more interesting finding might be diagnostic rather than a benchmark score. The researchers tried using PPO, a reinforcement learning method, to let the system learn better retrieval weights automatically. It failed, because raw BM25 keyword scores dominated the signal so heavily that the learning process had nothing else to adjust against. Scaling the underlying generator model from 7 billion parameters up to a 284 billion-parameter mixture-of-experts model didn't fix this either. The retrieval architecture is the bottleneck, regardless of how large or capable the underlying model is.
LycheeMemory V2, from Harbin Institute of Technology in Shenzhen, takes a different angle: instead of consolidating memory turn by turn, it batches multiple exchanges into larger segments before encoding them. Under GPT-4.1-Mini, this produces strong scores on both the LoCoMo and LongMemEval-S benchmarks while cutting the token cost of building memory substantially compared to the earlier A-Mem system. The insight driving this result is that how finely memory gets chunked before it's stored (its consolidation granularity) shapes the accuracy-to-cost tradeoff as much as which retrieval strategy sits on top of it.
Memanto takes a third path, showing that tightly optimized semantic retrieval, combined with structured typing of memories and automatic conflict resolution, can match or beat hybrid systems that combine graphs and vectors, while cutting out ingestion overhead and reducing a retrieval operation down to one single query.
Two lessons repeat across all three systems. First, the expensive work (extracting facts, resolving which entity is which, generating embeddings) belongs at write time, running in the background, because memories get written once but read many times in situations where speed matters. Second, precision beats coverage: MEMTIER's own measurements show that roughly three LLM-extracted facts per question outperform roughly five hundred heuristically-extracted facts by nearly three times on F1. Flooding a retrieval system with undifferentiated content hurts results more than it helps them.
Production memory providers in practice
The research above answers which architectural choices work. The next question for anyone building a production system is which provider already implements them, and the honest answer is that no single provider leads on every tradeoff at once. Four providers show up repeatedly in production: Mem0, Zep, Letta, and Cloudflare's Agents primitive, and each one accepts certain tradeoffs from the sections above in exchange for avoiding others.
Zep, built on the Graphiti engine, takes the temporal-knowledge-graph approach. Every fact it stores carries both a valid time and a transaction time, and when a fact gets superseded, Zep marks it as outdated rather than deleting it outright, preserving the history of what was once believed true. Retrieval combines embeddings, BM25 keyword scoring, and graph traversal together, which makes Zep particularly strong on questions that are entity-centric, that hinge on timing, that require resolving contradictions, or that need multiple hops across related facts. That strength comes with real operational cost: self-hosting Zep means running an actual graph database such as Neo4j, FalkorDB, Kuzu, or Amazon Neptune, plus the schema design and extraction work that a graph requires. The managed Zep platform, as of pricing dated July 31, 2026, offers a free tier, paid Flex tiers billed annually, and a custom Enterprise tier that includes SOC 2 Type II and HIPAA BAA coverage for regulated deployments.
Cloudflare's Agents primitive, built on Durable Objects, takes a different position entirely: it's a stateful foundation to build a memory system on top of, a layer teams assemble into a memory product rather than a packaged one. That makes it a fit for teams who want full control over their own memory logic and are already working inside the Cloudflare ecosystem, accepting more build effort in exchange for more control over how the four failure modes above get addressed.
Choosing among these means matching a provider's position in the tradeoff space to the workload at hand, because picking the wrong fit doesn't just cost performance, it reintroduces the exact failure modes the earlier sections describe.
What multi-tenant and long-running production deployments require
Everything above assumes a single agent remembering things for a single ongoing session. Production systems rarely stay that simple. Once many users are each running their own long-lived agent against the same backend, two questions appear that flat-file and single-tenant designs were never built to answer: whose memory is this, and whose credentials is the agent acting on right now?
Memory isolation between users is a structural requirement of the architecture itself. Without strict separation, one user's stored memories can surface in another user's retrieval results, and worse, the four compounding failure modes described earlier don't just happen once: they happen inside every single user's session, independently and at the same time, multiplying the surface area for things to go wrong.
Credentials raise a parallel problem. The pattern used to handle this safely is brokered access: the agent passes along a user identifier, such as alice@example.com, rather than a token, and that identifier selects whose OAuth grant to use. A single function resolves the identifier to an actual token only at the moment a call needs to be made, and the token itself never touches model inputs, never appears in a tool schema, and never lands in a log. Composio's documentation for its LangChain integration describes exactly this pattern: each application user gets a stable user_id, each user connects their own accounts once, and the integration layer executes every toolkit action using the credentials tied to that specific user_id.
Production deployments running at this scale carry a set of security requirements that aren't optional extras: token rotation, least-privilege scopes limiting what each credential can actually do, role-based access control, audit trails covering every action taken, tenant isolation enforced at the architecture level, and guarded execution that checks actions before they run. Compliance-sensitive deployments add one further constraint: if data legally cannot leave an organization's own infrastructure, the only viable platform is one offering self-hosted or airgapped deployment. A managed platform without that option simply isn't eligible for that workload, no matter how well it performs everywhere else.
Sources
- The Compaction Cliff in Long-Running AI Agent Memory
- MEMTIER: Tiered Memory Architecture and the Retrieval Bottleneck
- LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation
- Best AI Agent Memory Systems in 2026: 8 Frameworks Compared
- Memanto: Typed Semantic Memory with Information-Theoretic Retrieval for Long-Horizon Agents


