APIs, integration & security — in depth

Conversation History Truncation Strategies for Multi-Session Agents

Keeping agents sharp means managing what they remember, not just how much.

Senior Writer · · 11 min read
Cover illustration for “Conversation History Truncation Strategies for Multi-Session Agents”
Agent State & Memory · October 4, 2026 · 11 min read · 2,530 words

Multi-session agents rarely fail because they run out of tokens. They fail because the history piling up behind them wears down how well they reason, long before any cutoff ever comes. That distinction changes everything about how to build them, and it's the subject of this piece: five distinct strategies for managing conversation history, each one a different bet about what to keep, what to compress, and what to throw away.

Why multi-session agents degrade before they run out of context

A language model has no memory of its own. Every single call starts from zero: the model reads whatever sits in the messages array that turn, and nothing else. No state persists between calls on its own. Think of the context window as RAM, not disk. RAM holds only what's loaded right now, and when the call ends, that's it. Whatever the developer chooses to pass in is the entire world the model can reason about for that turn.

That history grows in ways that are easy to predict individually but hard to manage together. The system prompt sits there as a fixed cost, but it can be large. Conversation turns pile up one after another as a session runs longer. Tool results are the wild card: a single file read or API call can return thousands of tokens in one shot, dwarfing an entire conversation that came before it. Retrieved context adds another layer on top. All four compete for the same fixed budget, and none of them shrink on their own.

The real damage happens before any limit gets hit. Models attend less precisely to the middle of a long context, so a session can sit well inside its window and still give worse answers, because the useful signal is buried under turns that don't matter. The usable context is smaller than the advertised number, and it shrinks further as a session goes on. One analysis of coding-assistant sessions found that the overlap between a query and its own prior history has a median of only around one-fifth. In other words, on a typical turn, roughly four-fifths of the history the model is carrying has nothing to do with the question being asked right now, and the model still has to sift through it.

The real question an agent builder faces is what to preserve, compress, or throw away so the model keeps reasoning well three hundred turns into a session. Every strategy that follows is a different answer to that question.

What context engineering means as an architectural discipline

Prompt engineering is about getting the wording of one instruction right. Context engineering is about managing everything the model sees across an entire session: which messages stay, which get dropped, how history gets compressed, how much budget each piece gets, and in what order retrieved material shows up. It's a different job at a different scale.

For a simple chatbot answering one question at a time, none of this matters much. There's no history to manage because there's barely any history. The moment an agent starts running tools, accumulating results, making decisions based on earlier decisions, and carrying a session across dozens or hundreds of turns, context engineering becomes the central engineering problem.

A typical multi-turn agent is juggling several things that all draw from the same pool: the system prompt, which is fixed but can be sizable; the conversation history, which grows with every turn; tool results, which spike without warning; retrieved context, injected fresh each turn; and the current user input itself. Each of these grows at a different rate and for different reasons, and nobody manages this competition by accident. It has to be designed.

There's a structural choice underneath all of this that shapes which strategies are even available. When the service hosting the model manages history automatically, compaction happens without the developer having to think about it, but the developer also loses the ability to steer how it happens. When the client manages history instead, the developer has to handle compaction and decide what survives and what doesn't. Every strategy covered here runs into that same trade-off between convenience and control, so you pick the right one by knowing which side of that line your system sits on.

Sliding-window truncation and what it sacrifices

Sliding-window truncation keeps only the most recent N messages, so anything older than that just falls away. It makes one clean bet: recent context is almost always more useful than old context, and for a lot of agents, that bet is correct often enough to justify the simplicity.

This is the cheapest option by a wide margin. No extra model call to summarize anything, no retrieval infrastructure to build and maintain, no latency added for managing history. It's a fixed window, applied mechanically, turn after turn.

It fits support bots, CLI assistants, and short-lived task agents well, because a user's intent resets often in these systems, so the first few exchanges of a session stop mattering once it moves on.

Each dropped turn costs the agent permanently, not just on occasion. Once a turn drops out of the window, it's gone. Variable definitions, branch decisions, commitments made earlier in the conversation, all of it disappears permanently, and any causal chain running through that turn breaks with it. The Pull paper on lazy materialization names this directly: hard truncation and sliding windows work by irreversibly discarding history, and that's the defining limitation of this whole layer of session management. It handles the current query fine. It cannot recover a turn once that turn is gone.

Picture a long advisory or research session where a decision made in turn 3 constrains what's valid by turn 47. A sliding window will drop that early turn without any signal that it mattered, and the agent will keep reasoning as though the constraint never existed. That failure case is why the next four strategies exist.

Conversation summarization: trading specificity for thread continuity

Summarization makes a different bet than truncation does. Instead of assuming old turns are worthless, it assumes the thread of reasoning across a session affects coherence more than the exact wording of any one exchange, and that a generated summary can hold onto enough of that thread to keep the agent coherent. Once context hits a set threshold, the system swaps the older history for a summary block, but it leaves the most recent turns intact word for word. The prior conversation survives, just in compressed form, rather than vanishing the way it does under truncation.

Claude Code runs this exact pattern in production. It generates a human-readable summary once context crosses its compaction threshold, it supports a manual /compact [instructions] command so you can steer what gets kept, and it reloads the project-root CLAUDE.md file after compaction completes. Nested or path-scoped CLAUDE.md files don't get reloaded automatically, though, so they stay lost until something in the session reads that file again. Anthropic's compaction API, under the beta header compact-2026-01-12, is available across the Claude API, AWS Bedrock, and Microsoft Foundry, turning this from something each team builds by hand into a standard API feature.

The honest cost here: summarization is an irreversible merge. Specific variable states and exact branch decisions can disappear for good, along with anything that didn't look important to whatever did the summarizing. The summary makes the input smaller and more digestible, but it can also quietly erase the very information the agent needs three turns later.

Research on multi-agent collaboration found that having a coordinating agent curate what gets summarized, rather than handing the job to a generic summarizer, cut token costs substantially while holding accuracy steady or improving it. The policy governing what gets summarized matters just as much as the act of summarizing itself.

This approach fits long advisory, research, or coaching sessions, where preserving the thread of reasoning keeps the agent coherent better than preserving the precise wording of any one exchange, and where a human operator can give compaction instructions to guide what gets kept.

Selective tool-output pruning: handling the spikes that summarization misses

Tool outputs don't grow the way conversation turns do; they spike. A single file read, search query, or API response can return more tokens in one call than an entire conversation accumulated over dozens of turns, so tool outputs need their own handling rather than getting folded into general summarization.

Selective pruning extracts or truncates a tool result down to the relevant part before it reaches the model, but then you pay for an extra call just to decide what counts as relevant. That cost is usually far smaller than the cost of letting the model reason over a context bloated with raw tool output it doesn't need.

Claude Code solves this at the storage layer instead of through summarization. When a tool returns something large, the system saves the full output to a file and passes the model only a truncated preview plus the file path. You can still recover the complete result on disk, but it never floods the active context window.

Deciding what to prune is a judgment call every time. Prune too aggressively and the model loses details it actually needed. Prune too conservatively and the whole exercise is pointless. So you pair this technique with explicit budget allocation rather than use it as a standalone fix.

It fits agents built on verbose tools: search, file reads, large API calls, anywhere the raw output routinely exceeds what the model actually needs to reason well, and where the tool output feeds the session without being the main subject of it.

Token budget management: treating context as an explicit resource with eviction policy

Everything covered so far reacts to a problem after it shows up: history gets too long, so it gets truncated; context fills up, so it gets summarized; a tool result is too big, so it gets pruned. Token budget management flips that order. It allocates a set share of the context window to each component ahead of time and enforces that allocation before anything overflows.

In practice, the budget typically adjusts dynamically rather than staying fixed, and it triggers compression the moment any one component starts to overrun its share. The system prompt gets its slice, conversation history gets its slice, tool results get theirs, retrieved context gets theirs, and the whole thing is managed as a single resource with rules, not an open-ended pile that gets cleaned up whenever it gets too big.

TokenPilot treats cache-efficient context management for LLM agents as a formal engineering problem in its own right, and that shows where the field has moved. Token budgets are a managed resource with explicit policy, the same way a memory cache gets managed in any other system with finite RAM and a need to decide what stays resident and what gets evicted.

The agents that stay reliable over long sessions tend not to be the ones running on the largest context windows available. They treat context as a budget with rules, the same discipline any engineer would apply to managing a memory cache under pressure.

This approach fits any agent meant to run long sessions in production, especially where the system prompt, history, tool results, and retrieved context are all drawing from the same budget at the same time and none of them can be allowed to crowd out the others.

Retrieval-augmented context injection: selective recall instead of full-history management

Retrieval takes a completely different position than the first four strategies. It doesn't truncate the history, compress it, or budget it. It leaves the history where it is and pulls out only the pieces relevant to the current query, using semantic search to decide what counts as relevant. This approach bets that most of a long conversation has nothing to do with any single question asked later, and that a good retrieval step can reliably find the small part that does.

This approach needs real infrastructure that the other four strategies don't require. Redis lists, for example, offer a fast, ordered way to store conversation turns with built-in expiry and atomic operations, and context window management functions built on top can enforce token limits while favoring the most recent exchanges. Retrieval is a storage decision in its own right, not just a semantic search call bolted onto an existing pipeline.

Research on MultiSessionCollab showed that if agents carry memory built specifically to learn user preferences across sessions, they produce higher task success rates, more efficient interactions, and less effort required from the user. Retrieval pays off when the memory system is actually built around what the agent needs to recall.

Retrieval carries its own architectural limit, though, and it's shared across every system built this way: retrieved text comes back as a fragment separated from the temporal context it originally sat in. A memory system built purely on retrieval can lose track of branch topology and the lifecycle of entities across a conversation, and a fragment pulled without its causal context can mislead the model rather than help it.

The Pull paper's lazy materialization is a direct response to that limitation. Rather than retrieving dead text fragments, it keeps a live metadata directory tracking entity lifecycles, branch topology, and information density, and the model pulls only the specific turns it actually needs at query time. Because nothing gets merged or rewritten the way it does under summarization, any collapsed turn can be expanded again later. That reversibility is the key difference between this approach and summarization's one-way compression.

Retrieval fits agents that carry long cross-session histories where most prior content has nothing to do with most new queries, and where the retrieval mechanism is built to preserve causal structure, not just surface-level similarity.

How production agents combine these strategies

No serious production system deploys any of these five strategies alone. Sliding-window truncation handles the baseline case of recent-turn relevance. Once a session's history crosses a length threshold, summarization kicks in and preserves the reasoning thread while it compresses the specifics. Tool-output pruning runs on its own and catches the spikes that neither truncation nor summarization handles well, because a single large file read just doesn't behave like a normal conversation turn. Token budget management sits above all three: it allocates space to each component and decides which compression mechanism fires and when. Retrieval adds a fifth layer on top, for the cases where a session's relevant history lives further back than a normal window or summary would reach, and where pulling the specific fragment preserves what compressing everything that came before it would lose.

The choice isn't which single strategy to adopt. It's which combination fits the shape of a given agent: how long its sessions run, how verbose its tools are, how much specific state its reasoning depends on, and how much infrastructure a team is willing to build and maintain. If a support bot only answers short, self-contained questions, it has little use for retrieval infrastructure or careful budget tuning. A long-running research or coding agent, carrying state across hundreds of turns and multiple sessions, needs most of these working together, with clear rules for what gets kept, what gets compressed, and what gets thrown away for good.

Sources

  1. Beyond Frameworks: Unpacking Collaboration Strategies in Multi-Agent Systems
  2. Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations
  3. Chat History Storage Patterns in Microsoft Agent Framework
  4. MultiSessionCollab: Learning User Preferences with Memory to Improve Long-Term Collaboration
  5. TokenPilot: Cache-Efficient Context Management for LLM Agents

More in Agent State & Memory