State Persistence Patterns Across Agent Process Restarts
Different types of agent state need different storage to survive process restarts.

An agent that restarts from zero is suffering from one storage strategy applied to four kinds of state that don't behave alike. Conversation history, task progress, tool results, and working files each carry a different cost when lost, and each needs a different place to live. Treating them as one blob makes the agent either spend money saving things it didn't need to save, or lose things it desperately needed back.
Why agent state is several distinct problems, not one
Picture an agent mid-task when the container it's running in gets rescheduled. The process dies. On restart, the agent has no idea what it was doing thirty seconds earlier. That's not a bug specific to one bad deployment. Large language models carry no memory between API calls by default. Every single call starts from nothing unless something outside the model wrote down what happened and handed it back in time for the next call.
Production is not a gentle place to run a long task. Processes stop for reasons that have nothing to do with the agent's logic: an unhandled exception, an out-of-memory kill, a routine deploy, a container getting rescheduled onto a different host, a reboot. Any of these can happen at any moment, and an agent built without that assumption baked in will, sooner or later, forget everything it was in the middle of doing.
The instinct to fix this is usually to "add persistence," as if state were a single thing you either save or don't. That instinct causes its own damage. Lump every kind of agent state into one bucket, and the result lands on one of two bad outcomes: either everything gets written to durable storage on every step, driving up cost and latency for data nobody will ever need to read back, or nothing gets persisted at all because the team decided persistence was too expensive to bother with, and the agent silently restarts from scratch every time something interrupts it. The fix is matching each kind of state to the storage it actually needs, starting from the admission that agent state comes in more than one shape.
Conversation memory: the state type engineers get wrong first
Conversation memory is usually the first thing a team builds, and the first thing that breaks. The most common implementation is a Python dictionary or a JSON blob held in RAM. It's also the weakest: anything stored only in the process's own memory disappears the instant that process exits, whether from a crash, a routine restart, or a deploy. There's no warning before it happens and no way to get the data back after.
The fix starts with a mental shift about what conversation memory actually is. The agent's message log, the full append-only record of what was said and done, is the memory, not a side effect that happens to sit near it. Persist that log and recovery, session resumption, and later debugging all become possible from the same artifact.
Beyond the raw log, conversation memory splits into four tiers, each doing a different job at a different cost. Working memory is whatever the model can currently see in its context window: it's useful but volatile by design, and nobody should expect it to survive a restart. External key-value memory is a separate, persistent store, something like Redis or DynamoDB for key-value lookups, or PostgreSQL if a relational structure fits better, holding structured facts such as user preferences, entity attributes, and configuration that need to survive restarts without the agent noticing anything happened. Episodic logs are the append-only record of events over time, the foundation that crash recovery and audits are built on. Semantic vector memory is a retrieval index built from documents, so the agent isn't re-embedding the same material every time it boots up.
The twelve-factor app principle applies here with almost no translation needed: treat the running process as something that could disappear at any moment, and push anything that needs to outlive it into a separate backing service. Conversation memory is exactly the kind of state that has no business living only inside the process. A context window filling up over time is a related but separate problem from a crash: history will eventually outgrow what the model can see at once, and that has to be handled with a deliberate trimming or rollover plan built into the persistence design from the start, not bolted on after the first time it breaks something.
Task checkpoints: where the cost of getting it wrong compounds
Losing a conversation is annoying. Losing a task checkpoint is expensive. Anyone who has watched a three-hour research job or a multi-step workflow restart completely from the beginning after a timeout knows the difference in the gut, not just on paper. A workflow that touches several systems, waits on a slow external API, or runs across days needs somewhere durable to record exactly where it left off. Without that, a timeout or an unplanned restart doesn't just interrupt the task. It erases the hours already spent, and the tokens already paid for, forcing the whole thing to run again from step one.
It helps to separate the conversation from what might be called the case file. The conversation is simply what was said. The case file is what's still true once the talking stops: which steps finished, what the agent is waiting on, what it already tried and ruled out. That case file is the task checkpoint, and it needs a durability guarantee stricter than conversation memory gets, because the cost of losing it scales with how long the task has been running.
Checkpoints only meet that bar when they live in a durable, shared datastore that replicates across more than one failure domain. Local, in-memory storage fails the same test here that it failed for conversation memory, only with a steeper price attached. LangGraph's checkpointing approach uses a PostgresSaver or a RedisSaver backed by PostgreSQL or Redis to persist the agent's state at each step, then attaches that saved checkpoint to the OpenTelemetry trace as a span attribute. Done that way, an incident investigation later has the full decision tree in hand and can replay any past state exactly as it existed.
Four distinct failure scenarios call for four different checkpoint behaviors. A planned restart, like a deploy or scheduled maintenance, gives the agent advance notice, so checkpoints can be flushed cleanly before the process exits. A crash from an out-of-memory kill or an unhandled exception gives no warning and no chance to clean up. Checkpoints have to be written continuously as the task runs, not saved only once at the end. Context overflow, where the conversation history grows past what the model can hold, requires a checkpoint that carries enough task state to resume without replaying the entire conversation from scratch. A hang or timeout, where the process is alive but stuck, needs a checkpoint that lets a watchdog process restart the agent from the last known good state rather than waiting indefinitely.
The practical rule that follows from all four scenarios: write state at sensible points as the task runs, after each completed step, after each finished tool call, or on a fixed timer, rather than only at the very end. Just as important, and the step most implementations skip, is reading that state back in on startup. A checkpoint nobody loads on boot provides no more value than having no checkpoint.
Tool outputs and working files: the state types that don't fit a database
Tool outputs and working files behave differently from conversation logs or task checkpoints, and storing them the same way wastes money or breaks retrieval. Tool outputs are the results of API calls, searches, and computations the agent has already paid for once. Losing them means paying for the same expensive or rate-limited call again, which adds direct cost and slows every later run down.
Working files are often the biggest share of an agent's state and the most painful to lose mid-task: scratch files, downloaded data, generated artifacts, a cloned repository sitting under active edit. Container filesystems are ephemeral on purpose. Rebuilding or rescheduling a container means whatever was sitting in its local filesystem is simply gone. That's a known property of how containers work, not a defect to patch around, and the persistence design has to assume it from day one.
Three storage options map cleanly onto these needs, and none of them is a database. Mounted volumes, a disk that lives outside the process and attaches at runtime, are the simplest fit for working files, checkpoints, and small local databases: close to the old assumption of "save a file and trust it's still there next time," made true again. Object storage, services like S3, Cloudflare R2, or Google Cloud Storage, fits large or infrequently touched artifacts: generated files, backups of a vector index, large downloads. It's cheap per gigabyte and scales without a practical ceiling. Databases remain the right choice for structured state that needs concurrent access, but they're the wrong tool for large binary files or scratch data that gets rewritten constantly. And once more than one agent touches the same working file, the stakes rise further: two agents editing shared state without isolation and a consistent view of that state don't collaborate, they collide, producing race conditions instead of coordinated work.
What rollback restores, and what it doesn't
Persisting state correctly is necessary, but it isn't the end of the story. A checkpoint can be restored with perfect fidelity and still hand the agent a situation that never existed in any valid sequence of events. A systematic security study of checkpoint and rollback across five representative agent frameworks found five recurring failure modes: incomplete internal state coverage, inconsistent checkpoint state, external state mismatch, unbound nondeterministic replay, and unrecorded external effects.
A coding-agent scenario makes the problem concrete. Suppose an agent removes malware from a repository, scans the cleaned result, and records that the scan passed. If recovery later restores the workspace to its state before the cleanup, while keeping the agent's internal memory of the post-scan "passed" result, the agent can end up releasing the original malicious repository, believing it to be verified clean. Nothing here was forged, and the checkpoint and the scan result both matched what had actually happened at the time. Rollback simply reattached a genuine security judgment to a file it never actually checked.
This isn't theoretical. End-to-end attacks demonstrated on Hermes, Cline, and LangGraph produced a malware-verification bypass, unauthorized mail forwarding, and a double payment, all through rollback recovery failures rather than through prompt injection. The common thread across all three is what researchers describe as a break in execution continuity: the security-relevant states, decisions, and effects carried across a recovery have to stay consistent with some single valid history, and current systems don't guarantee that.
Part of the reason is structural. Three separate recovery boundaries exist in agent systems today: framework state, workspace state, and OS or VM state. These don't necessarily share the same recovery scope. A checkpoint that restores framework state cleanly can leave the workspace or the underlying machine stale, and the agent has no built-in way to notice the mismatch. The design lesson that follows is specific: persistence alone doesn't guarantee safe recovery. Anything that can't be rolled back, a sent email, a processed payment, a deleted file, has to be tracked explicitly as a committed, unrecoverable side effect, not left as an assumption buried inside the checkpoint.
Per-user isolation as the production constraint that reshapes every pattern
Every pattern described so far changes shape the moment an agent serves more than one user. Conversation memory, checkpoints, tool outputs, and working files all need to be isolated per user. Shared state produces two different kinds of failure at once: a correctness failure, where one user's context leaks into another user's session, and a security failure, where one user's credentials or decisions become reachable from someone else's conversation.
Per-user sandboxing is the baseline architecture for any agent-based product with more than one customer using it. In multi-agent systems, shared state without proper isolation produces race conditions and inconsistent views of task progress, and those failures trace back to a fragmented state design, not to any weakness in the underlying model.
Credential storage makes the stakes concrete. A user who connects their own accounts to an agent needs those credentials to survive restarts automatically, without that user's tokens ever becoming visible or usable in another user's session. The operational consequence follows directly: the repeatable unit of scale for an agent-based business is one isolated agent instance per customer, provisioned, secured, monitored, updated, and recovered on its own, rather than one shared instance serving everyone at once. That same logic extends to white-label products, where an operator or SaaS company activates a private, persistent, branded agent for each customer automatically at onboarding, with no hand-configuration required per account.
Matching state type to storage strategy: a working decision framework
Once the four categories are separated out, the storage decision for each one follows from two simple questions: how badly does it need to survive, and how is it going to get read back.
Conversation memory belongs in an append-only log, held in a relational or key-value store, something that survives restarts and gets read back before the agent takes any new action. Task checkpoints need to be written after every completed step rather than saved only at the end, backed by PostgreSQL or Redis, and replicated across failure domains for anything that runs long. Tool outputs can live in a database when they're small and need to be queried, but belong in object storage once they're large or rarely touched, and should never depend on an in-process cache as the only copy. Working files and artifacts belong on mounted volumes while they're in active use, and move to object storage once they're finished and need to outlive the session, never left sitting in a container's own ephemeral filesystem.
The step most implementations skip is the read-back. Writing state carefully means nothing if the agent never loads it back in on startup. The first thing an agent should do on boot is check whether unfinished work or existing memory is waiting for it, before taking any new action.
The rollback issue from earlier belongs in this framework too, not off to the side as a separate security concern. Any external effect that can't be undone, an email sent, a payment processed, needs to be recorded as a committed fact the moment it happens, not left as something the checkpoint merely implies. The boundary of what a checkpoint covers has to be drawn wide enough to include every dependency the restored state actually relies on. Get that boundary right, and an agent that restarts stops losing its place. It picks up exactly where it left off, because the state it needs was stored exactly where that kind of state belongs.
Sources
- Safe to Resume? Breaking Execution Continuity of Agent Execution via Rollback
- Agent Team Work Zone: An Automated, Persistent Workspace for Long-Lived Claude Code Agent Teams
- Hierarchical Memory Orchestration for Personalized Persistent Agents
- Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory
- Agent memory and state management - Agentic AI Lens
- Always-OnAgents:A Survey of Persistent Memory, State, and Governance in LLMAgents
- AgentRewind: Recoverable Execution for Long-Horizon LLM Agents


