APIs, integration & security — in depth

Per-User Agent Memory Isolation in Multi-Tenant Systems

Storage-layer isolation prevents silent data leaks that filters alone cannot catch.

Senior Contributor · · 10 min read
Cover illustration for “Per-User Agent Memory Isolation in Multi-Tenant Systems”
Agent State & Memory · October 6, 2026 · 10 min read · 2,251 words

A support SaaS company scoped its agent memory to a single shared bank, tagging each record with a customer_id. Hindsight's analysis treats that event as a breach, not a personalization bug, and that framing is not a rhetorical choice. It determines who gets called, what gets disclosed, and what actually needs to be rebuilt. A stricter filter would not have stopped this. Only a real storage boundary would have.

Agent memory has two jobs. It needs to recall the right thing, and it needs to never recall someone else's thing. Most teams build for the first job and treat the second as a filter to bolt on afterward, typically a user_id column added once multi-tenancy becomes unavoidable. When those two goals collide, isolation has to win, and that means the system has to be designed around the second job from the start, not patched toward it later. If a support agent surfaces a slightly wrong document, that is a quality problem. A support agent that surfaces the right document, just belonging to the wrong customer, is an incident.

This is not a narrow or isolated failure. Crew Scaler's security study reviewed sixteen security frameworks built for multi-agent systems and found that Data Leakage, alongside Non-Determinism, is one of the two domains those frameworks address the least. Sixteen frameworks, and leakage through shared memory is still a gap across most of them. That is a systemic weakness in how the field has approached agent memory, not an edge case any one team happened to miss.

The four ways shared memory leaks across tenants

Shared memory in a multi-tenant agent system fails in four distinct ways, and each one is harder to catch than the last. Standard unit tests, the kind that check whether a function returns the expected value for a given input, catch none of them reliably, because all four failures depend on cross-tenant state that a single-tenant test case never constructs.

Direct leakage is the easiest to picture and the easiest to catch: the agent returns a snippet, a summary, or a document fragment that belongs to a different tenant's conversation, and once someone sees it, the problem is obvious. That makes it the least dangerous of the four, not the most, because visibility is what gets it fixed.

Behavioral contamination is harder to catch, because nothing cross-tenant appears in the output. A reviewer checking outputs for leaked text will find nothing, because the leak lives in the decision path, not the words on the page.

Memory-based prompt injection is the version most teams have not priced in. A memory system treats stored memories as trusted context, not as untrusted input. If one tenant writes a careless or malicious instruction into shared long-term memory, say an instruction to export full logs, that instruction will not stay quarantined. Crew Scaler's taxonomy names this specifically: Memory Poisoning is catalogued as its own risk category, and latent poisoning through shared vector databases is flagged as an attack pattern that almost no framework currently reviewed has a real countermeasure for.

Retrieval contamination happens at the embedding layer itself. Vector search is probabilistic by design. Similarity collisions mean a semantically close item belonging to a different tenant can get pulled into a result set that was never supposed to include it. Multi-tenancy needs a hard, deterministic boundary at retrieval time, and vector search alone does not offer one.

Each of these four failures traces back to the same root decision: memory lives in one shared store, with tenant separation enforced by a filter applied at query time instead of being built into the structure of the store itself. That single choice is what the next section exists to unpack.

The single architectural decision: hard isolation boundary versus soft partition

Diagram: Hard Isolation vs. Soft Partition: Two Very Different Failure Modes. Visualizes: Show a side-by-side comparison of two architectural approaches to multi-tenant agent memory.

Nearly every guide to multi-tenant agent memory reduces to the same advice: put a user_id on each record and filter by it. That is a soft partition, not isolation, and the gap between the two is a difference in which layer of the system actually enforces the boundary.

Hindsight's analysis draws the line precisely:

| | Hard isolation boundary | Soft partition | |---|---|---| | What it is | A separate store per tenant or user | A label or filter inside one shared store | | What enforces it | The storage layer itself | The query, at call time | | Failure mode | Operational: cost, complexity | Silent data exposure |

The two failure modes are not symmetric. Hard isolation's failure mode is operational: it costs more to run, and it takes more engineering to manage entitlements across separate stores. Soft partitioning's failure mode is silent: the system looks correct, keeps serving requests, and passes every test, right up until the one query that forgot the filter runs in production. Soft partitioning's failure surfaces in a breach notification.

The tell is in how these systems default. In a soft-partition system, tag filtering by default includes untagged memories, because tags were designed as organizing hints for retrieval, not as walls for security. That default is correct for personalization and wrong for isolation, and most teams never notice the mismatch until an untagged record walks straight through it, exactly as it did in the support SaaS case.

The rule that follows is simple to state: a filter a developer can forget to pass is not isolation. If a product is per-user or per-tenant, the user or tenant itself needs a hard boundary. Everything else, categories of memory, topics, projects, can live behind tags inside that boundary. The same logic extends to how memory gets classified in the first place. Ephemeral context, user preferences, org-wide policies, derived summaries, and tool state all carry different isolation needs, but the test is the same for each: if an item must never surface in the wrong context under any circumstance, it belongs behind a storage-layer boundary. If it is merely organized for easier retrieval, a tag is fine.

Isolation and Partitioning in the Silo and Pool Models

Two implementation patterns carry out that decision in practice: the silo model, which implements hard isolation, and the pool model, which implements a soft partition. Neither one is categorically better. It comes down to which threat model a given system actually needs to satisfy.

The silo model gives each tenant a dedicated memory store, a separate bank, collection, or namespace that the storage layer treats as a genuinely distinct object, not a label inside a shared one. There is no built-in query capable of spanning banks. The tradeoff is cost and complexity, managing more stores and more entitlement logic, an operational and economic question.

The pool model keeps shared infrastructure and enforces separation through a hierarchical, namespace-based scheme: all tenant data is stored in one common store, with strict filtering based on composite identifiers combining a tenant identifier and a user identifier. Authentication carries tenant context as claims or tags on scoped credentials, and every memory operation gets prefixed with the right namespace path before it runs. Amazon Bedrock AgentCore Memory implements this pattern with hierarchical namespace isolation across four levels, global, strategy, actor, and session, with tenant and user boundaries enforced through composite actorId conventions and IAM policies. The pool model fits slices of data a system sometimes wants to cross-reference on purpose: shared projects, topics, or org-wide knowledge. It is not the right tool for the tenant or user boundary itself.

Production systems typically combine both, using three scopes. Memory private to a single user gets a hard boundary, something like bank_id="user:{id}". Memory shared within an organization or workspace gets its own hard boundary at the org level, bank_id="org:{id}". Global product knowledge, meant to be shared and cross-referenced by design, lives in a shared, mostly read-only store where a soft partition is the appropriate and intentional choice. Make the bank the thing that must never leak, and use tags only for the kinds of memory that live safely inside that boundary.

A concrete implementation of this logic appears in CrewAI's pull request #5967, closed without being merged into main, which threaded tenant_id and an optional user_id through the entire save, store, and recall path. Isolation was enforced at the storage layer itself through a ScopedStorage proxy built on three contracts. First, a stamp-on-write rule: records carrying the default tenant_id get overwritten with the actual bound tenant, while records arriving with a different, non-default tenant_id raise a PermissionError. The three contracts work as layered defense: a type-check catches a forgotten keyword argument, the backend pushes the tenant predicate directly into the vector query, and ScopedStorage stands as the runtime guard if either of the first two layers fails.

Compute sandboxing as the second enforcement layer beyond memory isolation

Storage-layer isolation protects memory, but code the agent generates and executes at runtime remains unprotected. Once an agent starts generating and executing code at runtime, against a tenant's data, the compute environment running that code becomes a second boundary, and it needs enforcement at the hardware or kernel level, not just at the level of a process.

Blaxel's analysis describes the mechanism behind this shift. Agent workloads flip the traditional SaaS trust model on its head: the vendor no longer controls what code actually runs. The agent generates that code from the customer's prompt, at runtime, and then executes it. That makes every tenant's workload a live, potential threat to every other tenant sharing the same execution environment. OWASP's guidance for LLM applications states the implication directly: treat the model as any other user, and validate everything it produces, because code the platform never reviewed is untrusted by definition, no different in principle than input typed by an unknown visitor.

This is the same structural argument as the memory section, applied one layer down. A query-time filter on a shared memory store is not a storage boundary. A namespace on a shared container is not a kernel boundary. Compute sandboxing needs to defend against five concrete threats: filesystem read or write access outside the intended scope, network egress to arbitrary domains, exposure of the host kernel's syscall surface to code the agent wrote, cross-tenant data leakage once isolation stops at the process level, and secret exfiltration through environment variables or the /proc filesystem.

Three isolation technologies map to three different threat levels. Firecracker microVMs give each tenant a dedicated guest kernel, enforced through hardware virtualization, so a kernel exploit inside one microVM cannot reach the host or any neighboring tenant. V8 Isolates, built around JavaScript, TypeScript, and WebAssembly (which also lets them run Python and Rust), suit latency-critical, lightweight tasks with a bounded workload and a lower threat profile.

Containers sit below all three in isolation strength, because they share the host kernel and rely on namespace isolation rather than hardware boundaries. The CNCF's analysis of the runc breakouts from November 2025 warned that these flaws pose a critical risk, especially in multi-tenant environments where users define their own containers or run images nobody vetted. AWS reached the same conclusion architecturally: Bedrock AgentCore runs each user session in a dedicated microVM, with isolated CPU, memory, and filesystem, the same hardware-level pattern Firecracker implements.

Credentials need their own isolation treatment on top of all this, because environment variables are readable from inside a sandbox, and even retrieving secrets from a vault at runtime still places them in the sandbox's memory at some point. The sounder pattern is proxy-based credential injection: a proxy applies tenant-scoped credentials at the network layer, outside the agent's own process, so the agent never holds the raw credential.

None of this requires picking one tier for an entire platform. Running customer-facing agents that execute arbitrary code inside microVMs, while keeping internal background jobs that run platform-written code inside containers, is a legitimate production pattern. The isolation tier matches the trust level of the workload it carries, not a single setting applied uniformly everywhere.

Diagram: Three Isolation Tiers for Agent Compute Sandboxing. Visualizes: Show a ranked hierarchy of three compute isolation technologies mapped to their threat levels, from strongest to weakest.

Requirements for holding this boundary in a trusted cloud environment

Storage-layer memory isolation and kernel-level compute isolation establish the technical boundary, but neither one accounts for every party with access to the environment they run in. A cloud deployment is a multi-party system: the cloud provider, the orchestration layer managing the agents, and other tenants sharing the same underlying infrastructure are all parties whose access has to be bounded for the isolation guarantee to mean anything end to end.

Research from the Technical University of Munich, with one co-author from INESC-ID, frames this as the core problem facing AI agents deployed as cloud services. Agents increasingly run inside a complex, multi-party ecosystem, and untrusted components anywhere in that chain can lead to data leakage, tampering, or behavior nobody intended. A memory store can be perfectly isolated by tenant and still sit inside an orchestration layer that was never designed to prove, to an auditor or to the tenants themselves, that the isolation held at every step.

This is what separates a technically correct architecture from one that a regulated customer can actually trust. Isolating memory by bank and sandboxing compute by microVM answers the question of whether a leak can happen. A trusted cloud environment also has to answer who else could have seen the data in transit, who orchestrated the call, and what guarantee exists that no intermediate party had the access needed to tamper with it. Building the storage boundary and the compute boundary is the foundation. Holding that boundary in a shared cloud, across every party touching the system, is the standard a genuinely multi-tenant, security-first agent architecture has to be built to meet.

Sources

  1. Security Considerations for Multi-agent Systems
  2. The Hitchhiker's Guide to Agentic AI: From Foundations to Systems
  3. Trusted AI Agents in the Cloud
  4. Building multi-tenant agents with Amazon Bedrock AgentCore

More in Agent State & Memory