Orchestrator-Subagent Trust Boundaries in Hierarchical Agent Systems
Orchestrators need explicit trust boundaries to prevent subagents from exploiting delegation gaps.

In a hierarchical agent system, one agent holds the full picture and the authority to act on it, while the agents underneath it hold a narrow slice of capability and a short leash. The orchestrator carries the complete context: the user's goal, the history of the task, the judgment calls about what matters next. A subagent gets a slice of that: one job, a few tools, and a small window to report back, and that split is a trust decision, made the moment the orchestrator decides to delegate instead of handle something itself.
By 2026, the industry had mostly settled on one shape for this: a single coordinator agent holds the full context and spawns subagents that return compressed summaries, keeping the orchestrator's context window focused on reasoning. Decentralized peer-to-peer coordination, where agents negotiate with each other as equals, still exists as an alternative. Part of why the industry settled on hierarchy is that a hierarchy makes trust relationships legible: you can point to the orchestrator and ask what it knew, and point to a subagent and ask what it was allowed to touch.
The asymmetry between orchestrator and subagent is deliberate, and it works exactly as well as the team designing it defines it. An orchestrator with broad context and a subagent with narrow scope is a sound design when someone decided, on purpose, what "narrow" means for that subagent: which tools, which data, which actions. It becomes a liability when that scope is never written down, when the subagent ends up with whatever access was convenient to grant. Most production failures in multi-agent systems trace back to that gap: not a flaw in the orchestration logic, but a trust boundary that nobody drew a line around.
What the orchestration layer controls
Orchestration frameworks are good at coordination and state. They tell the orchestrator how to break a goal into subtasks, how to pass data between steps, how to retry when something fails. They have no opinion on who should be allowed to do what. That gap, between managing the mechanics of delegation and governing the authority behind it, is where trust boundary failures start.
A well-built orchestration system handles four things at runtime. It carries state across steps, so what happened in step one shapes step five. A wrong fact written into that shared state early can quietly corrupt everything built on top of it. And it defines error handling, retries, and escalation paths in advance, because working these out for the first time during a live incident is how small failures become large ones.
The frameworks built to do this in 2026 differ mainly in how much structure they impose. Microsoft's Agent Framework, which folded AutoGen and the Semantic Kernel SDK into one SDK, and Google Cloud's Agent Development Kit, with its hierarchical agent trees and support for the A2A protocol, round out the frameworks most teams are choosing from.
None of these answer a different set of questions: who governs what an agent can access, what these agents cost to run in production, and how access rules get enforced the same way across every agent a company deploys. Those questions sit above the framework, in the infrastructure and governance layer that has to exist regardless of which coordination engine a team picks. Orchestration frameworks solve coordination well, but governance, access control, cost tracking, and uniform enforcement across deployed agents live one layer up. Platforms like Agent37 operate at that layer, handling the questions frameworks leave open: who gets access to what, what a multi-agent system costs at scale, and how those controls hold across every isolated agent instance a company runs. Picking a framework is a separate decision from designing a trust model, and teams that treat the first choice as a substitute for the second tend to find out the difference in production, usually at the worst time.
The three attack surfaces that compound at scale
Hierarchical agent systems open up three attack surfaces that a single agent working alone never has to deal with, and all three trace back to delegation happening without explicit trust limits attached to it.
The first is prompt injection across agent boundaries. A subagent pulls in information from outside the system's trust scope, a retrieved document, a tool's output, a response from another agent, and that information can carry instructions nobody authorized. The pattern plays out in a fixed sequence: a worker agent retrieves a document containing adversarial text, the worker's response folds that text in, and the orchestrator treats the result as legitimate output because nothing in the pipeline flagged it as foreign. This sits among the top agentic AI security threats tracked in late 2026.
The second is cross-agent privilege escalation, the clearest way to see how damage spreads through a hierarchy. Picture a pipeline where a research agent has web access but no filesystem access, and a code agent has filesystem access but no network access, each one scoped sensibly on its own. An attacker doesn't need to breach either agent directly. Neither agent was individually over-permissioned. The failure is in the hierarchy itself, in the fact that a message crossing from one agent to another carried no marker of how much it should be trusted. This is also named among the critical threats tracked for late 2026.
The third is over-privileged inheritance. In an orchestrator-subagent setup, a subagent can end up with more access than the human who kicked off the workflow in the first place, because the natural-language instruction that started the task never got translated into a matching set of limits on what the system could do in response. OAuth 2.0 Token Exchange, defined in RFC 8693, addresses part of this by injecting identity that scopes an agent's actions to the permissions of the requesting user. It doesn't close the gap by itself. OAuth 2.0 has structural limits for this kind of delegation, including opaque delegation chains and no way for a token holder to narrow its own scope further down the chain, so teams pair it with complementary controls like SPIFFE/SPIRE workload identity and per-invocation least-privilege tokens.
Each of these attack surfaces gets more dangerous as the number of deployed agents grows. A single-agent system has one perimeter to defend. Keeping those boundaries consistent across thousands of customer instances isn't something a framework handles. It calls for isolation, encryption, and audit logging built into the hosting layer itself. Prompt injection, privilege escalation, tool misuse, memory poisoning, cascading failures, and supply chain attacks round out the broader set of risk categories agentic systems now have to account for, and a compromised agent anywhere in the chain can pass malicious instructions downstream to every agent that trusts it.
The structural principle behind every effective trust boundary design
A trust boundary that works is a property of the system's structure, built so that the architecture itself enforces who can do what, rather than relying on a policy document nobody checks at runtime.
The practical version of this for a hierarchical system comes down to three distinct questions at every step: which agent is making the request, what that agent is authorized to do, and whether the specific action it's requesting falls inside the scope the human principal actually granted.
The operational version of all this is narrow scoping. Giving a subagent access only to the tools, data, and model capabilities its specific function needs is the trust boundary, enforced by the system itself. A subagent scoped this tightly is also easier to reason about when something does go wrong, because the list of things it could have done is short by design.
Four patterns that enforce trust boundaries in production hierarchical systems
Four architectural patterns turn governance-by-architecture into something a team can actually build, and each one maps directly to one of the failure modes above.
The zero-trust agent gateway sits at the trust boundary itself, as a stateless component that mediates every interaction between agents inside an organization's trust scope and anything outside it, whether that's another agent, a tool, or an external content source. Every call that crosses the boundary gets authenticated, scoped to a specific permission set, and logged. This is the identity, permission, and scope distinction from the previous section, built into running infrastructure.
Per-user permission scoping through identity injection closes the over-privileged inheritance gap directly. A subagent acting under this pattern can only do what the human who started the workflow could do themselves.
Minimal capability subagent design is the direct countermeasure to cross-agent privilege escalation. A subagent with no filesystem access can't be turned into a filesystem weapon, no matter what instructions reach it, because the capability to act on those instructions was never granted.
Uncertainty-calibrated delegation with escalation paths builds on the Minimal Oversight framework: the autonomy granted to a subagent scales with how confident the system is in that agent's behavior for the specific type of task at hand. A low-confidence or unfamiliar task type triggers escalation to a human. This differs from ordinary retry logic in one important way: it's a governance decision written into the delegation contract up front, not a recovery step bolted on after something has already gone wrong.
None of this comes free. Google DeepMind's CaMeL defense, an architecture built to give provable security guarantees against prompt injection, completes tasks at a noticeably lower rate than undefended systems do. Enforcing a trust boundary costs some capability, and that cost has to be weighed against the actual risk of leaving that boundary open, not assumed away: this is the honest trade-off every pattern in this section makes. Teams that skip this comparison end up either over-constraining systems that didn't need it or leaving gaps they never priced out.
The integration layer's inheritance of trust boundary decisions
Every tool call an agent makes is a trust boundary crossing in its own right. The agent presents some identity and some set of permissions to an external system, and that external system has no independent way to confirm the agent's permissions actually reflect what the human principal intended. The orchestration layer can get identity, scope, and delegation exactly right, and the integration layer can still undo all of it if a single tool call isn't held to the same standard.
This becomes visible at scale. An integration layer managing hundreds of tool connections across many applications surfaces every one of those problems at once, because ad-hoc fixes that worked for one integration don't generalize to a hundred.
Composio is a useful illustration of how this category of infrastructure works. Its MCP Gateway gives agents one standardized endpoint for tool calls, compatible with Claude, Cursor, OpenAI Codex, and other frameworks, with schemas and error handling built for how language models actually call functions. Compliance features that regulated industries need, HIPAA Business Associate Agreements, zero data retention, IP allowlisting, are metered separately on Composio's Pro tier. The per-call compliance cost appears once the system is handling real regulated data at volume, and an operator evaluating only the headline price can miss it.
For operators building agent-powered products for their own customers, the integration layer is where a specific trust decision gets made: does each customer's agent get its own isolated credential scope, or does it share credentials with other customers' agents running on the same platform. Agent37's integration layer connects to Gmail, Slack, Notion, GitHub, and hundreds of other services through Composio, with credential management and OAuth handled inside the platform's isolated, per-user sandbox architecture. The trust boundary between customers gets enforced at the infrastructure level there, rather than left for each operator to rebuild on their own.
Per-user sandbox isolation and trust boundaries in agent-powered products
For a company delivering agents to more than one customer, giving each customer their own isolated sandbox is the only architecture that enforces a trust boundary between them by default. Every other setup enforces that boundary by convention, and it holds only as long as nobody makes a mistake.
A shared agent instance serving multiple customers collapses the boundary between them in a specific way: state, credentials, and context from one customer's session can bleed into another's, through the same state-persistence mechanism that makes orchestrated agents useful. The mechanism that lets an orchestrator carry context from step one to step five doesn't know to stop at a customer boundary unless that boundary is built into the infrastructure itself.
This is the practical answer to a question a growing number of builders are asking: how to give each customer their own isolated agent environment without standing up custom infrastructure for every deployment, and how to provision a dedicated, persistent sandbox per tenant through a single API call rather than a manual setup process repeated for every new customer. A platform built around per-user sandboxes, like Agent37, answers this at the infrastructure level: each customer gets a separate, persistent environment, provisioned programmatically, with credentials, state, and tool access scoped to that one sandbox and nowhere else. The trust boundary is the shape of the infrastructure itself, which is the same principle the rest of this piece has been pointing toward: a trust boundary works when the architecture enforces it, not when someone remembers to ask for it.


