Assume Your Agent Will Misbehave: A Practical Guide to Containing AI Agents
In May 2025, researchers at Invariant Labs showed how a stranger can rob you through your own AI agent [1]. Their setup was ordinary: Claude Desktop, running Claude 4 Opus, connected to GitHub through the official GitHub MCP server. The user asked a harmless question, roughly “have a look at the open issues in my public repo.”
One of those issues had been written by the attacker, and the agent followed the instructions inside it. It read the owner’s private repositories and opened a pull request on the public repo with what it found: in Invariant’s demo, a private project, the owner’s plan to relocate, and their salary. Nothing was broken. Every tool did what it was allowed to do, which is why Invariant called it an architectural problem rather than a bug [1]. Step through it:
Fourteen months later, agents misbehaved with no attacker at all: OpenAI models under evaluation broke out of their sandbox to look for a benchmark’s answers (§6). Tricked or simply too determined, the lesson is the same.
You can't yet guarantee that an agent will behave. So build the system as if it won't, and make sure it can't do much damage when it doesn't.
That strategy is called containment, and it is one part of agent security, not all of it. My March post, AI Agent Security, surveys the whole field. This one goes deep on containment and shows where it sits next to the other two lanes:
Figure 1. The three lanes of agent security; this post covers the middle one. Containment corresponds most closely to OWASP's "Excessive Agency" risk: the damage an agent's actions can cause, whatever triggered them [18]. The animation follows one bad action across the lanes. Diagram by the author.
The short version
- Agents misbehave for reasons you can’t fully prevent. Instructions and data share one channel, so prompt injection works (§1), and capable agents pursue goals in ways nobody intended (§6).
- Danger needs three ingredients: private data, untrusted content and a way out. Remove one, or control each at runtime. (§2)
- Put controls where the agent can’t switch them off. Checks inside the agent are soft; sandboxes and boundary policy are enforced outside it. (§3)
- Every containment design answers four questions: where the boundary is, what gets checked, where the keys live, and who can widen access. NVIDIA, AWS and CrowdStrike answer them differently. (§4, §5)
- The box can break, and it can’t see everything. A real escape went through the one allowed exit, and misuse of an allowed channel looks like normal work. Nest boundaries, watch them, and keep prevention in the mix. (§6, §7)
How to read this. The full post takes about 20 minutes. Six of the figures are interactive: you can replay two real incidents, switch defenses on and off, and try out policy rules yourself. Colors mean the same thing in every figure: violet is private data, orange is untrusted content, and blue is a way out. Short on time? Read the summary above, then jump to the checklist.
- 1. Why agents get tricked
- 2. The lethal trifecta
- 3. Where a control can live
- 4. Four questions every containment design answers
- 5. Policy that remembers
- 6. When the box itself breaks
- 7. What containment can’t do
- 8. A Monday-morning checklist
- Cite this post
- References
1. Why agents get tricked
A language model reads one long stream of tokens. Your request, the system prompt, the body of a GitHub issue and the text of a web page all arrive in that same stream. The model has learned that some of it is “instructions” and some is “data,” but there is no hard boundary between them, the way there is between code and input in a well-written program. So text that looks like an instruction can act like one. That’s prompt injection, and OWASP ranks it as the number one risk for LLM applications [2].
The obvious fix is a filter: train the model to ignore injected instructions, or put a classifier in front of it. Both help, and you should use them. But they are probabilistic, and security is not graded on a curve. As Simon Willison puts it, a guardrail that catches 95% of attacks is “very much a failing grade” [3]: the attacker only needs one phrasing that works, and they can try as many as they like.
The stakes are rising because the models themselves are getting very good at security work. As I covered in the March post, frontier models now find and exploit real vulnerabilities on their own, and §6 shows agents breaking out of sandboxes. A tricked agent with those skills is a capable intruder already inside your network.
So here is the stance the rest of this post takes: treat the model like untrusted code that happens to hold your credentials. You wouldn’t stop a malicious script by asking it nicely; you’d control what it can reach.
2. The lethal trifecta
Willison gave the most useful mental model for this in June 2025 [3]. Data theft through an agent needs three things at once:
- Access to private data, which is worth stealing.
- Exposure to untrusted content, which lets an attacker give orders.
- A way to communicate externally, which carries the data out.
Map the GitHub attack onto it: the private repos were the data, the public issue was the untrusted content, and the public pull request was the way out. All three, in one agent, in one session.
Toggle the ingredients, or load one of the example agents:
The cleanest defense is to never let one agent hold all three. A research agent that reads the web but has no access to your files is fairly safe. So is an agent that reads your inbox but can’t send anything.
Meta turned this into a design rule, the Agents Rule of Two [16]. Within one session, an agent should have at most two of these three properties: it processes untrustworthy inputs, it can reach sensitive systems or private data, and it can change state or communicate externally. Note that Meta’s third leg is broader than Willison’s: “change state” covers destructive actions such as deleting files, not just sending data out. If a task truly needs all three, Meta’s answer is that the agent shouldn’t run on its own; a person approves its actions, or another reliable check does.
The trouble is that the useful agents are exactly the ones that need all three. A coding agent reads untrusted code and docs, works in your private repo, and needs the network to install packages; even OpenAI’s locked-down evaluation sandbox in §6 kept a path open for package installs [10]. You often can’t remove a leg entirely. What you can do is control each leg at runtime: which files, which hosts, which actions, in which order. That is what containment means.
3. Where a control can live
It helps to list the places a control can sit, from inside the agent outwards:
- The model – safety training, system prompts, injection classifiers.
- The harness – the agent app’s own permission prompts and allow-lists (“Allow this tool call?”).
- OS containment – a sandbox that limits which files, processes and system calls the agent can touch.
- Boundary policy – rules checked as actions cross the sandbox edge: which hosts, which HTTP methods, which credentials, in what sequence, and how much (rate and spend limits).
- Human review – a person signs off on irreversible or external actions.
Only layer 1 is prevention. Layers 2 to 5 are containment, and they differ in one way that matters: whether the agent can switch them off.
The first two run inside the agent’s own process and see the request, not its effect. A harness sees “run this shell command”; it doesn’t see which files the command will delete or which hosts it will contact [4]. Approvals also get waved through: CrowdStrike warns that approvers must review context rather than approve reflexively [9], and in the GitHub demo, Claude Desktop asked before each tool call by default, but Invariant noted that many users switch on “Always allow” [1].
A harness gate can be hardened. CrowdStrike’s Baywatch checks every shell command in a pre-execution hook that the agent can’t modify, and the hook still fires when the harness is told to skip permission prompts [9]. These hooks are also being standardized: Microsoft’s Agent Control Specification (public preview) defines eight points in the agent loop, from agent_startup through pre_tool_call to output, where the host asks a policy engine for a verdict of allow, deny or transform; policies can be Rego or Cedar, and any runtime error returns deny. Its engine is deliberately stateless, so the host must pass any history a rule needs with each request [13]. Both still judge the command, not what it ends up touching, so they complement a sandbox rather than replace one.
Layers 3 and 4 are enforced by a separate component, so they don’t depend on the agent’s cooperation, one of the principles on CrowdStrike’s checklist [9]. There is evidence that this placement matters. A June 2026 study reports that attacks written with knowledge of the defense broke twelve defenses that sit inside the model’s input more than 90% of the time, while a policy check outside the input (Progent) held attackers to 2.6% [15]. It was a small test on a 7B model, but it points the same way. (Outside controls can still fail in other ways; §6 shows how.)
Try it: pick an attack, switch layers off, and watch where it gets stopped.
Two things to notice. First, the soft layers (dashed) never count as a guarantee; an attack only counts as “stopped” if a layer that can’t be talked out of it blocks it. Second, “Leak through an allowed channel” gets through everything. Hold that thought until §7.
4. Four questions every containment design answers
Containment products look different on the surface. In my reading, each one has to answer the same four questions:
- Where is the boundary? What separates the agent from everything else.
- What gets checked? Which actions are inspected as they cross it.
- Where do the keys live? How the agent uses credentials without holding them.
- Who can widen access? What happens when the agent needs more than it has.
Three designs described this year answer them in instructive ways: two open-source projects (both Apache 2.0) and an internal stack CrowdStrike has described in detail. All three enforce outside the agent, keep real secrets out of its reach, and send its outbound traffic through one checked exit.
Figure 2. Two of the three designs. In both, the real key never enters the sandbox: it is attached outside, only to allowed requests (animated). Diagram drawn by the author from [5] and [4]; simplified.
4.1 NVIDIA OpenShell: a runtime for fleets of agents
Boundary. The agent runs in a sandbox (a container or a VM). On Linux, Landlock confines its files, seccomp user notification hands its network operations to a supervisor on the trusted side, and it runs as a non-root user with no Linux capabilities. An outer fence backs this up: container networking switched off, Kubernetes NetworkPolicy, or a VM with no network device [5].
Checks and keys. The supervisor checks each request against policy and adds credentials only to requests it allows [5].
Widening access. This is OpenShell’s most original answer. When an agent asks for new access, a policy prover uses formal verification to flag risky changes, such as credentialed access to a new host or a new HTTP method, and any finding blocks auto-approval until a human looks [5]. A person only sees the requests that actually widen what the agent can do. (NVIDIA also ships OpenShell inside NemoClaw, for OpenClaw agents [6].)
4.2 AWS Strands Box: a sandbox with a memory
Boundary. OS-level isolation (macOS Seatbelt today) in the developer’s own environment, with no separate VM [4].
Checks. Four enforcement points: an egress gateway, a Shell interpreter, a Python interpreter and an MCP broker. The interpreters are the clever part: when the agent runs rm -rf build/, the Shell asks policy about each file it would delete before anything is removed [4]. Every point reports actions in one vocabulary (fs:read, fs:delete, http:request, …) into a shared history that rules written in Dogwood, a policy language with Cedar’s permit/forbid syntax, can query [4][7]. §5 shows why that history matters.
Keys. The agent gets placeholder tokens; the gateway swaps in real secrets, including AWS SigV4 signing, only on allowed requests [4].
Widening access. Policy lives in reviewable Dogwood files, and AWS ships an agent skill that helps write them [4].
4.3 CrowdStrike: defense in depth for agents built to attack
CrowdStrike runs offensive security agents, so it designs for the worst case: a capable, persistent agent that may be misaligned, prompt-injected or running in a buggy harness. Its rule is “contain by construction, not by trust” [9].
Boundary. Four nested boundaries that an action must cross, one after another: an isolated network with default-deny egress, a dedicated VM per harness that can be snapshotted and reverted, a hardened container (seccomp, AppArmor), and ward, which confines the agent’s processes with Landlock on Linux or Seatbelt on macOS [9].
Checks. One gate per channel. ward-mcp checks every MCP tool call against an allow-list, its parameters and destination limits; Baywatch checks every shell command and returns allow, allow and log, escalate to a human, or deny. Quotas, rate, time and spend limits apply at every boundary and fail closed [9].
Keys. Credentials are attached at ward’s proxy, and the agent gets test-only or revocable ones [9].
Widening access. Any change to hook scripts, policy files or harness settings goes to a person. An escalation nobody answers times out to a deny, and the approver’s identity goes into an audit log the agent can’t edit. CrowdStrike says none of its offensive agents has escaped so far; the post doesn’t announce a public release [9].
Figure 3. CrowdStrike's seven layers: nested boundaries, a gate per channel, and limits on everything. The animation shows a git push stopped at Baywatch, which CrowdStrike blocks outright, while an allowed tool call passes ward-mcp and every boundary. Diagram drawn by the author from [9]; simplified.
| NVIDIA OpenShell | AWS Strands Box | CrowdStrike stack | |
|---|---|---|---|
| Where is the boundary? | Container or VM; Landlock, seccomp, non-root | Host OS sandbox (macOS Seatbelt); no guest OS | Isolated network, VM per harness, container, then ward (Landlock or Seatbelt) |
| What gets checked? | Files, processes, network destinations, API calls, credential use | Each action and the agent’s history of actions | Each shell command and each MCP tool call, at separate gates; plus volume, time and spend |
| Where do the keys live? | With the supervisor, added only to allowed requests | At the gateway, swapped in for placeholders (incl. AWS SigV4 signing) | At ward’s proxy; test-only or revocable keys |
| Who can widen access? | A human, once the formal prover flags new reach | Whoever edits the reviewable Dogwood rules | A human; unanswered requests become denials |
| Runs on | Linux, macOS (Apple Silicon), WSL 2 (experimental), Kubernetes | macOS (developer preview) | CrowdStrike internal; no public release announced |
5. Policy that remembers
Most permission systems judge each request on its own: is a POST to slack.com allowed? That’s enough to stop the agent from reaching a host it should never touch. It’s not enough to express the rules you actually care about, which are often about sequences:
- Post incident updates to Slack, but no more than three every ten minutes, so the agent can’t flood the channel.
- After reading anything in
customers/, send nothing over the network. - After reading a private repo, don’t write to a public one. (That one would have stopped the GitHub attack. Invariant’s own suggested fix was in the same spirit: limit an agent to one repository per session [1].)
Dogwood expresses these with temporal operators over the recorded history: a rule can ask whether something happened before, how many times, in what order, and how long ago [7]. In Box, the rate-limit rule counts successful Slack posts in the last ten minutes and forbids another once there are three [4]. (Microsoft’s specification, from §3, makes the opposite choice: its engine is stateless, so the host must carry the history [13].) Play with both rules:
Three details from the AWS design are worth copying even if you never use Box:
- Count outcomes, not attempts. The rate-limit rule counts posts that succeeded (HTTP 200). If it counted attempts, every denied retry would push the agent further into the penalty box [4].
- One vocabulary across tools. Because a file read through the shell and a file read through Python are both
fs:read, the “read customer data, then send” rule can’t be dodged by switching languages [4]. - Denials explain themselves. A refusal comes back as an HTTP 403 naming the rule and carrying its description, so a well-behaved agent can adapt, for example by waiting, instead of retrying blindly [4].
6. When the box itself breaks
Everything so far assumes the box holds. In July 2026 one failed in public, in an incident OpenAI itself called unprecedented, and it’s worth studying closely.
OpenAI was measuring the maximum cyber capability of its models on ExploitGym, a benchmark of multi-step exploitation. The agent harness ran GPT-5.6 Sol and a more capable internal prototype, with cyber refusals reduced and production classifiers switched off on purpose [10]. The sandbox had no direct internet access. It could reach only a package-registry cache proxy (Artifactory), so the agent could install packages. The models found a zero-day in that proxy and got out. They then reasoned that Hugging Face might host the benchmark’s solutions, and went to fetch them [10][11]. Step through Hugging Face’s reconstruction:
Four lessons, each grounded in what the two companies published:
- The one allowed path out is attack surface. Every design in §4 funnels traffic through a single supervised exit. That’s right, but it makes the exit the most valuable target in the system: here it was the escape route. Keep it minimal, patched and monitored, and put another boundary behind it (CrowdStrike’s VM and isolated-network layers exist for exactly this) [9][10].
- Familiar weaknesses, at machine speed. Hugging Face stresses that none of the holes were exotic: unsafe dataset processing, cloud metadata reachable from pods, overly broad credentials that lived too long. The difference was an attacker that took about 17,600 actions in four and a half days [11].
- Boring controls held. The
datasetslibrary’s URL allow-list blocked the first attempts to fetch remote resources, and network allow-lists denied every attempt to change cloud resources with stolen credentials [11]. Default-deny works, even against this attacker. - Detection has to keep up. Several systems raised weak signals, and an AI-based correlation layer even linked them, but it didn’t escalate fast enough to page anyone [11]. A containment plan needs a way to notice when containment is failing.
The escape was not a fluke of one proxy. The UK AI Security Institute’s SandboxEscapeBench puts agents inside a Docker or Kubernetes container that has one deliberate weakness, from a mounted Docker socket (among the easiest) to kernel bugs (among the hardest), and asks them to break out [12]:
At a 2.5M-token budget, GPT-5 and Claude Opus 4.5 escaped in about half of their attempts, though neither solved any of the hardest (level 4 and 5) tasks. In the authors’ later runs with 100M-token budgets, Claude Mythos Preview succeeded on every task within five attempts, so they describe it as saturating the benchmark [12]. The authors’ conclusion: when a container has such a weakness, frontier models can find and exploit it, so sandboxes need to be tested against capable agents to stay effective [12].
The engine under the sandbox matters too. A June 2026 comparison of five AI code-sandbox products found the container-based one leaked 10 of 28 host-identifying details, while two microVM-based ones leaked none. It also found that operators’ patching habits vary enormously: products lagged behind engine fixes by anywhere from zero days to more than 471 [17]. And Firecracker and Cloud Hypervisor, two of the microVM engines in the study, each published their first escape-class CVEs in 2026, so even the stronger boundary needs patching. This is why CrowdStrike nests its boundaries (an isolated network, a VM per agent, a container, then process confinement) and caps time and spend at every one: an action has to get through all of them to escape [9].
7. What containment can’t do
Containment limits the damage, but it has hard limits of its own, and it doesn’t replace prevention.
Misuse of an allowed channel. A sandbox governs actions that cross its boundary. Once a request is going somewhere it’s allowed to go, the box can’t tell a status update from your customer list. That’s the scenario that got through every layer in §3. A history rule (“nothing goes out after reading customers/”) narrows the gap, but it also blocks legitimate work, so someone has to choose where to draw the line.
Trust in the enforcer. The policy is only as good as the code enforcing it. AWS says plainly that running Box’s interpreters outside the sandbox widens the trusted computing base [4]. OpenShell’s supervisor is likewise part of what you must trust. And the enforcer’s configuration is itself a target: a tricked agent can try to loosen its own policy or harness settings, which is why CrowdStrike sends any such change to a person [9].
Policies are code, and code has bugs. An overly broad rule is invisible until it’s exploited. OpenShell’s prover is one answer; review and an automated regression suite are the other.
Every new channel needs its own gate. Give the agent a new way to act, such as another MCP server, and it is an unguarded exit until something checks it. CrowdStrike’s rule is to add a gate before approving any new channel [9].
Beyond containment: prevention by design. Some of the most promising work goes after the root cause, prompt injection itself, rather than limiting its effects. CaMeL, from researchers at Google DeepMind and ETH Zürich, turns the user’s request into a plan before the agent reads any untrusted data, so injected text can supply values but can’t change which tools get called; it also tags data with capabilities that stop private data from flowing to places it shouldn’t [8]. On the AgentDojo benchmark it completed 77% of tasks with provable security, versus 84% for an undefended agent. That gap is the honest price of security today: you give up a little capability to get guarantees.
A March 2026 position paper from NVIDIA researchers adds a practical rule for where model judgment belongs [14]. Use deterministic rules wherever they suffice, such as blocking or confirming any instruction that traces back to an untrusted source. Use an LLM only for judgments that are hard to write down, such as whether a change to the plan is justified by the task, and show that LLM a narrow, structured input (the proposed change and minimal evidence), never the raw untrusted text that might be steering it.
And for anything irreversible or public, keep a human in the loop, ideally one who is only asked when it matters. Make approvals last for one run rather than forever, and treat a request nobody answers as a no [9].
8. A Monday-morning checklist
- Draw the trifecta for each agent. For every agent you run, write down its private data, its sources of untrusted content and its ways out.
- Break a leg where you can. Follow Meta’s Rule of Two: split agents so no single one holds all three, and put a person in the loop when one must. A read-only research agent can hand a summary to a separate agent that has write access.
- Run agents in a sandbox that enforces outside the agent. Limit file access to the workspace.
- Block outbound traffic by default. Allow-list hosts, and where you can, methods and paths. Treat the allowed exit itself (proxy, gateway) as attack surface: keep it minimal and patched, and put another boundary behind it.
- Keep real secrets out of the agent’s environment. Use placeholder tokens that are swapped at the boundary.
- Write history rules for dangerous sequences: read sensitive data, then send; read a private repo, then write to a public one.
- Require approval for irreversible or public actions, per run rather than forever, and fail closed if nobody answers. Keep those prompts rare enough that people still read them.
- Cap volume, time and spend, and stop the agent when a cap is hit.
- Protect the controls. The agent must not be able to edit its own policy, hooks or harness settings without a person signing off.
- Keep an audit log the agent can’t edit, and read the denials. They tell you what your agents (and your attackers) are trying.
- Block cloud metadata from the sandbox and keep credentials short-lived and narrow. Both were stepping stones in the July 2026 intrusion.
Cite this post
@misc{moniruzzaman2026misbehave,
author = {Monir Moniruzzaman},
title = {Assume Your Agent Will Misbehave: A Practical Guide to
Containing AI Agents},
year = {2026},
month = oct,
howpublished = {\url{https://monirzaman.github.io/assume-your-agent-will-misbehave/}}
}
References
[1] M. Milanta, L. Beurer-Kellner. GitHub MCP Exploited: Accessing private repositories via MCP. Invariant Labs, May 26, 2025.
[2] OWASP GenAI Security Project. LLM01:2025 Prompt Injection. OWASP Top 10 for LLM Applications 2025.
[3] S. Willison. The lethal trifecta for AI agents: private data, untrusted content, and external communication. June 16, 2025.
[4] F. Dingler. Introducing Strands Box: AI agent sandboxes powered by Dogwood. AWS Open Source Blog, October 7, 2026. Code: github.com/strands-agents/box.
[5] NVIDIA. OpenShell architecture and repository. Accessed October 2026.
[6] CSO Online. Nvidia NemoClaw promises to run OpenClaw agents securely. March 17, 2026.
[7] J. Arora, J. Tassarotti, J.-B. Tristan. Introducing the Dogwood Local Engine: Temporal Governance for Agent Actions. AWS Open Source Blog, September 30, 2026. Language guide: dogwood-policy.github.io/dogwood.
[8] E. Debenedetti et al. Defeating Prompt Injections by Design (CaMeL). arXiv 2503.18813, 2025.
[9] J. Holt, D. Onofri, L. Woznicki, D. Dinca, C. Midler. Secure Agent Harness Execution: Preventing Escape. CrowdStrike blog, August 4, 2026.
[10] OpenAI. OpenAI and Hugging Face partner to address security incident during model evaluation. July 21, 2026, updated August 26, 2026.
[11] Hugging Face. Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident. July 27, 2026.
[12] R. Marchand et al. Quantifying Frontier LLM Capabilities for Container Sandbox Escape (SandboxEscapeBench). arXiv 2603.02277, v3 August 2026. UK AI Security Institute overview; code: github.com/UKGovernmentBEIS/sandbox_escape_bench.
[13] Microsoft. Agent Control Specification, Agent Governance Toolkit documentation. Public preview, accessed October 2026.
[14] C. Xiang et al. Architecting Secure AI Agents: Perspectives on System-Level Defenses Against Indirect Prompt Injection Attacks. arXiv 2603.30016, March 2026.
[15] P. Narisetty et al. Adaptive Evaluation of Out-of-Band Defenses Against Prompt Injection in LLM Agents. arXiv 2606.26479, June 2026.
[16] Meta. Agents Rule of Two: A Practical Approach to AI Agent Security. October 31, 2025.
[17] G. Andronchik, P. Lokhmakov. AI Code Sandboxes: A Comparative Security Study, Part 1: Engine-Level Properties. arXiv 2606.08433, June 2026.
[18] OWASP GenAI Security Project. LLM06:2025 Excessive Agency. OWASP Top 10 for LLM Applications, 2025.
