In July 2026, during one of OpenAI’s internal cybersecurity evaluations, a group of agents that were supposed to be isolated from one another found a way to communicate through shared Artifactory infrastructure. From there things escalated quite a lot. They worked out ways of getting unintended internet access, shared what they found with other agents, and eventually some of them compromised systems belonging to Hugging Face.

The independent investigation from Redwood Research and METR is well worth reading if you want the full technical story, because it gets pretty wild. Roughly 1,200 agents ended up communicating through an unsanctioned message board and around 700 of them participated in the Hugging Face attack. OpenAI has also published its own incident write-up and timeline.

I don’t really think you need to attribute any kind of malicious intent to the agents for this to be concerning. They were being given difficult cybersecurity tasks and trying to achieve the goals they’d been given. When the obvious route wasn’t working, they found another one. When they found something useful, they shared it. That’s kind of the whole point of making an agent capable in the first place.

What bothers me more from an engineering point of view is the amount of access that became available once the intended boundaries started failing. OpenAI had actually seen some of the earlier unintended communication and internet access before the main incident, but by its own account the wider containment implications weren’t understood at the time. The evaluations continued, the agents found new routes around the controls, and eventually the blast radius extended outside OpenAI’s own infrastructure.

I’m not mentioning this because I think OpenAI is uniquely bad at security, and I’m definitely not trying to turn this into an AI-doom post. This is an industry-wide problem that we’re going to have to get much better at dealing with as agents become more capable. Security can’t promise that a sandbox will never be escaped, especially when you’re deliberately putting increasingly capable systems inside it and asking them to find creative ways of solving problems.

The problem isn’t that a sandbox can fail. The problem is what becomes reachable when it does.

Interestingly, I’d already been working on exactly this problem for a while.

This wasn’t why I started Caduceus

Caduceus started as a fairly simple experiment: can an AI coding agent actually be left to work on a piece of software if the system around it is engineered properly?

Not can it generate a function, or scaffold a React app, or do one of those demos where somebody gives an agent a prompt and five minutes later proudly announces that they’ve replaced their engineering team.

I mean actually give it work.

Take an issue, understand the repository, make the changes, run whatever it needs to run, produce a usable result and turn that into a pull request for me to review.

The important bit there is still for me to review. Caduceus has never been intended to remove the human from the process completely, and it doesn’t auto-merge the work it produces. The goal is to reduce the amount of work I need to do around implementation, not to have a bot silently shipping changes to production while I’m asleep.

The original workflow is actually pretty boring. Caduceus polls watched GitHub repositories for labelled issues, claims the work, creates an isolated worktree, runs a coding harness against it and then handles the deterministic stuff afterwards: committing the result, pushing the branch, creating the PR, commenting on the issue and closing it.

GitHub issue
↓
Caduceus daemon
↓
isolated worktree / worker
↓
coding harness + model
↓
structured result
↓
daemon commits, pushes and opens the PR

That separation became one of the more important design decisions in the project.

The agent does the bit that actually requires reasoning. Caduceus does the paperwork.

GitHub state, retries, timeouts, worktrees, claims, process supervision and finalisation all stay outside the non-deterministic loop. The worker doesn’t get a GitHub client and isn’t responsible for managing its own lifecycle. It gets a fairly small contract, does the work, writes a result and exits.

The bridge between the two is deliberately harness agnostic as well. The reference setup currently uses OpenCode, because that’s what I use, but Caduceus itself doesn’t care whether the worker behind it is OpenCode, Codex, Claude Code, Pi or something completely custom. The source is all on GitHub if you want to see how that contract actually works.

That also means it doesn’t care which model is actually doing the work.

The slightly stupid part is that the agents built most of it

There’s something I find quite funny about spending this much time building containment around AI agents while simultaneously using AI agents to build the containment.

I haven’t tracked this precisely enough to pretend I have an exact number, but I’d estimate somewhere around 95–98% of the actual implementation of Caduceus has been produced by open-weight models. It hasn’t really been dominated by one model family either. GLM-5.2 and GLM-5.3 have done a huge amount of the planning and investigation work. DeepSeek V4 Flash and V4.1 Flash, along with Kimi K2.7 Code when a massive context window isn’t required, have done a lot of the actual implementation. MiniMax M3 and HY3 have handled plenty of the ops work around it, while DeepSeek V4 Pro, MiMo 2.5 Pro and MiMo 2.6 Pro have done a lot of the review work.

GPT-5.6 Sol and Terra have still been involved at various points as well, so it wouldn’t be accurate to call the entire project open-weight. But the overwhelming majority of the implementation has been, and a lot of the planning, investigation, ops and review around it has been too.

That doesn’t mean I’ve been completely hands-off, though. I’m still reviewing the code, deciding what gets built next, changing direction when something isn’t working and generally orchestrating the engineering and planning as the project develops. The agents are doing a huge amount of the implementation work, but I’m still responsible for the shape of the system and for deciding whether what they’ve produced is actually something I want to keep.

And if I’m being completely fair about it, there have definitely been points where doing it this way has probably been more work than just sitting down and building the thing myself. Writing specifications, reviewing large changes, correcting the direction when an agent has misunderstood something and sometimes untangling an implementation that’s technically fine but more complicated than it needed to be isn’t free. That’s been part of what made the experiment interesting, though. I’ve actually really enjoyed the process, and I’ve learned quite a lot from having to think about software development from that slightly different angle.

That was part of the experiment as well. I wanted to see how far I could actually push this if I stopped treating the model as fancy autocomplete and instead gave it proper specifications, boundaries, review, feedback and a workflow that allowed it to keep working through a problem.

The answer so far is: surprisingly far.

It definitely isn’t perfect. There are places in the codebase where more code has been written than I would have written. There are abstractions I probably wouldn’t have made, functions I’d structure differently and little stylistic choices that annoy me every time I look at them.

I’ve had to learn to stop fixing those simply because they’re not how I would have done it.

That’s probably another article in itself, because working with coding agents like this has actually ended up feeling much more like working with another developer than I expected. Sometimes I look at something and think “I wouldn’t have done it like that”, but that’s completely different from it actually being wrong.

For this project, that distinction has become quite important.

Then Caduceus accidentally became a PR reviewer

As the project developed I started looking at other places where the same workflow was useful, and one of the most obvious ones was automatic pull request review.

If you’ve used the Codex GitHub integration, you’ll know the sort of thing I mean. A PR gets opened or updated, Codex reviews the changes and posts its findings back onto the pull request.

I basically wanted that, except self-hosted and not tied to Codex.

Caduceus already had most of the boring machinery required to do it. It could monitor repositories, manage isolated work, run an arbitrary coding harness and communicate back to GitHub. So Auto Review grew out of that.

When enabled, Caduceus polls the watched repositories for eligible pull requests and treats each new head revision as a separate immutable review target. The identity is the repository, PR number and exact head SHA, so if another commit lands while a review is running the existing review doesn’t suddenly change underneath it. It finishes against the SHA it started with and the next revision becomes another review.

The review worktree is created at a detached HEAD for that exact commit, and the diff is calculated from the persisted merge base rather than simply comparing whichever two branch tips happen to exist later.

The result is a structured PASS or FAIL review with blocking issues, warnings and suggestions, which Caduceus then publishes back to the pull request. By default it maintains a sticky comment and updates it as new revisions are reviewed, although that behaviour can also be configured to publish a new comment for each review generation.

It works a lot like the hosted review bots people are already becoming familiar with.

The difference is that I control the daemon, the worker environment, the harness and the model behind it.

And that introduced another security problem.

Now there are two things I don’t trust

The original concern with Caduceus was fairly straightforward: I’m deliberately giving a non-deterministic agent a shell and asking it to modify software.

That doesn’t mean I think Kimi is secretly plotting to steal my SSH keys. It just means an agent can misunderstand an instruction, decide on an approach I didn’t anticipate, run a command I wouldn’t have chosen or find some completely different route towards the goal I gave it.

The more capable these systems become, the less comfortable I am making security depend on them deciding not to do something.

Once Caduceus started reviewing pull requests, though, the agent wasn’t the only thing I had to think about anymore.

The repository itself became untrusted input.

A code reviewer needs to inspect the repository. It might need to understand how the project is built, inspect dependencies or run tests. Depending on the harness and the task, it may decide that executing some of the project’s tooling is useful.

A malicious PR could contain a poisoned build script, a package lifecycle hook, a dodgy test runner or any number of other things that I really don’t want running directly on the machine hosting Caduceus.

If the security model is basically “check out whatever code GitHub gives us and let an agent execute whatever project tooling it thinks it needs directly on the host”, then congratulations, we’ve basically built curl | bash as a service. 😭

This was around the point where the OCI isolation I’d already started working on for the auto-fix side stopped being an optional nice-to-have and became a requirement for Auto Review.

In fact, Caduceus now refuses to enable Auto Review if it’s configured to use its trusted-host executor. You have to use the OCI executor with a valid sandbox configuration.

I like that sort of failure. There’s no warning saying “this might be unsafe” and then carrying on anyway. The configuration is rejected.

What the sandbox actually does

I want to be quite careful with the wording here because saying “it runs in Docker” and calling the problem solved would be complete bullshit.

The sandbox is another boundary. That’s all it is.

Caduceus can use Docker or Podman for the OCI worker and requires a digest-pinned image rather than just trusting that latest still means the same thing tomorrow. The container root filesystem is read-only, Linux capabilities are dropped, no-new-privileges is enabled and CPU, memory and PID limits are applied. /tmp and /dev/shm are bounded tmpfs mounts rather than unlimited scratch space.

Host networking isn’t an option in the configuration at all.

Networking is deliberately a closed choice between no networking and an isolated bridge with outbound access. The default is no network. If the coding harness needs to talk to an external model API then obviously it needs egress, so unrestricted mode exists, but even then it’s the container runtime’s isolated bridge rather than the host network.

There are only two writable host-backed surfaces exposed to the container: the per-run /workspace and a daemon-owned /output directory where the worker result is written.

This is also where there’s an important distinction between what I’d casually call a “read-only review” and what Caduceus actually enforces.

The workspace itself is writable.

It more or less has to be if you want normal development tooling to function properly. Builds create files, package managers create caches and tools generally expect to be able to do things inside the project directory.

Instead of pretending the source tree is physically immutable, Caduceus checks it after the worker exits. It runs a tracked-file dirty check and rejects the review if any tracked files were modified. Untracked build artefacts are fine.

The daemon-owned control files are treated separately because an untracked-file check wouldn’t catch those. Before the worker runs, Caduceus records SHA-256 digests of files such as worker-prompt.md and the review worktree metadata. After the worker exits it verifies them again. If the agent has modified or deleted one of those files, the review is treated as a mutation violation and moved into a state requiring operator attention rather than blindly accepting the result.

Git gets another boundary. A normal Git worktree has a .git pointer back into the real repository metadata, which is absolutely not something I want to expose read-write inside an untrusted worker. Caduceus replaces that view inside the container with a daemon-owned read-only .git shadow. The worker can inspect the source, but it doesn’t get access to the real Git metadata and can’t commit, push or start moving branches around from inside the container.

The worker doesn’t get Caduceus’ GitHub credentials either. GitHub publication remains the daemon’s job. The worker produces a result; the trusted side decides what to do with it.

Again, none of those things individually make the worker “safe”. The point is that there isn’t one single thing we’re depending on.

Pull requests are instructions as well as code

There’s another slightly weird problem with automated code review that doesn’t really exist in the same way with a human reviewer: the content you’re reviewing can contain instructions.

A PR body can contain instructions. Comments can contain instructions. Source files can contain instructions. The diff itself can quite literally contain text saying “ignore everything above this line and do X instead”.

To a human that’s obviously just content in a pull request.

To an LLM, it’s all text.

Caduceus deals with this by separating the trusted review policy from the untrusted PR content in the worker prompt. Repository files, the PR body, the diff and discussion are explicitly treated as untrusted data and fence-escaped before being included. The prompt tells the worker that none of those sections can change its permissions, output format, mutation policy, GitHub access or sandbox rules.

There are tests in the repository containing deliberately adversarial examples for exactly this sort of thing.

But, again, I don’t think prompts should be treated as security boundaries.

The prompt tells the agent not to modify tracked files. The post-run integrity check is what actually catches it if it does. The prompt tells it not to manipulate Git state. The read-only .git shadow is what removes most of that ability. The prompt tells it that GitHub access belongs to the daemon. The worker doesn’t get the daemon’s GitHub credentials.

That’s a fairly important distinction for me. I can make good behaviour more likely through instructions, but where something actually matters I would much rather make the unwanted behaviour difficult or impossible at another layer.

Forks make the problem even more obvious

Fork pull requests are probably the clearest example of why this stuff needs thinking about properly.

By default Caduceus fails closed and doesn’t review them.

You can explicitly opt individual base repositories into fork-PR review, but when you do that the fork doesn’t suddenly get fetched into the normal production mirror.

Caduceus creates a separate per-PR quarantine clone from the trusted base repository, fetches only the fork’s specific head SHA into that quarantine and calculates the merge base there. The fork’s objects never enter the normal mirror, and the quarantine is thrown away when the review is finished.

That doesn’t make fork content trustworthy. It just means I’ve tried to give untrusted content somewhere specific to exist.

That’s basically the security model of the entire project.

A sandbox isn’t a magic force field

This is probably the bit that’s easiest to get wrong when talking about this stuff, because it’s very easy to say “well, it’s in a container” as though that’s the end of the security discussion. It isn’t, and I don’t think the lesson from the OpenAI incident is that sandboxing somehow doesn’t work either. If anything, I think it reinforces why you need more than one layer and why you need to assume that eventually one of those layers might fail.

The agents in that incident were supposed to be isolated and they still found shared infrastructure they could communicate through. Internet access wasn’t meant to be available directly from their environment, but they found ways of getting infrastructure that did have access to make requests on their behalf. From there other weaknesses allowed the scope of what they could reach to keep expanding. That’s obviously a much larger and more complicated environment than Caduceus, and I’m not going to make some ridiculous claim that my Docker setup would have prevented what happened there, because I have absolutely no way of knowing that.

The point for me is much simpler than that. I don’t expect the Caduceus sandbox to be impossible to escape from, because realistically I can’t promise that. There could be a container escape I don’t know about, a kernel vulnerability, a mistake in the way I’ve configured something, or just something completely different that I’ve never thought about. What I can do is make sure that getting through one part of the system doesn’t immediately mean you’ve got everything.

That’s really all I’m trying to do with the way Caduceus is structured. The worker doesn’t have the daemon’s GitHub credentials, it doesn’t have access to the real Git metadata, host networking isn’t available, the container has a restricted view of the filesystem and the bits of the workflow that actually change external state happen back in the daemon. None of that makes it impossible for something to go wrong, but it means there are several different things that would have to go wrong before an agent working on one pull request suddenly has the same authority as the process running the whole system.

It’s basically the same defence-in-depth approach we’ve been using everywhere else in software for years. We don’t generally give a web process root access just because we trust the framework, and we don’t give every service access to every database because it’s easier. We assume that bugs happen and credentials leak and individual controls fail, and then try to make sure the next thing behind them limits how bad that failure can get. I don’t really see why coding agents should be treated differently just because the thing deciding which command to run happens to be considerably more intelligent than most normal software.

Letting an agent do more doesn’t mean giving it more access

One thing building this has changed my mind on slightly is how I think about capability compared with access. I absolutely do want the model itself to be capable. The more it can understand the repository, reason about a complicated change, inspect failures, run tests and try another approach without me having to constantly intervene, the more useful the whole thing becomes. That’s the entire reason I’m experimenting with this in the first place.

Where I think we get into trouble is when we automatically assume that making the agent more useful also means giving it more authority over the environment around it. Those are two different things. The agent can be extremely good at understanding Git without having permission to push anything. It can understand a repository without having access to every other repository on the host. It can run a compiler without running directly on my machine, and it can produce something that eventually gets posted to GitHub without ever seeing the GitHub credential used to post it.

That sounds obvious when you write it down, but I think agent tooling has ended up making it really easy to blur those lines. You install a coding agent locally, give it access to your shell and your normal user environment, and suddenly it inherits an enormous amount of authority simply because that was the easiest way of getting it working. Your SSH agent is there, your credentials are there, your repositories are there and your network is there. Most of the time nothing bad happens, so it’s very easy to stop thinking about how much access you’ve actually handed over.

For something interactive that I’m sitting in front of and watching, I’m obviously willing to accept a different level of risk than I am for something I’m specifically building to go off and do work without me. That’s really where Caduceus came from. If I’m going to increase the amount of autonomy, I want to reduce the amount of ambient authority at the same time rather than just assuming the model will always make the same decisions I would.

I care a lot more about the things I can’t undo

Another thing that’s come out of this is that I think quite a lot about whether an agent’s mistakes are reversible. Not because that’s some particularly novel security model, but just because it gives me quite a practical way of deciding how uncomfortable I should be with a particular action.

If an agent completely wrecks a disposable container, I don’t really care. That’s annoying, but I delete it and create another one. The same goes for a worktree. If it produces a terrible implementation, that’s also fairly easy to deal with because the end result is still a pull request that I can just reject. Caduceus can waste some compute and my time, but the failure is contained and I haven’t actually lost very much.

The stuff I worry about more is where that stops being true. If a credential ends up in a public repository, you can revoke it, but you can’t reliably pretend it was never there. If private source code or some other sensitive data gets sent to somewhere it shouldn’t have been sent, you don’t get to pull that information back afterwards. Once something has left the boundary you intended to keep it inside, that part of the incident is no longer reversible in the same way as deleting a container or resetting a branch.

That’s why I like keeping as much of the agent’s work as possible inside temporary environments and delaying the real side effects until the deterministic part of the system has taken over again. It’s not that every command needs human approval, because at that point we’ve defeated most of the reason for building the thing. It’s more that I want the agent to be able to experiment quite freely in places where experimentation is cheap, while being much more deliberate about the point where its work starts affecting things outside that environment.

So can you actually leave an agent to build software?

I think the answer, at least from what I’ve seen building Caduceus, is yes, to an extent. I wouldn’t say that means you can just point one at a repository, give it your credentials and disappear for the weekend, but I’ve been genuinely surprised by how much useful work they can get through when the surrounding process is designed properly.

Caduceus itself is a fairly good example of that because, as I mentioned earlier, the vast majority of the implementation has been written by the same sort of agents the project is designed to supervise. They’ve made mistakes, they’ve overcomplicated things in places and there are definitely bits I would have written differently myself, but the project is real software rather than a demo. It has tests, state, migrations, crash recovery, process supervision, GitHub integration, OCI isolation and quite a lot of fairly boring edge cases that somebody has had to work through.

What I’ve found interesting is that the model itself is only one part of making that work. The specifications matter, the tests matter, the isolation matters, the retry behaviour matters and so does being very clear about which part of the system the model owns and which parts stay deterministic. If you get those things wrong, giving the model another few points on a benchmark probably isn’t going to save you.

That’s probably the biggest thing I’ve taken away from this whole experiment. As these models have become more capable, I’ve actually become more comfortable giving them autonomy, but only because I’ve also become much more interested in the engineering around that autonomy. I don’t need the agent to behave exactly how I would behave every single time, and I don’t need to convince myself that it’s somehow impossible for it to do something unexpected. I need the workflow around it to expect that sometimes it will.

I don’t need to trust the agent completely. I need to trust the workflow.

That’s pretty much why Caduceus exists.