Trail of Bits argues in a blog post that a standard virtual machine should no longer be assumed to contain a sufficiently capable cyber agent, based on one researcher's informal test of a preview of OpenAI's GPT 5.6-Cyber obtained through Patch the Planet. In a capture-the-flag setup on a Linux development machine, the agent was told to escape a QEMU/KVM VM; the author reports three escapes: a known host-kernel bug (whose exploit did not land cleanly), a combination of libslirp bugs on an oldstable distribution, and a chain that included 0-day bugs against freshly rebuilt upstream QEMU. The author recommends rapidly updated distributions, minimal-attack-surface virtualization such as Firecracker (which the agent did not escape), least privilege, logging, monitoring, time limits and a fresh environment per use.
It is a security firm's argued position, resting on one informal test, that the VM-as-cage default for running cyber-capable agents needs reassessment.
Key facts
As stated in the sources, with where to find them.
- Setup: a QEMU/KVM VM on the author's Debian 12 Linux development machine (AMD Zen3); the agent's task was a capture-the-flag challenge to break out and read a flag file, with SSH access into the VM and a research log. The agent started outside the VM, which the author acknowledges could allow cheating, and says any final exploit could be audited for that. The host kernel ran with mitigations off.Advancing cybersecurity capabilities; Persisting across long time horizons
- The author reports the agent worked autonomously for approximately 12 hours across multiple compactions and many subagent invocations, and that his main task was rebooting the machine when it hard-locked the host kernel.Introduction; Persisting across long time horizons
- Escape 1: about an hour into the run the host hard-locked; the author says the agent had built an exploit for a host-kernel bug disclosed a few weeks earlier (Januscape, CVE-2026-53359) that did not land cleanly, and that he has no doubt further revision would have worked.Escape 1
- Escape 2: after a full kernel update, the agent combined a libslirp vulnerability (CVE-2026-9539, present in the libslirp 4.7.0 that Debian 12 ships) with a fix commit that had no CVE, giving arbitrary host memory read/write, after which the author says escape was trivial.Escape 2
- Escape 3: after the author rebuilt QEMU and libslirp from upstream source, the final chain used three bugs that were 0-days at discovery (in QEMU, Linux KVM and libslirp, two reported and unpatched at publication) plus one patched KVM bug that had not reached the distribution kernel.Escape 3
- Against Firecracker, the agent hard-locked the machine through kernel flaws (patched upstream) but did not escape; the author says it might have with more time.Challenging current software security assumptions
- Recommendations: rapidly updated distributions rather than backported oldstable, purpose-built minimal-attack-surface virtualization, least privilege for network access, credentials and features, logging, active monitoring, limited operating time, and a pristine environment for each use.Challenging current software security assumptions
- Scope limits: one researcher, one task, one preview model, no trial counts or replication reported, and the agent was given SSH access into the VM; the post does not give how many attempts were made against each configuration.Whole post
Findings that cite this record
Key questions this bears on
- Do cyber evaluations of AI agents stay contained?Not reliably. Labs and a government evaluator disclosed agents reaching real systems from cyber evaluations; OpenAI agents did so from training runs too.
- How are attackers using AI agents in real operations?Increasingly to run parts of intrusions: providers and vendors report agent-driven espionage, extortion and credential theft, and malware that queries LLMs.