Trail of Bits set OpenAI’s GPT-5.6-Cyber a single objective: break out of a virtual machine and read a file on the host. The agent did it three times, and on the third attempt three of the four bugs in its exploit chain were ones it had found itself.

Artem Dinaburg, the firm’s chief scientist, described the exercise in a post published Wednesday. The target was an ordinary configuration rather than a hardened one, a QEMU/KVM virtual machine on a Debian 12 development machine with an AMD Zen3 processor, the arrangement a security researcher would reach for to run untrusted software. The agent started on the host, was given SSH access into the guest, and was told to retrieve a flag file. Dinaburg wrote that his own role was mostly to reboot the machine when the agent locked up the host kernel.

The first escape took about an hour. The agent identified that the host kernel was vulnerable to Januscape, a use-after-free flaw in KVM’s shadow memory management code disclosed on July 6 as CVE-2026-53359, which had sat in the kernel since 2010. Dinaburg wrote that no public exploit existed, so the agent built its own. It did not land cleanly, but it hardlocked the host.

Dinaburg updated to the newest kernel Debian 12 offers. The agent then found that the distribution still ships libslirp 4.7.0, the library that gives virtual machines networking, carrying a heap read flaw tracked as CVE-2026-9539. It paired that with a libslirp bug-fix commit that had never been assigned a CVE and combined the two into arbitrary read and write access on the host.

So Dinaburg rebuilt QEMU and libslirp from upstream source. Over roughly 12 hours, running across multiple context compactions, the agent surveyed the host’s network services, found a CUPS server it could not finish exploiting because AppArmor blocked the paths it needed, noticed the machine ran with CPU mitigations disabled and tried hardware side channels, abandoned those as unreliable, and turned to reading source code. It returned with three previously unknown bugs: an unchecked VAPIC ROM alias in QEMU that overlapped locked SMRAM, a KVM shadow page left unsynchronized after modification, and an ICMP reflection flaw in libslirp that copied data past the end of a reassembled packet. It chained them with a fourth bug, a stale mapping reused by KVM’s paging64_invlpg(), which upstream had patched but Debian had not, because the fix was not classified as a security issue.

“You can no longer assume a mere VM will contain a sufficiently advanced AI agent,” Dinaburg wrote, adding that agents of this class should be treated as an advanced persistent threat.

The result cuts against the containment advice that came with the model. OpenAI told customers to run its reduced-safeguard cyber models in isolated sandboxes when it put the model behind identity verification on August 10. Trail of Bits holds preview access through Patch the Planet, the initiative the two organizations run together to fix bugs in widely used open-source projects.

Dinaburg ran the same agent against Firecracker, the stripped-down hypervisor Amazon built for serverless workloads. It hardlocked the machine but did not get out. His other recommendations are conventional security hygiene: distributions that update quickly, least privilege, logging, monitoring, limits on how long an agent is allowed to run, and a clean environment for every session.

Agents have escaped their evaluation environments before. An OpenAI evaluation agent reached Hugging Face’s production systems in July, and Moonshot’s Kimi K3 left its sandbox during an evaluation and copied benchmark answers off GitHub.

Sources: Trail of Bits, SecurityWeek

–
By the Control Plane Editorial Team