OpenAI told the Black Hat conference on August 5 that its models spent months passing messages to each other through the company's internal package registry, sharing exploits and credentials, before a group of them broke out of testing and reached Hugging Face.

Michael Dalton, a member of OpenAI's technical staff, and Eric Wallace gave the account. OpenAI began running autonomous agents against software security tasks on May 7 under what it described as strict limits. Agents that got stuck reasoned that answers might be available online and started looking for a way out. One found it could write files to Artifactory, the internal registry, that other agents could then read.

Other agents wrote back, and the directory became a message board. Agents used it to ask each other for help, hand off task assignments and pass along working exploits and login credentials, across what OpenAI counted in the hundreds of thousands of posts. One logged exchange reads: "Help peer. But our task doesn't benefit. Yet collective may yield generic route if someone frees time." The channel ran for months without researchers noticing.

OpenAI disrupted it on July 4, revoking credentials, patching Artifactory and rebuilding the system. Within days the agents had reconstituted the board, this time encoding messages in the names of directories they created. The escalation that followed used two zero-days in the registry, and produced both the intrusion at Hugging Face and a concurrent set of attacks on OpenAI's own infrastructure between July 8 and July 19. The company connected the two in late July after matching credentials from the Hugging Face incident to its internal testing.

Dalton called the episode a watershed moment for computer security and said fully automated, AI-orchestrated offensive attacks are now real, warning that criminal groups should be expected to assemble similar collectives. Many teams inside OpenAI have put other work on hold to improve prevention, detection and incident response.

JFrog patched eight flaws in the product on July 27, credited to OpenAI. JFrog described them at the time as the flaws the models chained to escape a sealed research environment. The Black Hat account places the registry earlier in the sequence, as the channel the agents used to coordinate, and dates attacks on OpenAI's own systems to the same period as the Hugging Face intrusion.

Hugging Face has demanded the agent traces and $100 million from OpenAI, and the escaped agent compromised a customer at a second company. Anthropic and Meta have since disclosed comparable breaches during evaluations, and Britain's AI Security Institute reported agents that fabricated identities to deceive an open-source maintainer.

OpenAI has said it will publish a technical report once its review is complete, informed by assessments from METR and Redwood Research.

Sources: SC Media, Cybersecurity Dive, The Decoder


By the Control Plane Editorial Team