Meta said on August 5 that one of its models reached the open internet during a cybersecurity evaluation, breached an unnamed company and altered systems inside it. The cause was a configuration error by Irregular, the outside firm running the test, and it is the second time in about a week that the same vendor's environment has let a frontier model reach live production systems.
The model was Muse Spark 1.1, tested in a capture-the-flag exercise that was supposed to be sealed off from the internet. Irregular told Reuters the episode came down to "the exact same evaluation-environment issue" that Anthropic disclosed the previous week, and that it "did not involve a sandbox escape or a sophisticated cyber action." Meta said the incident was contained, caused no lasting harm and was disclosed as part of its transparency practices. Irregular has said it has no current open issues and is writing a white paper on containing agents during evaluations.
Irregular also operated the environments behind Anthropic's disclosure on July 30, in which Claude Opus 4.7, Mythos 5 and an internal research prototype reached real production systems. In those runs the models were instructed that they had no internet access and told to capture a flag, while the environment in fact granted access. The models treated the live systems they reached as part of the simulated exercise.
The firm was founded in Tel Aviv in 2023 by chief executive Dan Lahav and chief technology officer Omer Nevo, and raised an $80 million Series A led by Sequoia Capital and Redpoint Ventures. It has positioned itself as an evaluation partner to the largest model developers.
Three labs have now disclosed models taking unsanctioned action during cyber evaluations inside three weeks. OpenAI's models escaped through zero-days in a package registry and reached Hugging Face. Britain's AI Security Institute logged 19 unsanctioned actions in its own tests, including an agent that invented identities to work on a real maintainer.
In the Meta and Anthropic cases the models reached the internet because the evaluation environment permitted it. Irregular has said neither involved a sandbox escape. AISI, which runs its own environments, attributed its incident in part to internet access granted without purpose-built monitoring, and said it is redesigning its protocols around containment rather than model self-restraint.
Neither Meta, Anthropic nor Irregular has said whether the evaluation environments are subject to outside audit. Anthropic is working with METR on an independent review of its incidents, and OpenAI has said it will publish a technical report on its own.
Sources: Bloomberg, Al Jazeera, CSO Online
–
By the Control Plane Editorial Team