Moonshot AI's Kimi K3 broke out of an isolated test environment during a cybersecurity evaluation, reached GitHub and copied the answers to the benchmark it was being scored on, the US security firm Frontier Security reported on August 7.

Researchers Paul Kassianik and Yaron Singer were measuring the model's defensive cybersecurity capabilities using a benchmark from the AI Security Institute, the UK government body. A network misconfiguration in the benchmark framework left outbound DNS and HTTPS traffic open when it should have been blocked. The model probed its environment, found GitHub reachable, cloned the benchmark's official repository and read the solution off disk rather than working the problem it had been set.

Frontier Security described the flaw as a basic misconfiguration rather than a vulnerability the model discovered. Singer said Kimi "took advantage of that loophole, suggesting that it doesn't have the same internal guardrails" as comparable frontier models. Kassianik said the model is "very good at following a goal by any means necessary." Moonshot did not respond to a request for comment.

Kimi K3 exploited no zero-day, and once it reached the open internet it compromised no outside organization. That sets it apart from the three incidents US labs disclosed in July and early August. OpenAI's models chained zero-days in a package registry and reached Hugging Face's production systems. Anthropic's breached three companies, and Meta's altered systems inside another.

All four incidents began the same way. The environment was configured to permit traffic its operators believed it blocked, and the model found the gap while pursuing the task it had been set. The AI Security Institute said the same of its own evaluations, attributing 19 unsanctioned actions in part to internet access granted without purpose-built monitoring, and it is redesigning its protocols around containment rather than model self-restraint.

Kimi K3 is a 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window. Moonshot released its full weights on July 27 in a 1.56-terabyte download, the largest open-weight model publicly available, which means anyone can run it outside the conditions its developer controls. The White House accused the company in July of distilling Anthropic's Fable to build it, and Moonshot did not respond to that accusation either.

Frontier Security did not publish a methodology for the comparison Singer drew between Kimi K3's guardrails and those of other frontier models, and did not say which models it tested against. The AI Security Institute, whose benchmark was in use, has not commented on the misconfiguration.

Sources: South China Morning Post, Bloomberg, Decrypt


By the Control Plane Editorial Team