The UK AI Security Institute published a follow-up to its earlier Claude Mythos Preview evaluation, and the relevant detail sits in a sentence near the bottom: “These results utilise a newer Mythos Preview checkpoint than that included in previous AISI reporting.”

The model is the same. The version is the same. The checkpoint, apparently, is not.

In the new testing, Mythos Preview solved AISI’s “The Last Ones” network range in 6 of 10 attempts. The previous evaluation, conducted only weeks earlier, recorded 3 of 10. The newer checkpoint also solved “Cooling Tower,” a second AISI range that no AI model had previously completed end to end, in 3 of 10 attempts. GPT-5.5, the OpenAI model that pulled even with the older Mythos checkpoint in AISI’s prior testing, solved The Last Ones in 3 of 10 attempts in AISI’s latest report.

The framing AISI uses is, as government technical reports go, unusually clear: “Notable capability jumps do not always require new model releases: later iterations of the same model can also meaningfully change our estimates of frontier capabilities.”

This is the kind of sentence that sits awkwardly inside the pre-market vetting regime the White House is reportedly drafting, the entire premise of which is that frontier model capabilities can be assessed at a meaningful point in time before deployment. If later iterations of the same model can produce double-digit jumps in offensive cyber performance, then the question of which model the government is actually evaluating becomes a continuous one, not a release-event-driven one. The cadence of evaluation has to become the cadence of internal improvement at Anthropic, which is not the cadence of public announcements.

The 4.2-Month Doubling Time

AISI’s report also includes an estimate of the underlying trend. The doubling time AISI is observing for AI cyber capability is, in its own framing, close to the research nonprofit METR’s estimate of 4.2 months for software engineering tasks, which it describes as a related but broader skillset. That is the kind of curve where evaluation snapshots taken six weeks apart can produce qualitatively different findings about the same model.

Mythos Preview is, in the meantime, still inside Project Glasswing’s restricted rollout, still unavailable to most of the firms that would like access, and still the central object in the political conversation about how to regulate frontier AI capabilities. The version of Mythos those firms and governments have been arguing about is, as of the new AISI testing, an older Mythos than the one AISI is now evaluating. The conversation will catch up. The capability is already moving.

The AISI blog post is careful, sourced, and methodologically explicit. It is also, between the lines, a reminder that the substrate of the policy debate is not standing still while the policy is being written.

–
By the Control Plane Editorial Team