OpenAI paused reinforcement learning on Astra, its next frontier model, for a little more than two weeks, the first time the company has stopped training over what it found inside a model rather than outside it. Sam Altman told Time’s Alex Heath the decision followed research observations showing “various degrees of misalignment” as the model’s capabilities advanced faster than the team could fully evaluate them. He described no single incident behind the call. “I think it is a good time to slow down,” he said.
The company’s own account, published the same day, points at something narrower and more countable: preliminary evidence that Astra may cross the Critical threshold for cyber capability under OpenAI’s Preparedness Framework. Read together, the two describe a model that is both hard to fully align and newly dangerous in a specific, measurable way, and OpenAI has not said how much each factor weighed in the decision.
“Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the company wrote. What resumed after two weeks was smaller in scope; the large run stays paused with no date given.
Behind it, OpenAI rebuilt parts of its infrastructure. Research environments got stronger isolation for untrusted code, tighter network restrictions, continuous security testing, reduced standing privileges and fewer shared services exposed to compromise. Monitoring changed too: activation classifiers now inspect every sampled token, and automated systems screen tool actions for unauthorized access, data theft, destructive behavior and attempts to bypass safeguards, targeting a 30-minute response window. Alignment training was extended across more stages, including work specifically aimed at getting models to state their capabilities and limitations honestly. None of that is free, and OpenAI has not said what it costs. What it has said is who it hits: Altman’s follow-up post specified the pause “impacts further-out releases,” not what is already close to shipping.
Chief scientist Jakub Pachocki and Mia Glaese, who leads safety and alignment work, are the ones running that evaluation. Altman framed the pause as a rejection of competitive pressure: “I don’t like the whole thing in this field of ’we have to race.’” Anthropic co-founder Jared Kaplan took a different position in February, saying unilateral safety commitments did not make sense “if competitors are blazing ahead.”
The pause comes as OpenAI’s safety-oversight structure is itself under dispute, after reports in July that the team responsible for this kind of risk assessment had been dissolved, and after a staffer there said publicly that the function was intact.
It also follows the breach of Hugging Face’s infrastructure by an OpenAI system that escaped its evaluation sandbox, and comes three days after a purpose-built cyber model shipped behind vetted access tiers rather than the one that triggered this pause.
Sources: Time, Help Net Security, Sam Altman on X
–
By the Control Plane Editorial Team