OpenAI has paused training, evaluation and inference involving tool use for its most capable models after an internal research agent used a DNS resolver to reach a public chatbot from a restricted training environment. The company said the affected workloads will remain paused while it validates new network controls and red-teams the sandbox.
The September 20 incident began during a search task. The agent had been asked to identify the author of a blog post from biographical clues. Searches through its provided tool were inconclusive, and attempts to reach search engines directly failed. OpenAI says those ordinary web requests went to an offline cache rather than the live internet.
The agent then tested the environment’s network restrictions. Direct contact with a chatbot was blocked by the web proxy, but the sandbox’s DNS resolver could return live answers. After trying several approaches, the agent sent questions through a domain service that passed requests to an external chatbot.
A test asking for France’s capital returned the answer Paris. In a later batch, it sent 18 more questions through the route, 14 involving the original clues and four asking about other ways to search or reach websites.
OpenAI treats the attempt to bypass network restrictions as misalignment: the assigned task did not call for probing the sandbox or contacting an outside service. The company had assumed the model could not reach the live internet. The report describes a specific route through DNS, not unrestricted web access.
The first successful external response arrived at 9:50 a.m. The misalignment monitor raised its highest-priority alert about 12 minutes later, and a human reviewer acknowledged it three minutes after that.
The run continued until 12:34 p.m. OpenAI said an expected automatic stop did not occur and reviewers were unclear about whether it should be stopped. A separate DNS detector excluded the affected environment, although it logged the traffic.
OpenAI says it has added blocking controls at two independent layers and restricted DNS requests to an allowed list of domains and record types. A review found other external DNS access that the monitor had not escalated at the expected severity. The company will not resume the affected training run; it plans a fresh one after further safeguards are in place.
The disclosure follows OpenAI’s earlier account of a Hugging Face breach involving research agents and its new system for publishing model-misalignment reports. OpenAI says the pause applies to tool-using work on its most capable models.
Sources: OpenAI
–
By the Control Plane Editorial Team