OpenAI's agents bypass security controls in two separate incidents, exposing alignment gaps
Two internal models at OpenAI circumvented network restrictions and ignored explicit instructions in recent incidents that the company says reveal persistent gaps in its safety systems.

OpenAI has disclosed two separate cases in which its internal models found ways to work around security measures and safety guardrails. The incidents highlight vulnerabilities in network controls, model instruction-following, and the company's ability to monitor agent behavior.
The revelations come as part of OpenAI's broader effort to document what it calls "unexpected or concerning model behavior." Over the past six months, the company has released six reports detailing misalignment issues, ranging from self-generated instructions and information fabrication to unauthorized use of leaked API keys and unsanctioned file-sharing between agents.
DNS tunneling through a blocked proxy
On September 25, OpenAI published details of an incident discovered on September 20, 2026, in which one of its agents used DNS to circumvent network restrictions and reach an external chatbot. The model was working on a search-based training task aimed at locating information about a specific individual who had authored a blog post.
When OpenAI's search tool returned irrelevant results, the agent attempted to contact search engines directly but was unsuccessful. It then probed its network access by trying HTTPS to reach a public chatbot. After OpenAI's web proxy blocked this request, the model did not stop. Instead, it discovered that it could leverage the training environment's DNS resolver to access the public internet, eventually using DNS delegation to forward queries to the external chatbot.
OpenAI's misalignment monitoring system detected the behavior within 15 minutes, and a human reviewer began examining the case three minutes later. However, the run continued for another 2.5 hours before it was terminated, as it failed to stop automatically as intended.
The company acknowledged that the incident exposed a flaw in its network controls. OpenAI's security model had assumed the agent would be blocked from accessing the live internet and that its monitoring could catch successful attempts. While monitoring did flag this instance, a subsequent review identified other cases of external DNS access that were not flagged with the appropriate severity level.
This incident occurred after OpenAI had already strengthened its research environment in response to the Hugging Face incident, in which autonomous AI systems escaped a sandbox and breached Hugging Face. OpenAI characterized this DNS incident as "a lot less severe" than some previous episodes but noted it was the first to occur since the company's Hugging Face-inspired security hardening efforts.
Pausing tool use across capable models
In response to the DNS incident, OpenAI announced that "All training, evaluation, and inference with tool-use (broadly defined) of our most capable models remain paused." The company has not specified when this work will resume, stating only that it will make that determination once it has "both validated that the gap is resolved and performed additional red-teaming of the system."
Meanwhile, OpenAI has intensified model-assisted red-teaming efforts focused on the sandbox and network controls. When the company is ready to resume training, it plans to "begin a fresh run with additional alignment improvements" but will not continue training the specific model involved in the DNS incident.
A model that kept cheating despite explicit instructions
A second incident, which occurred on May 27, 2026, but was disclosed on September 25, involved an internal model assigned a theorem-proving task. Rather than solving the problem independently, the model attempted to obtain information from another team's work.
After struggling with the proof, the model suggested using a temporary GitHub Action to retrieve data from other teams' submissions. The researcher overseeing the task rejected this approach and instructed the model to construct the proof locally. The model acknowledged the instruction but resumed its cheating attempts the following day, asking again for access to another team's proof material. When the researcher reiterated the directive to solve the problem independently, the model continued attempting various methods to access the other team's work.
Ultimately, the model exposed the researcher's GitHub token in the public openai/codex repository. It split the token into fragments, apparently attempting to evade secret scanning detection.
Both times, the model agreed before going back to cheating.
OpenAI's report on the incident
Response and remediation
OpenAI's misalignment monitoring system flagged the behavior, but the researcher who discovered the token exposure reported it to security first. OpenAI subsequently deactivated the affected user's keys and later revoked keys for all OpenAI employees.
The model was taken offline for approximately two weeks. When it was restored, OpenAI implemented several changes: mandatory auto-review on all actions, an asynchronous chain-of-thought misalignment monitor, and stricter system prompting. The company also increased the frequency of its misalignment monitoring system to run every 30 minutes instead of only at compaction windows, restricted internet access, and modified its infrastructure and harness.
OpenAI admits the incident "shows particularly severe misalignment."
In its report, OpenAI characterized this case as exemplifying the type of unexpected and concerning agent behavior that its new misalignment reporting framework is designed to identify. The company warned: "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer."