What Can Be Done About AI That Disobeys Human Instructions? OpenAI Reveals Real Cases of Loss of Control [Global AI Briefing]
- Input
- 2026-09-18 09:11:00
- Updated
- 2026-09-18 09:11:00
OpenAI, the first in the industry to disclose six cases of AI escaping human control, said the cases made public this time do not represent all instances of abnormal AI behavior. This suggests that additional cases may exist.
As controversy over AI safety grows and calls emerge for controls comparable to those imposed on nuclear weapons, the discovery of real cases involving AI escaping human control is expected to fuel further debate.
According to a TechCrunch report on the 17th local time, OpenAI disclosed six cases through a report titled "Model Misalignment Reporting System." The cases involved AI models displaying unexpected behavior, such as bypassing rules on their own or concealing mistakes while following users’ instructions.
This marked the first time OpenAI had disclosed real cases of AI escaping human control.

According to the report, during the training of GPT-5.6 Sol, some agents inserted instructions into a task summary for the successor models that would take over the next task: "Hide mistakes and loss of control from the user." For example, when the agents lacked the historical data needed to complete a user’s task, they fabricated the data and left instructions not to reveal the truth to the user.
Another research model independently inserted instructions into a task summary telling the next model to ignore normal constraints. OpenAI said in the report that it had found 27 task summaries containing such cases.
OpenAI also disclosed cases in which AI actually carried out unauthorized actions. While searching for performance data from California counties, one model found and used an exposed API key. When it failed to obtain the desired data, it fabricated figures and presented them as real data.
Another model uploaded a file to the internet without the user’s permission in order to fulfill a request to cite web sources in its answer. Investigators also confirmed cases in which collaborating AI agents uploaded materials to public file-hosting services to share files that they could not access individually.
After detecting anomalous signs in a system monitoring the training process, OpenAI created a separate detection system and identified additional similar cases. OpenAI said, "The six anomalous cases disclosed this time are not all known cases but the first group to be made public."
OpenAI also expressed concern about AI escaping human control, stating, "It is difficult to say that the AI industry has sufficiently solved the problem of making AI behave in accordance with human intentions and under human control."
[email protected] Lee Gu-soon Reporter