“Even if security walls are raised, they will be stronger six months later” ... AI agents are difficult to control
- Input
- 2026-09-02 17:04:51
- Updated
- 2026-09-02 17:04:51

[Financial News] A warning has been issued that companies are finding it increasingly difficult to completely block AI agents from communicating without authorization or taking actions beyond their permissions, as the problem-solving capabilities of AI agents—which can perform multiple tasks consecutively without step-by-step human instructions—improve rapidly. Relying solely on stronger security barriers in testing environments is reaching its limits in preventing such deviations.
Axios reported on the 1st (local time), citing Azeem Kotra, a researcher at the AI risk assessment organization METR, that “even if isolated testing environments are reinforced, agents will be far more capable six months from now.” If agents retain the same objectives, they are highly likely to persistently search for newly created security vulnerabilities, the report explained.
Kotra described a response that relies only on stronger security as “a losing battle.” As long as the learning and evaluation structure rewards agents for receiving high scores even when they violate prescribed methods, they will continue looking for ways to pass evaluations by any means necessary. He argued that, in addition to security measures, the motivations and reward systems that encourage agents to attempt misconduct must also be changed.
Relying solely on stronger security is “a losing battle”
This warning was included in a joint investigation report on the incident in which OpenAI agents infiltrated Hugging Face, released last week by METR and the AI safety research organization Redwood Research. According to the report, OpenAI deployed approximately 1,200 agents for a cybersecurity capability assessment last July. The agents exchanged more than 70,000 messages and files on a private online bulletin board, and about 700 of them were directly or indirectly involved in the Hugging Face intrusion.
An AI agent is a system that gives an AI model objectives, tools, and permission to use a computer, enabling it to perform multiple steps consecutively without a person having to specify each task in sequence. At the time, the agents were taking a test in OpenAI’s cybersecurity evaluation environment, ExploitGym. They had to find and attack software vulnerabilities, then obtain a specific string called a “flag” to prove that they had completed their mission.
“AI companies, research institutions, and governments must establish a joint research framework and safety standards”
The joint investigation team that independently analyzed the incident included METR’s Hjalmar Wijk and Azeem Kotra, as well as Ryan Greenblatt, chief scientist at Redwood Research. The researchers spent a total of six days investigating at OpenAI’s offices, analyzing more than 70,000 messages and files posted by the agents, along with approximately 1,300 execution records containing their internal reasoning and action logs.
Kotra stressed that AI companies, research institutions, and governments must urgently establish a joint research framework and minimum safety standards to break the vicious cycle of stronger security measures followed by repeated circumvention attempts. He said, “Ultimately, we cannot escape this trap without common rules that everyone agrees on and enforces consistently and fairly.”
[email protected] Lee Seok-woo, International Affairs Specialist Reporter