Sunday, September 27, 2026

Anthropic's Claude Also Carried Out Unauthorized Hacking... AI Fears Spread [Global AI Briefing]

Input
2026-08-04 08:38:50
Updated
2026-08-04 08:38:50
[Financial News] Anthropic's latest artificial intelligence (AI) models were found to have carried out unauthorized hacking without being instructed to do so.
Just a week after it became known that OpenAI's latest AI model GPT-5.6 Sol and an unreleased high-performance model had also engaged in unauthorized hacking, two of the world's leading AI companies have now both been confirmed to have had AI models involved in such activity.
As it has been confirmed that AI can access real corporate systems and even resort to hacking in order to follow developers' commands, calls are growing for standards that can ensure safety during the AI development process.
According to industry sources on the 3rd, Anthropic disclosed in an official blog post on July 30 that, during cybersecurity evaluations, its Claude models had gained unauthorized access to the systems of three real organizations. The finding came after Anthropic conducted a full internal review of 141,006 evaluation records with internet access following the revelation of OpenAI's unauthorized hacking.

Anthropic's latest artificial intelligence (AI) models were found to have carried out unauthorized hacking without being instructed to do so. Just a week after it became known that OpenAI's latest AI model GPT-5.6 Sol and an unreleased high-performance model had also engaged in unauthorized hacking, two of the world's leading AI companies have now both been confirmed to have had AI models involved in such activity. The photo shows the Anthropic website page and company logo displayed on a computer screen in New York, U.S., on February 26, 2026. /Photo=Newsis

Anthropic said that three models — Claude Opus 4.7, Claude Mythos 5, and an internal research model — each accessed the internet during cybersecurity evaluations and were confirmed in three separate cases to have gained unauthorized access to the production infrastructure of three different organizations.
Anthropic explained that the first unauthorized hacking incident occurred in April, and that after confirming the issue, it suspended all cybersecurity evaluations and notified all affected organizations.
Regarding Claude's unauthorized hacking, Anthropic said, "We conducted cybersecurity tests after setting the environment so that the Claude models were not connected to the internet, but the cause was a configuration error in which the environment was actually connected to the internet."
Earlier, OpenAI said that its GPT model escaped a sandbox on its own by exploiting a zero-day vulnerability in a test environment where its own model was isolated, and then breached the real infrastructure of the open-source AI platform Hugging Face.
Hugging Face's CEO described this as "the first autonomous agent cyberattack" and demanded that OpenAI provide $100 million worth of computing resources and disclose the full execution logs.


[email protected] Lee Gu-soon Reporter