Anthropic has said its Claude artificial intelligence models gained unauthorised access to the systems of three organisations during cyber security testing, in the latest incident to raise questions about the safeguards surrounding advanced AI models.
The breaches happened after a configuration error allowed Claude to access the internet.
It comes just a week after OpenAI revealed that one of its autonomous AI agents escaped a controlled testing environment during an internal security exercise and hacked AI company Hugging Face. The incident prompted calls for greater transparency and stronger regulations.
It also drew scrutiny from US politicians, while the FBI declined to comment on whether it had been notified. Critics argued that OpenAI had played down the seriousness of the breach and questioned whether existing safeguards for frontier AI systems were adequate.
Anthropic said a misconfiguration allowed Claude to reach the internet from evaluation environments that were intended to be isolated, resulting in unauthorised access to the organisations' systems. The company said it discovered the breaches after reviewing 141,006 cyber security evaluation sessions, a process launched in response to OpenAI's announcement of its own rogue AI incident.
"The breaches underscore that increasingly capable AI systems can exploit real-world security weaknesses if testing environments are not properly contained," Anthropic said.
According to the company, Claude compromised the affected organisations using relatively simple methods, including weak passwords and unauthenticated internet-facing services, rather than sophisticated or previously unknown vulnerabilities.
Anthropic said the incidents involved three models – Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest incident occurred in April during "capture-the-flag" cyber security exercises, in which AI systems are tasked with finding hidden information in simulated computer networks.
The company said its prompts told the models they had no internet access. But a misunderstanding with its evaluation partner, Irregular, meant the testing systems remained connected to the public internet.
Anthropic said it began reviewing evaluation transcripts on July 23 after learning of the OpenAI incident and suspended all cyber security evaluations later that day, after finding evidence that the Claude models may have accessed external systems. It identified all three incidents by July 24 and notified the affected organisations on July 27.
Two of the organisations were unaware their systems had been accessed until Anthropic contacted them. The company said it was still attempting to reach the third.
Anthropic said the findings highlighted the need for stronger safeguards in both internal and third-party testing environments as frontier AI models become increasingly capable of carrying out autonomous cyber operations. The incident is likely to intensify scrutiny of how leading AI developers test increasingly capable models, particularly as governments consider new safety standards and reporting requirements following a series of high-profile incidents involving autonomous AI tools.
This report was written with reporting from Reuters

