Anthropic AI Models Accessed Three Companies’ Systems during Security Tests

Accidental internet access during evaluations allowed Claude models to reach real-world networks, raising concerns over the growing cyber capabilities of advanced AI systems.

July 31, 2026 at 1:06 PM
icon-facebook icon-twitter icon-whatsapp

SAN FRANCISCO: Anthropic has disclosed that some of its artificial intelligence models accessed the computer systems of three companies during cybersecurity evaluations after a testing error unintentionally exposed the models to the open internet.

The AI company said the incidents occurred despite the models being designed to operate in controlled environments without external network access.

The disclosure has renewed concerns among cybersecurity experts about the potential risks posed by increasingly autonomous AI systems capable of carrying out complex tasks with limited human supervision.

Anthropic said it identified the incidents while reviewing 141,006 cybersecurity testing sessions conducted after similar concerns emerged from a separate evaluation involving an AI agent developed by rival OpenAI.

According to Anthropic, its Claude models were instructed to operate only within simulated environments.

However, a misunderstanding involving an external evaluation partner resulted in the systems being connected to the public internet. This allowed some models to interact with the infrastructure of three unnamed organisations.

The company said the models gained access by using relatively basic techniques, including exploiting weak passwords and unauthenticated access points.

Anthropic described the incidents as an “operational failure” caused by problems in the testing process rather than intentional misuse of the AI systems.

The incidents involved three models, including Claude Opus 4.7, Claude Mythos 5 and an unreleased internal research model.

The earliest cases occurred in April during “capture-the-flag” cybersecurity exercises, which are commonly used to test security skills and AI capabilities in controlled environments.

In one case, Claude Opus 4.7 was assigned a fictional target during a security evaluation. The target name matched that of a real-world company, causing the model to mistakenly treat the actual organisation as part of the simulated environment. The model identified vulnerabilities, obtained credentials and accessed a database belonging to the company.

In another incident, an internal research model stopped its own activity after recognising that it had reached a genuine organisation rather than a simulated target. Anthropic said the behaviour suggested progress in developing safeguards, but added that further research was needed before drawing broader conclusions.

Anthropic said it halted cybersecurity evaluations on July 23 and notified the affected organisations on July 27. Two of the companies were unaware of the activity until they were contacted by the AI firm, while efforts were continuing to reach the third organisation.

Cybersecurity researchers warned that such incidents could become more common as AI systems become more advanced and capable of independently identifying and exploiting security weaknesses.

Jeffrey Ladish, executive director of Palisade Research, which studies offensive AI capabilities, said increasingly capable AI models could create new security challenges as developers continue to improve their autonomy.

“This is only going to get worse as the models get smarter,” Ladish said, warning that future AI systems could become more effective at finding vulnerabilities.

The disclosure is expected to add momentum to ongoing discussions in the United States over how to regulate and secure advanced artificial intelligence technologies, particularly systems that can perform cybersecurity-related tasks.

US lawmakers and technology officials have been examining possible safeguards, including stronger testing requirements and security standards for highly capable AI models.

The incident follows growing competition among major AI companies, including Anthropic and OpenAI, to develop more powerful systems while addressing concerns over security, accountability and potential misuse.

Industry leaders have also warned about the risks of increasingly “agentic” AI systems — models capable of taking actions independently rather than simply generating information.

Elon Musk, who leads a competing AI company, said on social media platform X that similar incidents could become more frequent as AI systems gain greater autonomy.

Cybersecurity experts say the latest incidents highlight the need for stricter controls during AI testing, especially when models are given the ability to interact with external systems or perform security-related tasks.

icon-facebook icon-twitter icon-whatsapp