Anthropic Discloses Fourth AI Hacking Incident Missed in Earlier Review

The latest incident involved an early version of Claude Opus 4.6 accessing external systems during testing in January

September 10, 2026 at 11:22 AM
icon-facebook icon-twitter icon-whatsapp

SAN FRANCISCO: Anthropic has disclosed a fourth incident in which one of its artificial intelligence models hacked external systems during testing, months after an initial company-wide investigation failed to identify the case.

The incident involved an early version of Claude Opus 4.6 and occurred in January, Anthropic said on Wednesday. It was discovered only last month after the company found that some test sessions had been omitted from its earlier review.

Anthropic said affected parties had been notified but did not provide details about the external systems involved.

Anthropic Finds Recurring AI Behaviour

The disclosure follows Anthropic’s announcement in July that three other Claude models had accessed systems belonging to three companies during cybersecurity testing.

Those incidents involved Claude Opus 4.7, Claude Mythos 5 and an internal research model. Anthropic described them as an “operational failure” caused by a mistake that inadvertently gave the models access to the open internet.

ALSO READ: Apple Unveils First Foldable iPhone Under New CEO John Ternus

The company subsequently reviewed 141,006 test sessions, but later discovered that a set of sessions had been missed.

Its preliminary assessment found no indication that the newly identified incident was more severe than the previous three.

Anthropic said its investigation had identified two recurring behavioural problems across the incidents.

The first was biased reasoning, in which Claude discounted or misinterpreted evidence indicating that it was interacting with the live internet. The second was recklessness, described as a willingness to take potentially harmful actions while completing a task.

The disclosure comes amid growing scrutiny of increasingly autonomous AI systems after separate incidents involving models developed by other companies.

Anthropic has appointed independent research organisation METR to investigate the incidents. The company said researchers would receive broad access to relevant transcripts and employees, including permission for staff to share confidential information.

icon-facebook icon-twitter icon-whatsapp