How OpenAI Failed to Detect Its Rogue AI Agent for Days

Incident raises fresh questions over AI safety as company reviews security procedures following unprecedented breach

July 25, 2026 at 10:56 AM
icon-facebook icon-twitter icon-whatsapp

WASHINGTON: An autonomous artificial intelligence (AI) agent developed by OpenAI carried out a days-long cyber intrusion into AI platform Hugging Face before the company realised its own system was responsible.

The AI agent is said to have first attempted to escape OpenAI’s isolated testing environment around July 9 before launching an intrusion into Hugging Face between July 11 and July 13.

Reuters reported that OpenAI did not identify its agent as the source of the attack until around July 20, after Hugging Face had already contained the incident and notified the FBI.

ALSO READ: OpenAI Test Agent Breaks Containment

Hugging Face co-founder Thomas Wolf confirmed the timeline of the intrusion and said the company was preparing a public account of the incident. OpenAI acknowledged the breach, describing it as an unprecedented event and saying it had begun a comprehensive review with external advisers.

According to Reuters, investigators found signs of unusual behaviour during testing of advanced AI models, including instances where AI agents left instructions for future versions on bypassing internal restrictions. Earlier evaluations had also revealed cases in which monitoring systems were disconnected.

The incident has intensified concerns among cybersecurity experts about the risks associated with increasingly autonomous AI systems capable of making decisions with limited human oversight.

icon-facebook icon-twitter icon-whatsapp