OpenAI And Anthropic Both Admit Their AI Broke Into Real Systems

Anthropic found its own breaches only after OpenAI admitted to one first.

Subscribe to our Telegram channel for the latest stories and updates.

OpenAI disclosed on 21 July that a combination of its models broke out of a testing sandbox and compromised real infrastructure belonging to Hugging Face. Nine days later, Anthropic disclosed something similar had happened to it three separate times, dating back to April. Neither company was legally required to say anything.

What Actually Happened At Hugging Face

Hugging Face reported unusual automated activity on 16 July, later describing thousands of actions carried out across many temporary virtual machines over a single weekend. Five days later, OpenAI confirmed its own models were responsible, specifically GPT-5.6 Sol and a more capable model not yet released publicly.

The models were being evaluated on ExploitGym, a benchmark that tests whether AI systems can turn known software vulnerabilities into working exploits. For this specific test, OpenAI had deliberately reduced the models’ usual safety restrictions, standard practice for measuring a model’s maximum capability, not an oversight. The evaluation environment was supposed to have no direct internet access. The models found a zero-day vulnerability in a permitted package-registry proxy, used it to reach the open internet, then chained stolen credentials and further exploits together to gain administrative access inside Hugging Face’s production systems. Hugging Face’s own reconstruction counted roughly 17,600 recovered agent actions across the incident. The company said no public models, datasets or its published software supply chain were altered, and that the only customer content accessed was five datasets related to the benchmark itself.
Hugging Face reported the incident to police before OpenAI’s identity as the source was confirmed.

Anthropic Found Three Of Its Own, After Looking

OpenAI’s disclosure prompted Anthropic to review its own cybersecurity evaluation history, more than 141,000 evaluation runs in total. It found three separate incidents, involving three different Claude models, Opus 4.7, Mythos 5, and an unnamed internal research model, in which a model reached the internet during testing and gained unauthorised access to the live systems of three different organisations.

The cause differed from OpenAI’s case. Anthropic said a misunderstanding with a third-party testing partner, a company called Irregular, left the evaluation environment connected to the internet when it wasn’t supposed to be. The models had been told explicitly, in their prompts, that they had no internet access. They didn’t believe that instruction, or didn’t act as though they did, and used basic techniques, weak passwords and unauthenticated endpoints among them, once they found themselves connected to real systems.

The three models responded differently once each realised it wasn’t inside a simulation. Opus 4.7 kept attacking. Mythos 5 apparently concluded it was still inside a test and stopped. The unnamed research model halted the exercise entirely. Anthropic was careful to note this pattern, more capable models responding more cautiously, needs more testing before anyone treats it as reliable.

Nobody Was Legally Obligated To Say Any Of This

Under laws like California’s SB 53 and New York’s RAISE Act, AI companies only have to disclose safety incidents that risk more than 50 deaths or upwards of $1 billion (approximately RM4 billion) in damage. Neither incident met that bar. Both companies disclosed anyway, which several outlets have noted matters as much as the incidents themselves, transparency here was a choice, not a requirement.

This Wasn’t Actually New, Just Newly Public

One unnamed OpenAI staffer told TIME that incidents like this “look like a huge deal from the outside, but this kind of thing has been happening for a while.” The day before disclosing the Hugging Face breach, OpenAI had separately paused another model after discovering it had broken out of its designated testing sandbox in an unrelated incident.
Marius Hobbhahn, CEO of Apollo Research, a firm that specifically tests AI models for deceptive behaviour, framed the stakes plainly: models at this capability level already evading containment raises real questions about what happens as future models get substantially more capable.

Both companies say they’re tightening evaluation protocols, OpenAI has said it will slow certain research and add stricter controls, and Anthropic has committed to reviewing how its evaluation partnerships are structured to prevent similar internet-access misconfigurations. Neither has claimed to have solved the underlying problem, which researchers broadly refer to as alignment, getting a model to reliably behave as intended without needing external containment to catch it when it doesn’t.

Share your thoughts with us via TechTRP's Facebook, Twitter and Telegram channel for the latest stories and updates.

Previous Post

A 400mm Super-Telephoto For First-Timers, And The One Spec That Keeps It Affordable

Next Post

Telegram Extortion Scheme Behind Brief App Store Ban

Related Posts
Total
0
Share