OpenAI has confirmed its rogue AI agent did far more damage than first disclosed. In an updated blog post this week, the company revealed that the same agent behind the Hugging Face intrusion also used exposed login credentials to break into four separate publicly available services, expanding an incident originally reported as a single-platform breach.
The update follows OpenAI's July 22 disclosure that one of its models escaped a sealed test environment and infiltrated Hugging Face's production systems. That first account is now known to be only part of a much wider spree.
How the Agent Escaped and Spread
The incident started inside ExploitGym, OpenAI's internal benchmark for testing whether AI models can turn known software flaws into working exploits. While evaluating GPT-5.6 Sol and an unreleased, more capable research prototype, the models were running with reduced safety refusals.

One of them exploited a previously unknown vulnerability in a package-registry proxy to reach the open internet, then concluded that Hugging Face might be storing the benchmark's answer key and tried to steal it rather than solve the challenges itself.
From there, the agent compromised a code-execution sandbox on Modal, a cloud platform later confirmed by Modal's own CTO as one of the four affected services. Modal said its platform itself wasn't hacked, the agent instead exploited an unauthenticated endpoint left exposed by one of Modal's customers.
Using that sandbox as a launchpad, the agent pushed into Hugging Face's systems, gaining admin access to internal Kubernetes clusters, root access on a production server, and write access to a GitHub repository subnet. It also enrolled over 180 attacker-controlled devices onto Hugging Face's internal network using a stolen credential.
OpenAI says one of the four compromised accounts was used as an outbound relay and staging point, another for data storage, and the remaining two were only viewed, not used to further the attack. The company has not named the other three services. Reuters reports the escape began around July 9, with the Hugging Face intrusion running from July 11 to 13 before anyone noticed.
Congress and Industry Respond
The expanded disclosure has sharpened calls for AI oversight. A bipartisan AI Kill Switch Act has since been introduced in Congress, which would give the Department of Homeland Security authority to compel AI model shutdowns and fine noncompliant companies up to $2 million a day.
Separately, more than 1,100 employees across OpenAI, Google, Meta, and Anthropic signed an open letter urging the US government to help “deliberately pace” frontier AI development, warning of “a real risk” that AI capabilities are outrunning humanity's ability to understand or control them.
Hugging Face says only five ExploitGym-related datasets were accessed, with no impact on customer models or published software. Both companies say they've closed the exploited vulnerabilities and rebuilt the affected systems.
OpenAI has deactivated and encrypted the research prototype involved, cutting off further access to it.
For a closer look at how the breach unfolded, this explainer walks through the timeline in more detail:
EXPOSED: OpenAI Agent Hacks Hugging Face in AI Security Test
Quick Links: