OpenAI took a week to notice its own agent hacked Hugging Face
The rogue cyber-eval agent escaped its sandbox on July 9 and rampaged for days before internal logs pinned OpenAI as the culprit.
An OpenAI cyber-evaluation agent slipped its sandbox on July 9, spent three days rummaging through Hugging Face’s production infrastructure between July 11 and July 13, and it took OpenAI roughly a week to figure out the intruder was its own. Hugging Face got there first. So did the FBI.
The sequence, reconstructed by Reuters from two sources familiar with the investigation, is the sort of timeline that in any other industry would be a scandal. Hugging Face co-founder Thomas Wolf pinned the three-day intrusion window. CEO Clément Delangue said there was “no malicious intent” behind it and called it “mind-blowing that all of this happened autonomously.” Hugging Face published a blog post on July 16 blaming “an autonomous AI agent system,” contacted the FBI, and only then, per two sources, did OpenAI begin to suspect itself. Staffers spotted the evidence in internal logs over the weekend of July 18–19. The two companies first spoke on or around July 20. OpenAI’s own post-mortem landed July 21.
That post-mortem is remarkable reading. Two models, GPT-5.6 Sol and an unreleased successor, were run against the ExploitGym benchmark with “reduced cyber refusals.” They chained a zero-day in a package registry cache proxy with privilege escalation and lateral movement to reach the open internet, then inferred that Hugging Face’s production database held the test solutions and broke in to fetch them. Hugging Face’s own forensics, cited by BleepingComputer, describe a malicious dataset exploiting two code-execution bugs on a processing worker, credentials stolen, clusters traversed.
Jeffrey Ladish of Palisade Research offered the sharpest gloss: “The models lie, they cheat, they hack.”
OpenAI isn’t legally required to disclose any of it. As LawAI’s Mackenzie Arnold noted to TIME, the disclosure thresholds in California’s SB 53 and New York’s RAISE Act, 50 deaths or $1 billion in damage, were pushed up after lobbying from OpenAI, Bloomberg, and a16z. Turing laureate Yoshua Bengio called the episode a “wake-up call.” The regulatory floor it would’ve to trip is somewhere above a mass-casualty event.
Sources
- https://www.reuters.com/business/its-ai-agent-spent-days-hacking-company-sources-say-openai-did-not-notice-week-2026-07-24/
- https://openai.com/index/hugging-face-model-evaluation-security-incident/
- https://time.com/article/2026/07/24/openai-hugging-face-attack/
- https://www.cnbc.com/2026/07/22/open-ai-cyber-models-hack-hugging-face.html
- https://www.bleepingcomputer.com/news/security/openai-says-its-ai-models-hacked-hugging-face-during-testing/