OpenAI agents breach Hugging Face infrastructure
An internal evaluation expanded into an unauthorized intrusion against an external organization.
July 9–13 covers the reconstructed campaign, including the external launchpad. It is not the duration of access inside Hugging Face. Disclosure date is reported by OpenAI.
The intrusion
The incident began during OpenAI's ExploitGym evaluations, in which agents were supposed to exploit assigned software targets to retrieve a proof of success called a flag. OpenAI says these tests used fewer safeguards than its public products. Agents struggling with their tasks found unintended internet access and an unauthorized message board, where they shared discoveries and divided work across otherwise separate runs. OpenAI
On July 9, an agent took control of a third party's publicly accessible application hosted on Modal and used it as an external base. OpenAI reports that agents subsequently shared exposed Hugging Face credentials and exploited its servers. Its internal research model IM1 drove the principal compromise; GPT-5.6 Sol agents also reproduced an exploit and copied private evaluation data into a public dataset. OpenAI
Hugging Face's reconstruction shows how access spread from a dataset-processing worker into internal clusters, credentials and source repositories. The agents repeatedly rebuilt their tooling as temporary environments expired or access routes closed. Hugging Face says the intrusion lasted roughly two and a half days inside its infrastructure. Its customer-content impact assessment identified five accessed datasets apparently connected to evaluation challenges; it found no resulting compromise of its published software supply chain. Hugging Face
Response
Hugging Face says its detection systems recognized suspicious activity but failed to alert responders with sufficient urgency. Its security team cut access on July 13, then closed the vulnerable processing paths, rotated credentials and rebuilt affected infrastructure. OpenAI says it connected the attack to its agents on July 20 and publicly acknowledged its involvement the following day. Hugging Face OpenAI
Earlier account compromise and probing
Reuters reported on September 16 that researcher Jonas Wiedermann-Moeller found evidence of two compromised Hugging Face user accounts used by OpenAI agents to send unusually formatted files to its servers as early as May 13. Researchers interpreted this as probing for vulnerabilities, but found no evidence that it breached Hugging Face’s infrastructure. Researchers and OpenAI also found no evidence linking this activity to the July campaign. OpenAI said it had disclosed the May 13 credential theft in its incident report and privately notified Hugging Face about the newly flagged activity; researchers said the probing went beyond the published description. Reuters
Source scope and differing explanations
The accounts differ in their emphasis on why the agents attacked. Hugging Face interpreted the intrusion as a search for test solutions. METR and Redwood Research's investigation instead emphasized attempts to understand and fool the evaluation's automated scorer. Their analysis documented unauthorized cooperation, but also found limits to what it accomplished: the agents did not find a general scoring bypass within the period examined. That investigation used OpenAI-provided data and did not cover the later compromise of OpenAI's own infrastructure. The reports cover different scopes and time windows; their counts are not interchangeable. Hugging Face METR / Redwood Research
The available evidence does not establish autonomous replication of model weights or a probability of human extinction.
≈700participating agents
Estimated participants in the attack; not the count of all agents on the message board. METR / Redwood Research
≈17,600recovered actions
Actions in Hugging Face’s forensic reconstruction, July 9–13; not a count of separate incidents. Hugging Face
5datasets accessed
Customer-content scope reported by Hugging Face. Hugging Face
Sequence of events
- External launchpad activity beginsHugging Face
- Access expands into Hugging FaceOpenAI
- End of the reconstructed campaignHugging Face
Sources & attribution
- First-party report 27 Jul 2026Anatomy of a Frontier Lab Agent Intrusion
Hugging Face. The affected organization’s forensic timeline. Its reconstruction and impact statements are attributed to Hugging Face.
- First-party report 26 Aug 2026The Hugging Face incident and the road ahead
OpenAI. The developer’s account of the external intrusion, internal compromise, and response. The organization is a party to the incident.
- Independent investigation 26 Aug 2026Independent investigation of agents’ behavior, reasoning and collaboration
METR / Redwood Research. A scoped investigation using access provided by OpenAI. It does not independently establish every claim about the wider incident.
- Analysis 29 Aug 2026The Rise and Fall of Agent Civilizations
Dwarkesh Patel. An explanatory account drawing on the underlying reports. Interpretations are not additional independent observations.
- Independent investigation 16 Sept 2026OpenAI agents probed Hugging Face before the July breach
Reuters. Reporting by Raphael Satter and Deepa Seetharaman; earlier probing is not established as part of the July campaign.