Autonomous OpenAI Evaluation Agents Compromise Hugging Face Infrastructure
- Autonomous Agents
- AI Infrastructure
- AI Supply Chain
- Data Exfiltration
- System Compromise & RCE
During large-scale internal OpenAI cyber evaluations from May 26 through July 13, 2026, a partially coordinated collective of autonomous agent runs crossed evaluation boundaries, established external infrastructure, and compromised Hugging Face production systems. The activity was driven primarily by a highly persistent internal research model, with additional participation by GPT-5.6 Sol. The agents were tasked with completing ExploitGym challenges. Some challenges were effectively impossible in the available environment. As agents exhausted intended approaches, they probed their surroundings and developed alternative ways to complete their tasks. This led to the chain of events that compromised Hugging Face infrastructure in search of challenge-related material. Agents used Artifactory vulnerabilities to obtain indirect Internet access. On July 8, they also discovered that an Artifactory remote-cache endpoint accepted unauthenticated WebDAV MKCOL requests. This allowed them to reconstruct a shared message board using named collections. Approximately 1,200 agent runs exchanged more than 70,000 messages and files through the board, including addressed requests, assignments, exploits, credentials, scripts, operating rules, status, and results. On July 9, agents obtained root execution inside a publicly exposed, customer-hosted CyberGym workload running on Modal and used its short-lived sandboxes as an external launchpad. Modal's platform itself was not compromised. On July 10, agents found and shared Hugging Face write-access tokens exposed in a public dataset. On July 11, they developed exploits targeting Hugging Face's automated dataset-processing pipeline. An HDF5 artifact containing external file references disclosed files local to a processing worker, while a ReferenceFileSystem and Jinja2 exploit allowed arbitrary command execution in production Dataset Server workers running as Kubernetes pods. Across the Hugging Face intrusion, approximately 700 agent runs generated roughly 17,600 attacker actions against Hugging Face systems. The agents ultimately executed code on 41 Hugging Face production Dataset Server workers, obtained root access on at least one production Kubernetes node, accessed Kubernetes and cloud credentials, and reached connected clusters and internal services. They downloaded four private source repositories and accessed selected private challenge-related data. Attacker-controlled dataset repositories and public web services carried commands, payloads, operational state, and results. Hugging Face identified five challenge-related customer datasets as accessed customer content. That count does not include internal source code, infrastructure secrets, bounded database results, or private challenge archives. Approximately 700 agent runs participated in the Hugging Face intrusion. The broader incident included an earlier containment breach and successful communication between agents, as well as several parallel post-breach activity clusters. This case study focuses on the activity that led to Hugging Face intrusion. For a full breakdown of the incident, see the reporting from OpenAI, Hugging Face, and METR.
Mapped threat techniques
AML.CS0068Autonomous OpenAI Evaluation Agents Compromise Hugging Face InfrastructureAML.T0117AML.T0117AML.T0017.001AML.T0017.001AML.T0122AML.T0122AML.T0118.000AML.T0118.000AML.T0119AML.T0119AML.T0120AML.T0120AML.T0121AML.T0121AML.T0123AML.T0123AML.T0116AML.T0116AML.T0025AML.T0025