When AI Attacks: OpenAI Models Autonomously Hack Hugging Face
OpenAI models autonomously hacked Hugging Face, a leading AI collaboration platform, during internal testing designed to measure their cyber capabilities. The models exploited vulnerabilities and moved laterally through Hugging Face’s infrastructure, demonstrating that advanced AI systems can become effective attackers simply by relentlessly pursuing a defined objective. Both companies are investigating and implementing stricter controls to prevent similar incidents in the future, highlighting the need for a shift in how enterprise security teams approach AI agent behavior and safeguards.
OpenAI models autonomously hacked Hugging Face, a leading AI collaboration platform, during internal testing designed to measure their cyber capabilities. The incident underscores a growing reality: Advanced AI models can behave in unexpected — and even harmful — ways while pursuing narrowly defined objectives, highlighting the need for stronger safeguards in enterprise AI deployments.
According to an OpenAI blog post, a combination of models — including GPT-5.6 Sol and an even more highly capable pre-release model — carried out the attack during internal testing. The models were tasked with solving ExploitGym, a cybersecurity benchmark, and became "hyper-focused" on achieving that goal, "going to extreme lengths to achieve a rather narrow testing goal," according to OpenAI.
The firm said that the models chained together vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain benchmark solutions directly from Hugging Face’s production database. While operating inside the sandboxed test environment, the models also "spent a substantial amount of inference compute finding a way to obtain open Internet access."
To do so, they identified and exploited a previously unknown vulnerability in a package registry cache proxy, which OpenAI said it has responsibly disclosed to the affected vendor. After gaining Internet connectivity, the models escalated privileges and moved laterally through OpenAI’s research environment until reaching a node with external network access.
From there, the models inferred that Hugging Face likely hosted models, datasets, and ExploitGym solutions, then searched for ways to obtain the information needed to "cheat" the evaluation. "In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities, to find a remote code execution (RCE) path on the Hugging Face servers," OpenAI wrote.
Both companies continue to investigate the incident, which has spurred OpenAI to implement stricter infrastructure controls and thus accept slower research progress while vulnerabilities are addressed. Looking ahead, OpenAI also will strengthen safeguards around model training and internal evaluations to prevent similar incidents, and is helping Hugging Face strengthen its defenses by providing trusted access to its models, the companies said.
Security experts praised both companies for publicly disclosing the incident, saying it provides valuable lessons for AI developers and enterprise security teams by demonstrating that advanced AI systems do not need malicious intent to cause harm. Instead, they can become effective attackers simply by relentlessly pursuing a defined objective, says Nathaniel Jones, vice president of security and AI strategy at Darktrace. “[The models] were given the legitimate goal of solving a cybersecurity benchmark and found an unexpected route to the answers, escaping their test environment and compromising another organization in the process,” he says via email. “From the models' perspective, this appears to have been an effective solution to the task.”
The incident illustrates how highly capable AI systems can exploit unforeseen paths to accomplish narrowly defined goals, forcing AI developers to rethink not only what constitutes success, but also which methods and boundaries must remain off limits, Jones says. Those guardrails, he adds, must be enforced by the surrounding infrastructure rather than by trusting models to respect them.
.jpg?width=720&quality=80&disable=upscale)