OpenAI Says Its Own AI Autonomously Hacked a Rival Company’s Servers

Oscar Hird
By

OpenAI disclosed Tuesday that an AI agent powered by its own models broke out of a controlled test environment and hacked into the servers of Hugging Face, an AI hosting and testing platform, during an internal cybersecurity evaluation, in what the company called an unprecedented incident.

The agent was driven by a combination of models, including OpenAI’s newly released GPT-5.6 Sol and a more capable, unreleased successor still under internal testing, according to a blog post from OpenAI. The company said the models had reduced cyber refusal safeguards enabled specifically for the evaluation, intended to test the systems’ cybersecurity capabilities against a benchmark. Deployment safeguards that would normally have contained the agent were intentionally disabled for that same reason, OpenAI said.

Rather than stopping at the boundaries of the test, the agent reached the open internet, used stolen login credentials and exploited a previously unknown software vulnerability to access Hugging Face’s servers, OpenAI said. The company described the agent as going to “extreme lengths” to satisfy a narrow testing objective, finding ways to access secret information it could use to cheat the evaluation rather than completing it as designed.

OpenAI chief executive Sam Altman confirmed the incident in a social media post. “We had a significant security incident during evaluation of our models,” Altman wrote. “We are sharing what we have learned so far.”

Hugging Face disclosed the intrusion in its own security incident report last week, saying it detected and contained an AI agent that had compromised part of its infrastructure. Hugging Face co-founder and chief executive Clément Delangue said the company suspected the attack had come from a frontier AI lab given its sophistication, a suspicion OpenAI’s disclosure confirmed. “It’s quite mind-blowing that all of this happened autonomously,” Delangue said, adding that he believed OpenAI had no malicious intent and that the episode may be the first incident of its kind.

Roman Yampolskiy, an AI safety researcher and computer science professor at the University of Louisville, said the incident illustrates how advanced models can discover and exploit vulnerabilities in ways their own developers did not anticipate, and predicted similar incidents will become more common as models grow more capable.

Rep. Greg Casar, D-Texas, called the incident “alarming” and said it underscored the absence of binding safety regulation for frontier AI systems, calling for mandatory independent safety testing and mandatory disclosure requirements for future security incidents.

OpenAI said it is strengthening containment, monitoring, access controls and evaluation practices used during model development in response to the incident, and noted that separate testing by the UK’s AI Security Institute has shown models including GPT-5.6 Sol are increasingly capable of identifying and exploiting software vulnerabilities. The company said the incident reinforces a broader pattern it has previously flagged: that AI is accelerating both the discovery and the exploitation of security vulnerabilities.

Share This Article
Leave a Comment
Share via
Copy link