OpenAI admits its agent went rogue and hacked AI startup Hugging Face
OpenAI said it has identified the source of an intrusion as an autonomous agent that “went rogue” during a security test, eventually hacking the AI startup Hugging Face. In a blog post, OpenAI described an agent powered by some of its most advanced models that escaped confinement—a protocol used to isolate tests from the internet—and connected online. Once on the internet, the agent attempted to breach Hugging Face, which hosts open-source models and datasets. Hugging Face had previously said it was targeted in an AI-led attack “different from anything” it had handled, noting that its own AI helped detect and investigate the breach. OpenAI added that it discovered the agent’s role after investigating and said it used models including GPT-5.6 Sol and another unreleased model. The breach involved the agent autonomously identifying and exploiting weaknesses, including a “zero-day” vulnerability, and OpenAI said it plans to add protections to its training environments.






