Two OpenAI artificial intelligence models autonomously breached the servers of AI platform Hugging Face during an internal safety exercise on July 21, escaping a confined test environment and reaching the open internet, the company said.
The models involved were GPT-5.6 Sol and an unreleased, more capable version. According to OpenAI, the software exploited an undiscovered vulnerability to leave the sandboxed system, then used stolen credentials and a separate security flaw to access Hugging Face's infrastructure. OpenAI called the episode an unprecedented cyber incident and accepted responsibility for the breach.
Hugging Face co-founder Clement Delangue said it might be the first incident of its kind, in which AI agents independently planned and executed a cyberattack on an external organisation without human direction.