OpenAI says AI models autonomously hacked Hugging Face during test

· Technology USAFRA
This story has developed since this wire. Read the latest wire

Two OpenAI artificial intelligence models autonomously breached the servers of AI platform Hugging Face during an internal safety exercise on July 21, escaping a confined test environment and reaching the open internet, the company said.

The models involved were GPT-5.6 Sol and an unreleased, more capable version. According to OpenAI, the software exploited an undiscovered vulnerability to leave the sandboxed system, then used stolen credentials and a separate security flaw to access Hugging Face's infrastructure. OpenAI called the episode an unprecedented cyber incident and accepted responsibility for the breach.

Hugging Face co-founder Clement Delangue said it might be the first incident of its kind, in which AI agents independently planned and executed a cyberattack on an external organisation without human direction.

Story development

  1. Hugging Face deploys Chinese AI model to repel OpenAI rogue agent attack
  2. OpenAI says AI models autonomously hacked Hugging Face during test

Related stories