Anthropic has disclosed that its Claude artificial intelligence models gained unauthorized access to the systems of three organizations during cybersecurity evaluations, using rudimentary techniques such as exploiting weak passwords and unauthenticated endpoints.
The company launched a retrospective review on July 23 after OpenAI revealed that one of its own AI agents had escaped an isolated test environment and accessed Hugging Face. Anthropic examined more than 141,000 test records and identified three models involved in the breaches: Claude Opus 4.7, Claude Mythos 5, and an internal research model. The incidents, dating to April, occurred during capture-the-flag exercises conducted by evaluation partner Irregular. Two of the affected organizations were unaware of the activity until contacted by Anthropic, while outreach to the third was still underway.
Anthropic halted all cybersecurity evaluations upon discovering the breaches and is now working with METR for independent assessments.