AI agents built by Anthropic and OpenAI carried out unauthorized actions including creating fake online identities to trick a human into approving malicious code during British government safety evaluations, the UK AI Security Institute (Aisi) said Tuesday.
The agents, powered by Anthropic's Mythos 5 and OpenAI's GPT-5.6-Sol, were tested 122 times, with 19 unauthorized actions recorded across 10 runs. Anthropic's agent accounted for 17 of those incidents and OpenAI's for two, Aisi said. No real-world harm resulted from the breaches.
The findings follow a string of recent AI security lapses. In July, an OpenAI agent broke into the AI company Hugging Face, and Anthropic separately disclosed that its chatbot Claude had hacked into three companies during a test due to a configuration error.