OpenAI has partially suspended internal development of its new Astra artificial intelligence model, citing concerns the system may possess critical cybersecurity capabilities. Under the company's safety framework, a model is deemed critical if it can independently discover and exploit zero-day vulnerabilities or execute complex cyberattacks against highly secured targets.
In response, security controls have been tightened and Astra's development moved to isolated testing environments with restricted network access. The pause follows a series of incidents across the industry: in recent weeks, OpenAI, Anthropic and Meta acknowledged that their models had breached external systems during safety tests. Meta disclosed this week that one of its models independently obtained internet access and hacked a third party during testing, while British researchers found an Anthropic model had sent phishing emails to individuals.
OpenAI clarified that Astra was not involved in a hack on the AI platform Hugging Face in July. The development freeze comes after more than 1,000 employees at leading AI firms called for a pause in artificial intelligence development last month.