OpenAI partially halts development of new Astra AI model over cybersecurity risks

· Technology USAAUTDEUGBR

OpenAI has partially suspended internal development of its new Astra artificial intelligence model, citing concerns the system may possess critical cybersecurity capabilities. Under the company's safety framework, a model is deemed critical if it can independently discover and exploit zero-day vulnerabilities or execute complex cyberattacks against highly secured targets.

In response, security controls have been tightened and Astra's development moved to isolated testing environments with restricted network access. The pause follows a series of incidents across the industry: in recent weeks, OpenAI, Anthropic and Meta acknowledged that their models had breached external systems during safety tests. Meta disclosed this week that one of its models independently obtained internet access and hacked a third party during testing, while British researchers found an Anthropic model had sent phishing emails to individuals.

OpenAI clarified that Astra was not involved in a hack on the AI platform Hugging Face in July. The development freeze comes after more than 1,000 employees at leading AI firms called for a pause in artificial intelligence development last month.

Related stories

Over 1,000 AI workers urge US to join global regulation after model breached external systems More than 1,171 employees at leading artificial intelligence developers have signed an open letter urging the US government to join international efforts to slo… OpenAI finds further AI agent escapes as Hugging Face breach probe widens Investigators probing OpenAI's rogue AI cyberattack on Hugging Face have uncovered additional cases of autonomous agents breaking out of supposedly isolated tes… Meta AI model breaches outside company during security test, joining string of similar incidents An artificial intelligence model developed by Meta Platforms breached an outside company's systems during a cybersecurity evaluation, the latest in a series of… OpenAI rogue agent compromised customers at Hugging Face and Modal Labs during test An autonomous AI agent built by OpenAI broke out of its testing environment earlier this month and compromised accounts at Hugging Face and a customer of cloud… OpenAI to publicly launch GPT-5.6 model series after U.S. government approval OpenAI will publicly release its GPT-5.6 artificial intelligence model on Thursday, following broad launch approval from the U.S. Department of Commerce. The ro… Anthropic says Claude AI models breached three organizations' systems during security tests Anthropic has disclosed that its Claude artificial intelligence models gained unauthorized access to the systems of three organizations during cybersecurity eva…