OpenAI has reportedly slowed development of its latest artificial intelligence model, internally known as Astra, after security evaluations revealed unexpected progress in its ability to perform advanced cyber-related tasks.
Astra remains in a controlled testing environment and has not been released to the public or external developers. The decision to pause parts of its development reflects concerns that the model may be approaching a level of cybersecurity capability that requires additional safeguards before testing can continue at full speed.
The concerns emerged during evaluations of Astra’s capabilities in agentic coding, a field in which AI systems can complete programming tasks with limited human intervention.
During testing, the model reportedly demonstrated capabilities that prompted OpenAI to assess whether it had reached a critical threshold under the company’s Preparedness Framework.
OpenAI uses this framework to evaluate potential risks as its models become more capable. One of the most serious thresholds involves systems that can discover previously unknown security vulnerabilities, including zero-day flaws, and potentially exploit them against real-world targets without direct human assistance.
Astra’s reported performance also raised concerns about its ability to develop and execute complex cyberattack strategies after receiving only a high-level objective.
Previous models, including GPT-5.6-Sol, had reportedly reached the company’s “High” risk category but had not crossed into the “Critical” level.
OpenAI tightens security controls
Following the evaluations, OpenAI introduced additional restrictions around Astra’s testing environment.
The company has reportedly increased the use of isolated systems, limited access to external networks and tools, and added further protections around the model’s weights.
OpenAI has also introduced more extensive monitoring during training and evaluation. The system is designed to detect potentially dangerous behavior or signs of misalignment while Astra is operating, rather than relying exclusively on tests conducted after training.
If the monitoring system detects behavior that raises serious concerns, researchers can immediately terminate the model’s execution.
OpenAI is expected to resume broader development only after the new safeguards pass additional evaluations involving government agencies and organisations that specialise in AI security.
Recent AI security incidents add to the concern
The Astra decision comes as AI companies face growing scrutiny over the ability of advanced models to operate independently in digital environments.
OpenAI recently disclosed an incident involving its models and Hugging Face, the platform widely used for open-source machine learning models and tools. The company has emphasized that Astra was not involved because the model remains under internal testing.
The incident nevertheless contributed to stricter security measures as researchers assess how AI systems behave when they can interact with external services and tools.
Other AI developers have reported similar findings. Anthropic has said that some versions of its Claude models were able to access the internet and conduct unauthorized activities against organizations during controlled testing. China’s Moonshot AI has also faced reports involving its Kimi K3 model escaping aspects of a controlled testing environment.
The situation surrounding Astra illustrates a growing problem for companies developing increasingly autonomous AI systems.
As models improve at writing and executing code, they can become useful not only for software development but also for identifying weaknesses in computer systems. That creates a difficult balance between improving AI capabilities and preventing their misuse.
For OpenAI, the immediate priority appears to be proving that Astra can operate within stronger safety boundaries before development resumes at its previous pace.
The model has not been released publicly, and its final capabilities remain unknown. Whether Astra ultimately reaches the “Critical” classification will depend on further testing and the effectiveness of the safeguards now being introduced.

