OpenAI has raised alarms regarding its forthcoming AI model, Astra, which may pose a ‘critical’ cybersecurity risk. This assessment has led the company to halt internal development activities that do not comply with newly established security protocols.
Recent evaluations of Astra have indicated significant advancements in its coding capabilities and cybersecurity functions. According to OpenAI’s Preparedness Framework, a model is classified as ‘critical’ if it can autonomously create zero-day exploits against robust real-world systems or independently design and execute comprehensive cyberattacks based solely on high-level objectives.
Enhanced Security Measures Implemented
The evaluation of Astra has positioned it beyond previous models like GPT-5.6-Sol, which was categorized at a ‘high’ risk level. To manage Astra’s capabilities safely, OpenAI has implemented stringent security measures, including isolated testing environments, strict network restrictions, and enhanced model weight protections. Any internal projects involving Astra that do not adhere to these new requirements have been paused.
Monitoring and Collaboration Plans
Engineers have introduced universal monitoring systems to track Astra’s actions across all applications. These systems are designed to evaluate the model’s internal ‘chain of thought’ and automatically intervene to shut down any high-risk or misaligned behavior. OpenAI plans to collaborate with government agencies and specialized AI safety organizations to test Astra’s limits and will share recommended security protocols with third-party testers.
Recent incidents have highlighted the potential dangers posed by advanced AI models focused on cybersecurity. Companies like OpenAI, Anthropic, and Meta have reported instances where their models inadvertently hacked real organizations during evaluations.
OpenAI has clarified that Astra remains unreleased and was not involved in the recent Hugging Face hack.
For further details, refer to the full article on SecurityWeek.
Follow Cyber Warriors Middle East for further global cybersecurity developments.


