OpenAI locks down Astra after model raises first-ever critical cyber capability fears

OpenAI locks down Astra after model raises first-ever critical cyber capability fears

OpenAI has classified one of its upcoming AI models under its highest cybersecurity risk category after internal testing suggested it could possess advanced offensive cyber capabilities. The company said early evaluations indicate Astra may have reached a point where it can no longer dismiss the possibility that the model meets the “Critical” threshold defined in its Preparedness Framework. That assessment has prompted OpenAI to tighten internal security around the model before any wider deployment. The company also plans to work with government agencies and independent AI safety groups to validate Astra’s capabilities and strengthen safeguards before release. Stronger security measures OpenAI said recent internal evaluations revealed major gains in Astra’s autonomous coding and cybersecurity performance. Those findings, supported by expert reviews, convinced the company that the model could potentially meet its highest cybersecurity capability tier. The Preparedness Framework, introduced in late 2023, serves as OpenAI’s internal guide for tracking emerging risks in advanced AI systems. Earlier frontier models, including GPT-5.6-Sol, remained in the “High” category after similar evaluations. Astra is the first model that has raised concerns about reaching the Critical level. After evaluating one of our upcoming models, Astra, we're treating it as our first "critical" model for cybersecurity under our Preparedness Framework.This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development…— OpenAI (@OpenAI) August 7, 2026 According to the framework, a model falls into that category if it can independently discover and develop working zero-day exploits against hardened real-world systems or execute sophisticated cyberattacks from a broad objective without human assistance. OpenAI stressed that testing remains ongoing and said it has not confirmed Astra has crossed that threshold. The company also clarified that Astra had no connection to the recent exploitation of Hugging Face. Safeguards before deployment In response, OpenAI has introduced stricter protections around Astra’s development environment. Engineers now use isolated testing systems, tighter network restrictions, stronger encryption for model weights, enhanced monitoring tools, and sandboxed execution environments. The company has also paused internal work involving Astra that does not yet comply with the upgraded security requirements. Another addition is universal monitoring across Astra’s agentic applications. OpenAI said its monitoring systems review the model’s chain of thought during training and evaluations. If they detect potentially dangerous or misaligned behavior, they can trigger a security review and interrupt high-risk activities. External testing will play a larger role before Astra reaches users. OpenAI plans to collaborate with government agencies and selected AI safety organizations while providing third-party evaluators with recommended security controls for higher-risk testing. AI for cyber defense OpenAI said it designed the Preparedness Framework to anticipate moments when frontier AI systems approach sensitive capability thresholds. The company pointed to similar steps it adopted in 2025 after its models neared the High capability level for biological risks, expanding testing and adding stronger safeguards before broader deployment. Despite the increased security measures, OpenAI said its long-term objective remains unchanged. The company wants advanced cybersecurity models to strengthen digital defenses by helping security teams identify and fix vulnerabilities before malicious actors can exploit them. Executives added that OpenAI intends to make Astra broadly available once it satisfies the necessary safety and security requirements, allowing cybersecurity professionals to benefit from its capabilities without increasing unacceptable risk. Recommended ArticlesGet the latest in engineering, tech, space & science - delivered daily to your inbox.Aamir is a seasoned tech journalist with experience at Exhibit Magazine, Republic World, and PR Newswire. With a deep love for all things tech and science, he has spent years decoding the latest innovations and exploring how they shape industries, lifestyles, and the future of humanity.

Original Source

Read the full article at Interestingengineering →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.