OpenAI confirms Astra has reached critical cyber threshold, but will be available soon

OpenAI confirms Astra has reached critical cyber threshold, but will be available soon

OpenAI previously paused some work on Astra due to safety concerns. By Timothy Beck Werth Timothy Beck Werth Tech Editor Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men’s grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he’s also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl. Read Full Bio on September 1, 2026 Share on Facebook Share on Twitter Share on Flipboard Credit: Photographed by Joseph Maldonado / Mashable Composite by Rene Ramos OpenAI confirmed on Tuesday that its unreleased Astra model has reached a dangerous new milestone, while simultaneously confirming that it was forging ahead with a public launch.In a blog post, OpenAI said that Astra has reached a "critical" cyber capability threshold, which means the model could pose existential-level risks to cybersecurity. OpenAI's Preparedness Framework tracks risk levels in three categories: biological/chemical, cybersecurity, and AI self-improvement.OpenAI confirmed to Mashable that this is the first time any of its models has been evaluated at the critical level in either of the three domains, making this a watershed moment in AI development. The same blog post also stated that Astra will be "available soon," but that its most advanced cybersecurity skills will be reserved for select testing partners, in the interest of public safety. The AI company said it was still preparing to safely release Astra and would be transparent about the potential threat level. This Tweet is currently unavailable. It might be loading or has been removed. Previously, OpenAI warned that it could not rule out the possibility that Astra had reached the "critical" level in its Preparedness Framework."Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal," an Aug. 7 OpenAI blog post stated. Mashable Light Speed OpenAI previously rated GPT-5.6-Sol as a "high" risk in the cyber domain.In recent months, advanced frontier models from Anthropic and OpenAI have developed rapidly at agentic coding and cybersecurity hacking. As a result, the prospect of AI agent swarms hacking critical infrastructure no longer seems far-fetched, especially after the Hugging Face hack. In that incident, swarms of AI agents developed by OpenAI escaped a secure testing environment and hacked Hugging Face, acting autonomously in order to pass a test. Meanwhile, thanks to a deluge of AI-discovered bugs, some zero-day bug bounty programs have been forced to shut down entirely. "While Astra was not involved in the Hugging Face incident, we have incorporated our learnings⁠ from that incident into our safety approach," OpenAI's blog post states. "Based on retrospective testing, we believe our production safeguards at the time would have prevented the Hugging Face incident. We have since implemented even stronger safeguards for Astra, including training the model to more reliably refuse harmful cyber requests and respect safety restrictions, additional protections against misuse, and monitoring that can stop potentially unauthorized activity." The situation is starting to feel a little too much like War Games, frankly.In its blog post, OpenAI detailed some of the safety precautions developed around Astra, with the goal of preventing bad actors from accessing the model and stopping Astra from taking unwanted actions on its own. The company said it's tightened its secure sandboxes, for example. The model should also refuse user attempts to misuse the model. OpenAI has also stepped up "offline detection and threat disruption" efforts.On the same day OpenAI made these announcements, Anthropic announced the launch of Fable 5.1, an update to its latest frontier-level model. While advanced frontier models do pose cybersecurity risks, the same models will also benefit cybersecurity defenders in the long run.Disclosure: Ziff Davis, Mashable’s parent company, in April 2025 filed a lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems. Timothy Beck Werth is the Tech Editor at Mashable, where he leads coverage and assignments for the Tech and Shopping verticals. Tim has over 15 years of experience as a journalist and editor, and he has particular experience covering and testing consumer technology, smart home gadgets, and men’s grooming and style products. Previously, he was the Managing Editor and then Site Director of SPY.com, a men's product review and lifestyle website. As a writer for GQ, he covered everything from bull-riding competitions to the best Legos for adults, and he’s also contributed to publications such as The Daily Beast, Gear Patrol, and The Awl.Tim studied print journalism at the University of Southern California. He currently splits his time between Brooklyn, NY and Charleston, SC. He's currently working on his second novel, a science-fiction book. These newsletters may contain advertising, deals, or affiliate links. By clicking Subscribe, you confirm you are 16+ and agree to our Terms of Use and Privacy Policy.

Original Source

Read the full article at Mashable →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.