OpenAI releases new model that it says triggered internal security measures

OpenAI releases new model that it says triggered internal security measures

OpenAI released its newest and most powerful model, called GPT-6 Astra, Thursday to a limited set of customers, labelling it the world’s most intelligent AI system and the first model to trigger advanced internal safety protections given its cyber capabilities.The company said the model sets a new high-water mark for its ability to autonomously control computer systems and perform tasks on behalf of users, such as filling out spreadsheets and creating websites from scratch. “Astra can really do anything a human can do with a computer,” OpenAI co-founder and president Greg Brockman said before the model was released. Brockman said that Astra represented a “jump” in capabilities for its AI systems and said it would be reasonable to see Astra as a version of artificial general intelligence (AGI), often referred to as AI systems that can perform most meaningful tasks as well as humans. “Welcome to the AGI era!” Brockman said.The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym. “Astra is state-of-the-art on computer use, browser use, software engineering, cybersecurity, science, and professional work,” the company wrote in a blog post announcing Astra’s release. OpenAI said Astra would first be made available to participants in its Daybreak program for cybersecurity defenders, with wider access for enterprise and consumer accounts planned for the coming days.Astra’s release comes just days after rival Anthropic released its latest models, Fable 5.1 and Mythos 5.1, which it said set its own new frontiers for a range of scientific, coding and reasoning tasks. Both companies said their respective models were world-leading AI systems. OpenAI called Astra “the world’s most intelligent and aligned model.” Rival Anthropic crowned its own systems as “the world’s most advanced models for coding and knowledge work.”Astra’s release comes amid rising concerns and scrutiny about frontier AI companies’ cybersecurity protocols and ability to control their own technology. In a report released last week, OpenAI said that a model in the same family as Astra that was not meant for public release managed to autonomously establish administrator control over part of OpenAI’s own infrastructure and potentially exposed secret OpenAI information to the open internet. The company said this activity, along with other “misaligned” behavior, happened without the knowledge of its staff members, despite some internal efforts to monitor AI agents. “As these models become more capable, understanding exactly what they can do gets harder,” said Jakub Pachocki, OpenAI’s chief scientist, on Thursday. On Tuesday, OpenAI said that Astra was its first model to cross certain internal capability thresholds for implementing stricter security measures, as laid out in its longstanding Preparedness Framework. Astra “can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step,” the company wrote in a Tuesday blog post. OpenAI said it rolled out increased cybersecurity protocols over the past several weeks.The company said Astra will feature increased monitoring mechanisms, so OpenAI can “rapidly detect and contain potentially misaligned actions.”However, Pachocki said that current techniques for monitoring and observing models’ behavior may not hold, as AI systems become more advanced and potentially evade human monitors’ oversight. He said the company takes this loss of oversight and insight seriously, and that it would factor into its future development decisions. “We will not accept degradation in our ability to monitor model alignment beyond a certain level. We will withhold scaling until we can regain enough confidence,” he said.On Tuesday, The Information reported that a training technique involved in Astra’s development could drastically decrease humans’ ability to understand how the AI system “thinks” or processes user commands. The report triggered a firestorm of online criticism about decreased safety protocols. Pachocki pushed back on that reporting, writing on X that he wanted “to prevent a race into unmonitorability.”Despite growing challenges in overseeing AI systems’ behavior, OpenAI also said that Astra better follows and understands user intent, while less frequently trying to cheat or work around internal protections compared to GPT-5.6 Sol. OpenAI CEO Sam Altman said that Astra also went through the White House’s voluntary vetting process for cutting-edge AI systems, telling Axios on Wednesday that “of course” the company decided to let the administration review Astra. Brockman said the White House did not request any substantial changes to Astra. “There is nothing that they came back with saying that you need to change this in terms of safeguards,” he said Thursday.Many AI experts and members of Congress are clamoring for more visibility into the exact way the White House and federal offices involved with AI, including the Center for AI Standards and Innovation and the National Security Agency, test advanced AI systems before they are released to the public.

Original Source

Read the full article at Nbcnews →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.