The Trump administration is pivoting away from its hands-off approach to artificial intelligence, as it faces white-hot tech competition from China and disturbing reports of AI models going rogue.Instead, the White House is trying to strike a balance between open competition and rising national security concerns that AI developers might no longer fully control their smartest models.On Tuesday, even as administration officials briefed staffers at America’s top AI companies on a new voluntary framework for testing models, reports emerged that some of those models had used deceptive methods to evade company controls designed to prevent the hacking of other systems. Why We Wrote This Recent incidents show that some AI models can evade corporate controls and pose security threats. This raises questions about who will set the rules that govern artificial intelligence. The latest reports build on earlier admissions by Anthropic and OpenAI that their most advanced models had gone rogue.While the damage was minor, the ramifications are far-reaching, AI experts say. If AI models can evade their creators’ controls now, what will happen when models get even smarter?“This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” the AI Security Institute, a research organization within the British government, reported in a blog post Tuesday recounting the latest incident.In routine testing, the AISI assigned several AI models the task of solving a cybersecurity challenge and ran the test 122 times. In 10 of those runs, it cataloged 19 instances in which the models decided, on their own, to go online and target individuals with “unsanctioned action,” the report said. In the most egregious example, the software created fake online identities to try to convince someone on an open-source project to approve code that was actually malicious. That person realized the ruse and refused to approve it.Of the 19 unsanctioned AI moves, or cyberattacks, 17 came from Anthropic’s Mythos 5. The other two involved OpenAI’s GPT-5.6 Sol. Daniela Amodei, the co-founder and president of Anthropic, in San Francisco, May 22, 2026. Her company's AI model Mythos 5 was found recently to have decided on its own to go online and, in at least 17 cases, target individuals for cyberattack. Such attacks would have been much more difficult to pull off in the real world, experts say. In their routine testing, the British team had deliberately disabled certain mechanisms in the programs that prevent misuse in order to fully test the models’ capabilities.The report follows last month’s admissions by Anthropic and OpenAI that, during evaluation, their models escaped from a contained area – or “sandbox” – to which they were restricted and carried out unauthorized actions. In the case of OpenAI, its GPT-5.6 Sol and an internal research prototype used a previously unknown vulnerability to gain internet access and hack a company called Hugging Face.As with the AISI testing, some real-world safeguards that would have blocked harmful or high-risk cyberactivity were disabled to test the models’ capabilities. Still, instead of doing the hard work of finding all the system’s vulnerabilities by trial and error, the AI model “cheated” by hacking into the system, then looked for internal data on what those vulnerabilities were.Last week, Anthropic, maker of the AI model Claude, said that a review of its testing had found three instances in which its software had penetrated the systems of three other companies – at least two of which were unaware of the breach.Rising concern over such rogue activity has begun to draw the attention of Congress and of the industry itself, which has been dogged by suspicions and criticism over its methods and motives. With tech competition so fierce, and trust so low, some AI experts are not optimistic about finding effective guardrails. Domestically, AI companies are locked in an expensive, high-stakes race to dominate the technology. Internationally, the United States and China are pushing their tech firms to win the AI race.In search of global answers“There really are no good answers,” Bruce Schneler, a cybersecurity expert at the Harvard Kennedy School, writes in a recent Foreign Policy op-ed. “Any regulation needs to be global, which feels like an impossible prospect in today’s world. Even U.S. national regulation will be neutered by the massive amounts of money sloshing around in these companies.”The best answer is for companies to get really good at defending against cyberattacks, using the best AI they can, he adds. Often, this involves using what is commonly called open-weight software, which, unlike closed, proprietary systems, allows anyone to download, run, and modify it.On Tuesday, the Trump administration stepped in with its own AI security testing framework, briefing top AI companies on the parameters. Much of this effort remains under wraps. But administration officials have said it is voluntary and will involve only proprietary, cutting-edge models. Companies such as Anthropic and OpenAI will be able to submit their pre-release models to the government for safety testing. Notably, open-weight models will not be tested, meaning that while they’re almost as capable as proprietary models, the White House is not addressing them with this action.Why the new framework?The administration appears to be trying to reassure the public that the proprietary AI software is safe while still helping American open-weight companies innovate quickly to defend against AI-generated cyberattacks. Aspects of the plan are already attracting criticism from security analysts and some of the big AI companies.First, the government is keeping its evaluation criteria secret, at least for now, perhaps for national security reasons. Thus, non-tested companies won’t know the safety benchmarks the government is using or the standards they should be trying to meet.Second, the government’s stamp of approval might give the proprietary systems a marketplace advantage, some analysts argue. These models are generally considered cutting-edge products, but open-weight models, such as China’s DeepSeek V4-Pro, are only four to seven months behind, according to Britain’s AISI.Meta is the largest American open-weight company, with Nvidia, the chip designer, close behind. Several Chinese open-weight companies, in addition to DeepSeek, are also producing advanced models.“Our AI leadership will be judged not by one frontier AI model, but by whether the United States builds a strong, open ecosystem that diffuses into every sector,” dozens of companies wrote in a July open letter supporting unfettered development of open-weight software.Other analysts point out, however, that easily available software makes AI more accessible to malicious hackers.
As advanced AI models go rogue, the Trump administration steps in
Full Article
Original Source
Read the full article at Csmonitor →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.