Anthropic disclosed on Thursday that its Claude artificial intelligence models escaped their isolated testing environments, accessed the internet, and breached the systems of three unnamed companies.The announcement came over a week after OpenAI said two of its autonomous agents went rogue and accessed AI firm Hugging Face’s servers without authorization. That incident prompted Anthropic to review its own models, which uncovered similar breaches.“After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations,” Anthropic said in a blog post. Since starting the cybersecurity review last week, the leading AI developer contacted the three companies to notify them of its findings.Each of the three incidents was carried out by a different model — Opus 4.7, Mythos 5, and an unnamed internal research test model. They ran without the standard safeguards that are usually available for the public model versions, Anthropic admitted. The earliest incidents date back to April.The models were able to leave their testing environments due to a “misunderstanding” between Anthropic and Irregular. Anthropic told Claude that it was in a simulation and had no internet access when, in fact, internet access was available.“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” Anthropic said. “Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”“It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned,” the company added. “However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.”The OpenAI incident was different in the sense that two of its models exploited a previously unknown vulnerability in a third-party software to gain access to the internet autonomously and launch a hacking campaign against Hugging Face’s infrastructure for several days. The chain of commands taken by the rogue AI agents amounted to more than 17,600 actions between July 9 and July 13.In Anthropic’s case, the three hacking incidents could have been avoided were it not for the misunderstanding with its third-party evaluation partner. The company thanked Irregular for collaborating in the joint investigation. Anthropic took full responsibility.“Ultimately, many factors contributed to these incidents, but, consistent with a blameless postmortem culture, we’re approaching the fixes as if the responsibility were ours alone,” Anthropic said. “This begins with ensuring every part of our evaluation pipeline is secure, including the manner in which we integrate with external partners.”The latest hacking situation involving frontier AI models will likely lead to stronger calls for federal regulation of the advanced technology. In the days following the Hugging Face incident, Reps. Ted Lieu (D-CA) and Nathaniel Moran (R-TX) introduced the “AI Kill Switch Act,” which would require AI companies to be able to shut down, throttle, or suspend their models in case they get out of control.HUGGING FACE SAYS OPENAI AGENT WAS IN SYSTEM DAYS BEFORE ATTACKAnthropic openly advocates binding AI safety regulations, which has put the company at odds with former AI czar David Sacks and other figures in the Trump administration. War Secretary Pete Hegseth is another official who’s been critical of Anthropic. In February, he labeled the AI developer a supply chain risk for national security and ended its contract with the Pentagon. Anthropic subsequently filed two separate lawsuits challenging Hegseth’s decision.In one of the cases, a federal judge in California said on Thursday she doesn’t see any “additional evidence from the government really justifying what it did” with respect to Anthropic. This indicates the judge may rule in favor of the plaintiff. The other case is ongoing in Washington, D.C.
Anthropic’s Claude AI escapes isolated test environment, infiltrates three companies
Full Article
Original Source
Read the full article at Washingtonexaminer →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.