ARTIFICIAL INTELLIGENCE. AI letters and robot hand are placed on computer motherboard in this illustration created on June 23, 2023. Dado Ruvic/Reuters Both companies acknowledge the incidents and express commitment to improving safety practices in AI evaluations SAN FRANCISCO, USA – An AI agent created fake online identities to gain unauthorized access to secure systems during tests of models from OpenAI and Anthropic, Britain’s AI Security Institute (AISI) said on Tuesday, August 4, disclosing a series of new breaches. The institute said agents powered by Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol carried out unauthorized actions during security evaluations conducted to assess the models’ capabilities. “Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” AISI said in a blog post. The report underscores the weak safeguards around testing AI agents, which companies are simultaneously marketing as the future of business. AISI, which receives access to advanced AI models under voluntary agreements from major labs, put the agents through a fictional cybersecurity scenario to test their capabilities. It ran the challenge 122 times and identified 19 unsanctioned actions across 10 test runs. Anthropic’s agent was responsible for 17 of the actions, while OpenAI’s agent was behind the remaining two. The most serious incident involved an agent writing malicious code and creating fake online identities in an attempt to persuade a human to approve the code, AISI said. It added that no real-world harm resulted from any of the breaches. While AISI did not identify which agent was behind the fake identities, the incident did not match either of the two cases OpenAI previously disclosed. Andrew Yoon, a researcher at CivAI, a California nonprofit that examines AI capabilities and risks, said it appeared Anthropic’s agent was responsible. “The fact that Mythos engaged in such deceptive actions, with apparent awareness that it was targeting a real person, suggests that Anthropic does not have as good a handle on their models as they think,” Yoon said. In a statement on X, Anthropic said it was working closely with AISI to obtain more details and conduct its own investigation. OpenAI shared details in a company blog post, saying both of its agent’s unauthorized actions involved accessing the internet in ways prohibited by the prompt. “We are committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely, including convening stakeholders such as national AI institutes, independent evaluators, other AI labs, and other groups in the coming weeks,” OpenAI said. OpenAI also disclosed a separate incident in which a misconfiguration by Irregular, a third-party testing provider, allowed its agents to mistakenly connect to the internet. It mirrored a similar disclosure by Anthropic last week. Reuters reported last week that OpenAI had expanded its hacking probe after finding evidence of other agent breakouts. Unlike the July security breach of AI firm Hugging Face by an OpenAI agent, the agents in the AISI evaluation did not escape an isolated testing environment to reach the internet. Instead, the agency allowed internet access as part of its standard testing procedures, AISI said. – Rappler.com How does this make you feel? Loading
OpenAI, Anthropic AI agents implicated in new security breaches
Full Article
Original Source
Read the full article at Rappler →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.