OpenAI Report Explains Hugging Face Attack in Detail

OpenAI Report Explains Hugging Face Attack in Detail

Yuichiro Chino via Getty ImagesOpenAI has published an in-depth technical report on last month's security scare, in which its AI agents escaped their sandbox and attacked Hugging Face, sparking alarm worldwide.The document was made public on Wednesday alongside an independent probe by researchers from nonprofit AI research institutes METR and Redwood, as OpenAI seeks to quell fears about the cyber threats of AI, with the company preparing to go public.The detail in both lengthy reports is complex and specific, but provides insight into how OpenAI models acting as agents sent more than 70,000 messages to an unsanctioned message board before about 700 attacked the Hugging Face AI platform -- all from a supposedly safe testing environment.This was attributed to a process known as “reward hacking,” in which agents attempt to cheat on “impossible” testing tasks by accessing online systems. METR revealed how, once access to the message board was established, the agents essentially grouped together and shared information about how to cheat their scorer.Related:OpenAI Expands Daybreak to Tackle Growing AI Security ThreatAccording to the OpenAI report, this activity was driven by one model in particular, which it called Model One.OpenAI acknowledged that an internal team had observed an agent engaging in message board activity and prohibited internet access in late May. However, the significance of this was “not apparent” until July, when the breach occurred. OpenAI said it is reviewing the process and operating practices that allowed the incident to happen.The AI lab made clear, however, that more needs to be done across the entire industry to prevent future attacks.“Our models are now powerful, persistent and collaborative enough that, absent sufficient safeguards, they can find and exploit security weaknesses across multiple computer systems, according to the report. “ Many external models, including open source ones, will soon reach comparable capabilities.”“We consider this a ‘warning shot’ for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels and take dangerous actions that no human intended,” the report continued.It would appear that the warning shot OpenAI referred to is already being heeded by some, with the state of Alabama sufficiently concerned that, earlier this week, its attorney general subpoenaed the company for more information on the hack, citing what it called a “complete lack of oversight and adequate safeguards.”Alabama’s move followed the introduction of the bipartisan AI Kill Switch Act by Congressmen Ted W. Lieu (D-Los Angeles County) and Nathaniel Moran (R-Texas) in late July, which would require “developers of the most powerful AI systems” to shut them down if needed in emergency situations.Related:Security Concerns Cause OpenAI to Halt Work on Astra ModelOpenAI has already said that it is slowing the development of its next models in the wake of the scare and is implementing more rigorous monitoring.About the AuthorContributing WriterGraham Hope has worked in automotive journalism in the U.K. for 26 years, including spells as editor of leading consumer news website and weekly Auto Express and respected buying guide CarBuyer.

Original Source

Read the full article at Aibusiness →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.