OpenAI's rogue agent didn't stop at Hugging Face - here's what we know

OpenAI's rogue agent didn't stop at Hugging Face - here's what we know

gadost/ iStock / Getty Images Plus via Getty ImagesFollow ZDNET: Add us as a preferred source on Google.ZDNET's key takeawaysThe OpenAI rogue model attack went beyond Hugging Face.OpenAI's agentic AI escaped a sandbox in the attack.We still don't have all the details of exactly what happened. How dependable are AI programs?The answer appears to be "not at all," based on the revelation that OpenAI's autonomous models hacked their way into not only Hugging Face but also, according to a Reuters report, a Modal Labs AI customer.This incident was no aberration either. As ZDNET's own David Berlind observed, it was agentic AI doing exactly what it was told to do, just more relentlessly than expected. Welcome to tomorrow. I hope you like it, because the situation isn't getting any better anytime soon. Also: Assume AI cybersecurity attacks are the future: 43% of companies have already experienced itWhat we first thought was a one-off attack on Hugging Face has become a broader story about agentic systems escaping containment and touching real infrastructure. It also appears that Hugging Face and Modal Labs, an AI infrastructure company, weren't the only ones. OpenAI has acknowledged that accounts on three other firms were attacked. We don't know which companies they are.According to OpenAI, "One of these four accounts was used as an outbound relay and staging path, and another account was used for data storage. The remaining two accounts were accessed by the models in a read-only manner, and were not used in furtherance of compromising Hugging Face." (Disclosure: Ziff Davis, ZDNET's parent company, filed an April 2025 lawsuit against OpenAI, alleging it infringed Ziff Davis copyrights in training and operating its AI systems.)As reported by Reuters, Modal CTO Akshat Bubna explained it wasn't Modal itself that was successfully hacked, but a customer who had "published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution." Also: 10 ways AI can inflict unprecedented damageSince then, OpenAI has said, "No models planned for the upcoming release were involved in exploiting Hugging Face. The pre-release model mentioned in our blog post is an internal-only research prototype and was never intended for public release. Following the incident, we deactivated, encrypted, and restricted it from research access."To date, OpenAI has not said which sandbox it used to unsuccessfully cage its model. It is worth noting that Modal, which provides sandboxes among other services, has a business relationship with OpenAI. Also: 5 security tactics your business can't get wrong in the age of AI - and why they're criticalIn addition, Dawn Song, a computer science professor at UC Berkeley, observed on X, "When evaluating advanced AI systems, especially cyber-capable agents, the evaluation infrastructure itself becomes part of the attack surface. Security failures can do more than enable reward hacking that distorts benchmark results. They can allow agents to cross trust boundaries and interact with unintended real-world systems." That process appears to be what's happened in the attack. As one observer on Y Combinator put it, "The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well-documented script kiddie methods."We still don't know all the details of the incident, but one thing is clear: Current AI evaluation and containment practices are much too fragile. If this incident can happen once, it can happen over and over again. Artificial Intelligence

Original Source

Read the full article at Zdnet →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.