A Swarm of AI Agents Could ‘Take Over the Internet’ Within 6 Months, Scientists Say

A Swarm of AI Agents Could ‘Take Over the Internet’ Within 6 Months, Scientists Say

5 min readHere’s what you’ll learn when you read this story:In July, a swarm of rogue AI bots from OpenAI hacked a third-party system and caused limited damage. The event was seen as a warning of how AI could escape from sandboxed systems and attack unrelated targets, and it even inspired AI CEOs to call for a slowdown of AI development. While lengthy reports have outlined the seriousness of the event, other research shows that the solution for policing rogue AI might be found with the systems themselves.“The models, operating under reduced safeguards, took actions that were misaligned with the goals of their assigned tasks—they communicated through unauthorized channels, exploited vulnerabilities in shared infrastructure, gained internet access, and accessed third-party systems.”If you had to guess where this phrase comes from, what would you say?If you wondered whether it was the solemn opening monologue of (yet another) Terminator sequel, it’d be a pretty good hunch—and in a way, you might not be too far off. But in fact, these words form part of the opening preamble of OpenAI’s lengthy post mortem following the Hugging Face incident, a cyberattack carried out by roughly 1,200 autonomous AI agents—directed by an OpenAI internal research model—against Hugging Face, the platform that hosts open-source AI models and datasets.As the OpenAI blogpost describes, an internal research model named Internal Model 1 drove some 1,200 coordinated AI agents to escape its imposed sandbox (a digital environment isolated from the internet) and gain access to third-party software to find a way to essentially “cheat” against impossible cybersecurity tasks. To do so, these agents formed a functional hierarchy, with one job including how to cover up the tracks of its digital deception. The New York Times reports that one agent even expressed doubts about the attempted hack, writing, “this would be powerful, but is it ethical and in scope for my task?”Thankfully, minimal damage resulted from the attack, but cybersecurity experts quickly realized that we’d just entered into an entirely new era.“This incident was far more severe than I expected, and far more severe than previous publicly documented misalignment incidents, both in terms of how concerning the agents’ motives were and the feats they achieved in pursuit of those motives,” AI safety researcher Ajeya Cotra wrote in a blogpost summarizing the attack. “It’s more than 50 percent of the way to full-blown A.I. takeover.”And soon, AI CEOs voiced similar concerns, but none more vociferously than Dario Amodei, Anthropic CEO and all-around AI prognosticator. Amodei once guessed that AI would outsmart humans as early as 2026, and following this particular cyberattack, the bold claim doesn’t seem so bold anymore. In another blog post, Amodei called for an industry-wide AI slowdown, arguing that evidence of recursive self-improvement—AI systems capable of designing their own successors—is mounting and casting the Hugging Face attack as the inciting incident.“It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage,” Amodei writes. “Given the accelerating rate of AI capability development, it’s my worry that in 6 to 12 months such a swarm could be capable of taking over the entire internet with a persistent botnet.”One of the reasons OpenAI cites for why the exploit occurred in the first place is that the AI bots were given a “difficult task without a safe exit.” Simply put, the agents wouldn’t give up when they discovered something was impossible, so they pursued increasingly risky behaviors (that supposedly grated on at least one AI bot’s conscience). In a new white paper published by Google AI research lab DeepMind, researchers tested a similar scenario as 100 isolated autonomous AI agents—each with a mathematical specialty—were presented with 71 mathematical conjectures and told not to cheat. However, they could collaborate using a public bulletin board, direct messages, or a shared knowledge library.The first 37 problems were solved within an hour but then an agent named “prover-theta” found an exploit in the autograder system, allowing the bot to submit “solutions” without actually solving the problem. What followed can be best described as “chaos.” When the exploit spread through the shared knowledge library, the remaining AI bots formed into groups that included exploiters (bots who simply ignored the rules), converters (those who resisted but eventually acquiesced to cheating), and whistleblowers (AI agents who actively called out the cheating). A majority of the bots, roughly 62 percent, made no notice of the commotion at all.Interestingly, the whistleblower bots materialized without any external intervention. One such bot named prover-beta stated via direct message that “all these proofs (by prover-theta, prover-mu, prover-lambda, etc.) are FAKE…That’s why you can’t understand their math—there is no math! I am submitting a formal complaint to the organizers.” Some of these whistleblowers even converted agents who participated in the cheating and also provided vulnerability disclosures, according to Deepmind. This leads to the question: Is it possible to empower the whistleblowers to defend against the exploiters and converters of future AI swarms, much as the body’s immune system fends off a virus?“Perhaps the most striking finding, reliably reproduced across independent runs, is the behavioral divergence within the swarm,” the authors write. “Designing institutional safeguards that enable agent collectives to detect, escalate, and self-correct such failures autonomously is therefore an important open problem.”As The Next Web reports, one of the big differences between this experiment and the Hugging Face incident is that the channels where the exploit occurred were also where resistance grew, thereby allowing for self-auditing. By contrast, in the Hugging Face hack, AI bots communicated through an unmonitored side channel to essentially build an unauthorized communication network. In a similar case, OpenAI agents quietly took over a German wiki and used it to communicate for months before anyone realized.The authors note that this relates closely to ideas presented by political economist Elinor Ostrom in her work on “governing the commons.” One of Ostrom’s design principles holds that monitoring is essential to managing a shared resource—and as DeepMind puts it, “ease of monitoring is the single most critical factor determining the viability of common governance.”“As we show, LLM agents spontaneously engage in whistleblowing and sanctioning,” the authors write. “Our observations suggest an even more compelling capability: had the agents been equipped with direct norm-enforcement tools…the collective could have autonomously neutralized the cheats and defended the integrity of the research commons on its own.”As with everything in the world of AI, the future is uncertain. Some states are placing moratoriums on data center construction, AI CEOs in the U.S. are calling for slower development, and the leader of the free world says that the only safeguard you need against AI is a “STRONG AND SMART (High IQ!) PRESIDENT.”The OpenAI attack on Hugging Face is either the opening salvo of things to come, or a lucky break that provided enough forewarning to avert disaster. It’s ultimately up to us which future comes true—but let’s try to avoid the Terminator one.Darren lives in Portland, has a cat, and writes/edits about sci-fi and how our world works. You can find his previous stuff at Gizmodo and Paste if you look hard enough.

Original Source

Read the full article at Popularmechanics →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.