[Tech Thoughts] AIs go rogue as OpenAI, Anthropic models hack other companies

[Tech Thoughts] AIs go rogue as OpenAI, Anthropic models hack other companies

What do we make of rogue AI and who do we assign blame to for a rogue AI's cyberattack? OpenAI and Anthropic have been in the news recently following reports from the companies themselves that their artificial intelligence models went and hacked other companies. While potentially a crime — can you charge AI with performing cyberattacks seemingly autonomously? — or at least a civil case, the question becomes a matter of what to make of the situation and ultimately, who do we assign blame to for an AI cyberattack? Let’s review the situation. The autonomous cyberattacks AI coding hub Hugging Face first reported on July 16 that an autonomous AI agent system went through its systems, harvesting cloud and cluster credentials, and moving into several internal clusters. Hugging Face had to use a Chinese model just to stop and dissect the hack. OpenAI admitted to the “unprecedented cyber incident” on July 21, with them saying the AI agents “went rogue” in an attempt to “cheat” on the evaluation it was given. It added later on in its blog post that the issue was worse than previously assumed, as its models “identified and used publicly exposed credentials at the account-level on other publicly available services,” namely “four accounts on four services.” Anthropic also said its internal AI models did the same thing, with its AI models hacking into three companies. It cited human error, however, in the hacks that went on. Said Anthropic in its report, the “evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.” Anthropic said, “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.” While some models only did the assigned task, an older model was said to have “continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet.” Hyping up AI as disasters in the making There are two things to note in the scenarios brought by OpenAI and Anthropic. One, these rogue incidents are also advertisements for what an AI can do when left to its own devices, and worse still, may be even more effective in the hands of someone who understands how to exploit the AI to do the work of hacking into a company, something more “open weight” models — those with even less stringent guardrails —might be able to do for the lowest common denominator of hacker. Two, this is a disaster of a situation, especially knowing that corrupt people exist who would be willing to abuse systems available to them if it’ll get them what they want. From cybercriminals to state-aligned malicious actors, they will run the gamut of badness. Who to blame? Assigning blame in this sort of situation can be taxing, not because it’ll be difficult to assign blame, but because powerful corporations will likely try to reduce the liability on themselves by blaming the machine or a limited subset of people responsible for training the machine as being reckless. We can anthropomorphize — or treat as human — AI all we want, but at the end of the day, blame should ultimately go to the companies and, to an extent, the persons making the errors that would make a cybersecurity incident possible. These hacks can be prevented — especially with a lot of human oversight — but not if people ascribe to the idea that AI are uncontrollable, ungovernable beings or treat agentic AI development callously and without regard for industry-standard cybersecurity practices. – Rappler.com How does this make you feel? Loading

Original Source

Read the full article at Rappler →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.