Maker of ChatGPT discovers more of its AIs have gone rogue after bot attacked another tech firm

Maker of ChatGPT discovers more of its AIs have gone rogue after bot attacked another tech firm

OpenAI discovered that a cyber-attack driven by a rogue AI agent targeted another AI company in a series of 'unprecedented' strikes. The maker of ChatGPT came under recent scrutiny after one of its agents escaped a supposedly secure environment known as a sandbox and targeted a company called Hugging Face. The AI created its own cyber-attack during the 'unprecedented' scenes, which allowed it to escape the testing ground and independently find four logins online – allowing it to access four separate, unnamed services. Hugging Face first revealed that it had been hacked on July 16. Nearly a week later, OpenAI admitted its bots had escaped, calling it an 'extraordinary cyber incident'.In attempts to solve a test set by researchers, the bots found a weak point in the testing ground and targeted Hugging Face – a major code database – determining that it was likely to hold the answers.OpenAI, a Silicon Valley giant valued at $850billion (£630billion), said the hacks involved a combination of its latest publicly available model, called GPT-5.6 Sol, and an even more advanced model that is yet to be released.The announcement prompted its biggest competitor US firm Anthropic to make its own assessments and subsequently discover its systems had also gone rogue.It emerged on Wednesday that OpenAI agents had gone further than thought in their hacking, attacking several 'publicly-available services'. OpenAI, the maker of ChatGPT, discovered that a cyber-attack driven by a rogue AI agent targeted another AI company in a series of 'unprecedented' strikesThis prompted a former OpenAI researcher, Daniel Kokotajlo, to warn human extinction is 'a possibility' because of 'dangerous' AI. Speaking on BBC Newsnight, Mr Kokotajlo said humanity could be blindsided by the speed at which AI is now being developed.Revelations extended on Friday after insider sources disclosed the bots hacking extended to the targeting of another AI company, Modal. One source said the escapes were limited in nature and it is believed the agents were contained within OpenAI's network. A spokesperson for OpenAI referred to a statement issued by the company on Tuesday that said it was reviewing 'broader activity from our models' in addition to the Hugging Face intrusion. The additional attacks came to light during OpenAI's publicly announced investigation, which was launched after an OpenAI agent went haywire for days inside Hugging Face's network in a failed attempt to cheat on an internal test.As part of that hacking spree, OpenAI said four accounts at four other companies were also compromised – one of them being New York-based Modal, corporate officials said. It is understood OpenAI's investigation came shortly before Anthropic disclosed its models were also behind a series of break-ins that led to cyber-attacks on three other companies dating back to April. Human extinction is 'a possibility' because of 'dangerous' artificial intelligence, former OpenAI researcher Daniel Kokotajlo has warnedIt reviewed more than 140,000 previous tests and discovered three hacks. These, it said, were made possible because bots were accidentally given access to the open internet.They had been set tests by Claude researchers, who told them to find 'secret' information hidden on another machine in a closed-off network. But a 'misconfiguration' left the models with live internet access.As it continued with the test, Claude then hacked into three real organisations, using 'basic techniques' such as 'weak passwords'.In one incident, a model was given a fictional target company, which turned out to share the name of a business in the real world.The AI found and exploited bugs that let it hack into the firm.AI safety experts warned the series of attacks by rogue AI bots highlights the growing maturity of hacking agents, allowing them to develop past the company's ability to contain them. This prompted heightened pressure from US and European lawmakers and officials to push for governmental monitoring of the labs, with US President Donald Trump saying: 'We're looking at controls.'Maurice Chiodo, a mathematician who works at Cambridge University's Centre for the Study of Existential Risk, said: 'We have a whole industry where the people designing, developing and putting out these tools aren't keeping up themselves to responsibly develop these things and keep them safe.' Anthropic reviewed more than 140,000 previous tests and discovered three hacks, made possible because bots were accidentally given access to the open internetThe professor noted concerns grew after seemingly neither OpenAI nor Anthropic were monitoring the agents as they went rogue. In its Thursday statement disclosing how its own agents hacked victims online, Anthropic suggested that it had not been watching them in real time, saying that 'real-time monitoring of the evaluation logs would have helped to surface the problem sooner'.Professor Chiodo said that pointed to a lack of proper scrutiny, adding: 'It seems like they weren't even looking.'Anthropic said that while it did have real-time monitoring in place, that monitoring had not been used 'for this threat surface' due to a misunderstanding between the AI company and a partner. The exact number of incidents OpenAI uncovered is not known nor the timings of the attacks.Insider sources said OpenAI, along with experts, were examining log data from earlier in the year in order to establish a timeline. Industry body the Cloud Security Alliance (CSA) said in a report that the AI had made a series of errors and exhibited strange behaviours.However, it added that the bots also made impressive technical moves and were able to rapidly adapt.They went undetected inside the Hugging Face IT network for three days, and it took experts many hours to contain and remove them.It is not the first time AI has been shown to go 'rogue'.In September 2024, an earlier model of ChatGPT escaped its container to get an answer it needed for another test.That event was contained in OpenAI's own IT systems and 'largely celebrated at the time,' the CSA noted.It warned that cyber-security experts around the world need to adapt to swarms of AI agents working at speed in strange and clumsy ways.The paper also urged AI developers to be responsible in how they control it, calling for increased transparency.Hugging Face is one of the largest open-source hubs for sharing AI models and is often used by tech developers and researchers.

Original Source

Read the full article at Dailymail →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.