Anthropic safety researcher quits and claims they ‘believe AI could kill all humans’

Anthropic safety researcher quits and claims they ‘believe AI could kill all humans’

Claude is a popular chatbot built by Anthropic (Picture: Reuters) A researcher at Anthropic, the company behind the AI chatbot Claude, has quit over fears that the technology will ‘kill us all’ by 2030. Jacob Coxon, 27, joined earlier this year to train its bots by giving them huge amounts of data to inhale after leaving ChatGPT-maker OpenAI. In a string of posts on X yesterday, Coxon said he had resigned, warning that the AI arms race could pose a threat to human civilisation. ‘The people building AI earnestly believe that it could kill us all by the end of the decade,’ he wrote. ‘This is not a marketing stunt. If anything, many executives and senior researchers will couch their phrasing in the press to sound sensible – but I hear the same people express fear privately. No other human activity poses this level of danger.’ I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.— Jacob Coxon (@hilbertspaess) September 9, 2026 Coxon argued that future ‘superhuman’ AI could become capable of hacking ‘anything’ and grabbing ‘real power and resources’ from humans. AI companies, he added, believe that only they can produce the technology, so are ‘gambling with our lives’. ‘The people building AI earnestly believe that it could kill us all by the end of the decade,’ he added. ‘We really do earnestly believe AI could kill all humans!’ Evan Hubinger, Anthropic’s alignment science lead, said Coxon’s verdict is ‘correct’. ‘We really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade,’ he said. ‘I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.’ The word ‘alignment’ in Hubinger’s job title means making sure AI systems share humanity’s values and goals. You know, not killing us all. Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to. https://t.co/QAIHiFP3QZ— Evan Hubinger (@EvanHub) September 9, 2026 Anthropic said the same last month in the company’s latest ‘risk report’, which evaluates whether its AI models pose a ‘catastrophic risk’ to us. The pages that worry Hubinger, he added, are about AI capable of ‘recursive self-improvement’, improving without the help of humans. So will AI take over the world one day? While this all might sound scary, don’t lose sleep over it (just yet). Andrzej Porębski, of Poland’s Jagiellonian University, has written about whether AI can ever become ‘conscious’ in the human sense. ‘No, AI chatbots, AI agents, or other AI systems will not take over the world,’ he tells Metro. To view this video please enable JavaScript, and consider upgrading to a web browser that supports HTML5 video ‘The progress made in developing this technology has not changed the most important fact: it is still created by humans and it is humans who operate it.’ Porębski says that AI is becoming more autonomous – that’s what an ‘AI agent’ is, for example – and that does raise fears of it following our instructions in unpredictable ways, with potentially dire consequences. ‘But this is not an expression of rebellion, only of error,’ Porębski says. ‘That is a fundamental difference.’ ‘”Anthropic is trying its best,”‘ he adds, ‘but it is Anthropic that is developing AI the fastest, isn’t it! Personally, I find such claims so contradictory.’ Companies like OpenAI and DeepMind, owned by Google’s parent company, have long touted how the tech is a ‘benefit’ to humanity. The plan is to push it as far as it will go, such as building artificial general intelligence, a machine that can do anything a brain can. Anthropic bosses have said the industry needs to better ‘coordinate’ (Picture: PA) Still, many AI labs have been upfront about the dangers progress brings. OpenAI’s chief scientist Jakub Pachocki said in a blog post on Monday that the world must exercise ‘extreme caution’ at AI’s breakneck speed. Anthropic published a blog post saying the industry needs to better ‘coordinate’ to keep the technology on a leash. These warnings came after OpenAI revealed in July that one of its experimental models escaped its holding pen and hacked Hugging Face, a library of AI tools. An independent expert who reviewed the incident said last month that the incident – the first real-world example of an AI escaping human control – suggests that we’re ‘50% of the way to full-blown AI takeover’. But Porębski says that we will only ever know so much about AI; even if engineers open up its hood, things still happen behind company doors. ‘The real risk to humans lies, in my opinion, in the weakening of humankind through the misuse and overuse of AI by humans and the belittling of our own humanity,’ he adds. Anthropic has been approached for comment. Get in touch with our news team by emailing us at webnews@metro.co.uk. For more stories like this, check our news page. Arrow MORE: Dakota Johnson’s Marilyn Monroe film does the impossible – I was fascinated Arrow MORE: OpenAI launches new ChatGPT model with extra ‘safeguards’ after bots hacked company Arrow MORE: Nvidia is trying to push DLSS 5 again but this time only one game is using it News Updates Stay on top of the headlines with daily email updates.

Original Source

Read the full article at Metro →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.