US models beat China’s Kimi K3 with a 76% score over 32% in cyber benchmarks, tests show

US models beat China’s Kimi K3 with a 76% score over 32% in cyber benchmarks, tests show

Despite fears that China could race ahead of the US in AI, a joint UK-US evaluation found that Moonshot AI’s Kimi K3 performs “significantly below” leading American models in offensive cybersecurity capabilities. The findings address growing concerns in Washington regarding China’s recent advancements in artificial intelligence following the release of Kimi K3, described as China’s most powerful large language model. The UK Artificial Intelligence Security Institute (UK AISI) and the US Center for AI Standards and Innovation (CAISI) found that Kimi K3 achieved an overall cyber capability score of 32.2%, compared to an average of 76.2% for top, unnamed US models. Evaluation identifies cybersecurity gap Researchers tested Kimi K3 using ExploitBench, a public benchmark developed by Carnegie Mellon University that measures an AI system’s ability to develop end-to-end exploits for software vulnerabilities. According to the evaluation, Kimi K3 “performs significantly below the most recent frontier cyber-capable models” when tasked with developing exploits. The model attained a score of 32 percent, outperforming GLM-5.2, another Chinese open-weight model, which scored 24 percent. However, Kimi K3 still lagged behind leading US systems. The most cyber-capable US models achieved an average score of 76.2%. The report also examined whether models could achieve arbitrary code execution (ACE)—a high-severity exploit outcome that allows attackers to take control of a target system. Kimi K3 achieved ACE on none of the 41 ExploitBench tasks, while “the most cyber-capable models achieved ACE on 20/41 samples on average,” according to the UK AISI and CAISI. Simulated attack shows remaining capabilities The evaluation also tested Kimi K3 on “The Last Ones” (TLO), a 32-step simulated corporate network attack designed to measure a model’s ability to conduct end-to-end cyber operations autonomously. Kimi K3 reached step 17 of the attack path on average, compared with the 28.5 steps achieved by the most cyber-capable US models. Despite this gap, the report noted that Kimi K3 completed the simulated attack in one of 10 attempts. The UK AISI and CAISI said this demonstrated that the model “is capable of autonomously attacking small, weakly defended, and vulnerable enterprise systems, when directed to do so and given initial network access.” Furthermore, Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during testing. Although the results were only preliminary, the research noted that Kimi K3’s overall cyber capability score was estimated from a single benchmark, whereas other models were evaluated across a larger number of cyber tasks. Findings reshape AI competition debate Kimi K3 has sparked attention, if not concern, following reports of its strong performance on other AI benchmarks. However, these new findings offer a different picture of the model’s cybersecurity capabilities, showing that it still lags behind US systems. The South China Morning Post reported that the evaluation “offers a stark reality check amid mounting panic in Washington over Beijing’s AI advancements following Moonshot AI’s latest flagship release.” David Sacks, former White House AI czar and current co-chair of US President Donald Trump’s Council of Advisers on Science and Technology, said the findings suggest that fears over Chinese model supremacy may be exaggerated. He urged policymakers to avoid over-regulating the sector. “The Kimi Panic needs to stop,” Sacks said in a social media post. “American frontier models are still ahead.” Recommended ArticlesGet the latest in engineering, tech, space & science - delivered daily to your inbox.Originally from LA, Maria Mocerino has been published in Business Insider, The Irish Examiner, The Rogue Mag, Chacruna Institute for Psychedelic Plant Medicines, and now Interesting Engineering.

Original Source

Read the full article at Interestingengineering →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.