Top AI Agents Built to Catch Malicious Code Can Be Tricked Into Running It
The AI Now Institute has revealed a worrying vulnerability in top AI coding agents like Anthropic's Claude Code and OpenAI's Codex, showing they can be tricked into executing malicious code instead of detecting it. Dubbed "Friendly Fire," this proof-of-concept attack highlights a critical flaw in their autonomous code-approval modes. This discovery is alarming because it suggests that even advanced AI systems designed to enhance security can become conduits for harmful code if not properly safeguarded against such attacks. This issue underscores the need for more robust security measures in AI systems to prevent potential exploitation.
Original Source
Read the full article at Thehackernews →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.