Show HN: I benchmarked LLM agents on fixing real-world security vulnerabilities
I built a benchmark with 20 real CVEs across 18 Python projects (Pillow, GitPython, yt-dlp, urllib3, etc). I've run it over 5 LLM agents (3 OpenAI, 2 poolside) and 3 different prompts (full advisory, locate, diagnose) with a total of 300 runs. The agents are tasked to fix security vulnerabilities in a sandboxed environment and they are scored against a hidden security tests from the maintainer's own fix.Best solve rate was 50%. On the other 50%, some fixes are sometimes coherent and pass all reg...
Original Source
Read the full article at Giovannigatti →KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.