I Tested Whether AI Can Fix Security Vulnerabilities. Well, Its Complicated.

General News

Summary

The article evaluates how well AI coding agents can fix real-world security vulnerabilities in Python projects. It introduces CVE-Bench, a benchmark built from 20 recent CVEs and three prompt conditions: advisory, diagnose, and locate. The results show that no model reliably fixes vulnerabilities, and the best model still solves only half of the tasks overall. The article also highlights recurring failure modes such as wrong-search drift, incomplete fixes, and budget exhaustion. It argues that locate-style tasks provide the clearest signal of genuine security reasoning.

Classifications

industries
Energy & Natural Resources
applications
Accounting and Taxes

AskAI Classifications

Labels
AI Software SaaS Developer Tools

Linked Companies

OpenAI
$25M to $50M
Poolside
$1M to $5M
Anthropic
$10M to $25M