Benchmarking 13 AI models on rediscovering known CVEs
ID: 1435bb73-19c8-5116-9b59-36b698297760
STIX ID: report--1435bb73-19c8-5116-9b59-36b698297760
Feed Name: Aikido Security's Blog
Threat Score
This benchmark evaluates 13 LLMs against 26 known CVEs using an AI code-analysis harness: GPT-5.6 achieved the highest recall (23/26), but cheaper mid-tier models can match results by pooling multiple runs, and open-weight models (GLM-5.2, Kimi K3) are rapidly improving—overall the report focuses on detection capability, cost-effectiveness, and variance rather than active exploitation.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
