logo

Benchmarking 13 AI models on rediscovering known CVEs

ID: 1435bb73-19c8-5116-9b59-36b698297760

STIX ID: report--1435bb73-19c8-5116-9b59-36b698297760

Feed Name: Aikido Security's Blog

Threat Score
15/100

Date Published: 2026-07-16

Date Updated: 2026-07-24

...
...

This benchmark evaluates 13 LLMs against 26 known CVEs using an AI code-analysis harness: GPT-5.6 achieved the highest recall (23/26), but cheaper mid-tier models can match results by pooling multiple runs, and open-weight models (GLM-5.2, Kimi K3) are rapidly improving—overall the report focuses on detection capability, cost-effectiveness, and variance rather than active exploitation.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.