logo

We burned 11.7bn tokens to find the best cyber AI model

ID: 922fb0f0-eb6c-533e-8f2f-76dee9b089f6

STIX ID: report--922fb0f0-eb6c-533e-8f2f-76dee9b089f6

Feed Name: Aikido Security's Blog

Threat Score
0/100

Date Published: 2026-08-21

Date Updated: 2026-08-21

...
...

This benchmark evaluates 10 AI models (three runs each) against 32 recently disclosed CVEs in source repositories, comparing pooled and single-run recall, precision, exploration behavior, and cost; DeepSeek V4 Pro achieved the highest pooled recall (28/32) while open-weight models closed much of the gap with frontier models, and repetition improved coverage but increased false leads, highlighting the need for harness engineering to balance recall, precision, and triage costs.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.