logo

Measuring LLMs’ ability to develop exploits

ID: 610073fe-231b-5d46-a892-964e9382d1a8

STIX ID: report--610073fe-231b-5d46-a892-964e9382d1a8

Feed Name: Anthropic Research

Threat Score
78/100

Date Published: 2026-06-05

Date Updated: 2026-08-04

...
...

This Anthropic analysis evaluates Mythos Preview against three rigorous benchmarks (ExploitBench, ExploitGym, and SCONE-bench) and finds the model can reliably derive exploit primitives and construct end-to-end exploits. Key findings include Mythos achieving arbitrary code execution on many V8 CVEs (ACE on 21/41 in combined tests), hundreds of successful exploit tasks in ExploitGym (157 intended-vulnerability successes, 226 total flag captures), and simulated smart-contract thefts totaling ~$35M in SCONE—demonstrating substantial advancement in automated exploit development and a corresponding increase in cyber risk.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.