logo

Sol Searching | Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?

Threat Score
30/100

Date Published: 2026-07-22

Date Updated: 2026-07-27

Author: Juan Andrés Guerrero-Saade & Gabriel Bernadett-Shapiro

...
...

SentinelLABS built an eight-stage reverse-engineering benchmark based on the 2005 sabotage implant “fast16” to evaluate frontier LLMs' ability to perform sustained, trustworthy malware investigations. The report details methodology, model runs, and findings: GPT-5.6 Sol completed the full benchmark and demonstrated project-scale recovery (withdrawing contradicted claims and propagating fixes), while GPT-5.5, GLM-5.2, and Opus 4.x showed local technical competence but failed to sustain closure across the entire investigation; the paper discusses costs, limitations, and the role of human supervision.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.