Sol Searching | Can Frontier Models Tackle Autonomous Long-Horizon Malware Analysis?
ID: 892a5749-6db7-5802-be89-eb5e7607521c
STIX ID: report--892a5749-6db7-5802-be89-eb5e7607521c
Date Published: 2026-07-22
Date Updated: 2026-07-27
Author: Juan Andrés Guerrero-Saade & Gabriel Bernadett-Shapiro
SentinelLABS built an eight-stage reverse-engineering benchmark based on the 2005 sabotage implant “fast16” to evaluate frontier LLMs' ability to perform sustained, trustworthy malware investigations. The report details methodology, model runs, and findings: GPT-5.6 Sol completed the full benchmark and demonstrated project-scale recovery (withdrawing contradicted claims and propagating fixes), while GPT-5.5, GLM-5.2, and Opus 4.x showed local technical competence but failed to sustain closure across the entire investigation; the paper discusses costs, limitations, and the role of human supervision.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
