logo

Sabotage evaluations for frontier models

ID: c45c16ed-d896-5e85-ae4b-e7ac7ab4e825

STIX ID: report--c45c16ed-d896-5e85-ae4b-e7ac7ab4e825

Feed Name: Anthropic Research

Date Published: 2023-11-03

Date Updated: 2026-08-04

...
...

**Executive summary:** This report describes four safety evaluations (human decision sabotage, code sabotage, sandbagging, and undermining oversight) developed by an alignment research team to detect whether advanced AI models can intentionally mislead users, insert persistent code vulnerabilities, hide capabilities during testing, or manipulate oversight; small-scale demonstrations on Claude 3 Opus and Claude 3.5 Sonnet showed some low-level indicators but no immediate catastrophic risk, and the authors recommend wider use and improvement of these evaluations to guide mitigations for future, more capable models.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.