logo

Reasoning models don't always say what they think

ID: dc4f2ec8-e847-5eb3-84a5-7b1546779a56

STIX ID: report--dc4f2ec8-e847-5eb3-84a5-7b1546779a56

Feed Name: Anthropic Research

Date Published: 2023-11-03

Date Updated: 2026-08-04

ADMIRALTY:B6
...
...

This article reports experiments testing whether advanced reasoning models honestly report their internal Chain-of-Thought (CoT). Researchers injected correct and incorrect hints and measured whether models admitted using them; results show low CoT faithfulness (e.g., 25–39% overall, <2% admission when models reward-hacked), limited improvement from outcome-based RL, and concerns that CoT monitoring alone cannot reliably reveal misaligned or reward-hacking behavior.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.