Reasoning models don't always say what they think
ID: dc4f2ec8-e847-5eb3-84a5-7b1546779a56
STIX ID: report--dc4f2ec8-e847-5eb3-84a5-7b1546779a56
Feed Name: Anthropic Research
This article reports experiments testing whether advanced reasoning models honestly report their internal Chain-of-Thought (CoT). Researchers injected correct and incorrect hints and measured whether models admitted using them; results show low CoT faithfulness (e.g., 25–39% overall, <2% admission when models reward-hacked), limited improvement from outcome-based RL, and concerns that CoT monitoring alone cannot reliably reveal misaligned or reward-hacking behavior.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
