I compared Claude Opus 4.8 with 4.7 in a 10-round honesty test - and a legal prompt broke it
ID: cdd35471-dbb5-5bb8-aa83-46a9f74b6796
STIX ID: report--cdd35471-dbb5-5bb8-aa83-46a9f74b6796
Feed Name: ZDNet Security
**Executive summary:** A ZDNET evaluation compares Anthropic's Claude Opus 4.8 to Opus 4.7 across ten prompts measuring honesty, accuracy, and calibration; Opus 4.8 generally shows modest improvements and better restraint in several traps but still makes notable judgment and confidence errors (particularly on a legal/insurance demand-letter test), leading the author to conclude Opus 4.8 is an upgrade but not infallible.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
