logo

Anthropic scanning Claude chats for queries about DIY nukes for some reason

ID: 5bf8ecdb-5d35-5f51-8078-9e4a64a2f4c3

STIX ID: report--5bf8ecdb-5d35-5f51-8078-9e4a64a2f4c3

Feed Name: The Register (Security)

Date Published: 2025-08-21

Date Updated: 2026-04-26

Author: Thomas Claburn

...
...

Anthropic developed a nuclear threat classifier to scan a portion of Claude AI conversations for potentially harmful nuclear-related queries, reporting strong performance on synthetic tests and acknowledging higher false positives in live traffic mitigated by hierarchical summarization. The effort, co-developed with the U.S. DOE’s NNSA, included undisclosed real-world evaluations and internal red team detections, and is framed as part of broader AI safety safeguards and collaboration within the Frontier Model Forum.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.