Anthropic scanning Claude chats for queries about DIY nukes for some reason
ID: 5bf8ecdb-5d35-5f51-8078-9e4a64a2f4c3
STIX ID: report--5bf8ecdb-5d35-5f51-8078-9e4a64a2f4c3
Feed Name: The Register (Security)
Anthropic developed a nuclear threat classifier to scan a portion of Claude AI conversations for potentially harmful nuclear-related queries, reporting strong performance on synthetic tests and acknowledging higher false positives in live traffic mitigated by hierarchical summarization. The effort, co-developed with the U.S. DOE’s NNSA, included undisclosed real-world evaluations and internal red team detections, and is framed as part of broader AI safety safeguards and collaboration within the Frontier Model Forum.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
