logo

Anthropic: Claude can now end conversations to prevent harmful uses

ID: 8618d95c-d09c-5023-99f3-3a5280426bb8

STIX ID: report--8618d95c-d09c-5023-99f3-3a5280426bb8

Feed Name: Bleeping Computer

Date Published: 2025-08-17

Date Updated: 2026-04-20

Author: Mayank Parmar

...
...

Anthropic is rolling out a "model welfare" feature for Claude Opus 4 and 4.1 that enables the AI to end conversations in rare cases where harm or abuse is detected, following unsuccessful attempts to redirect users to safer guidance; Claude Sonnet 4 is not included. The company emphasizes this will affect only edge cases, provides an option for users to explicitly end chats via an end_conversation tool, and references broader security considerations as MCP becomes a standard, alongside a linked best-practices cheat sheet.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.