logo

How Anthropic’s Jailbreak Challenge Put AI Safety Defenses to the Test

ID: f8e34840-06de-5520-aba9-30e3af4e011b

STIX ID: report--f8e34840-06de-5520-aba9-30e3af4e011b

Feed Name: HackerOne Blog

Date Published: 2025-07-08

Date Updated: 2026-06-12

...
...

**Executive Summary:** Anthropic partnered with HackerOne for a week-long AI red teaming challenge against a demo of Claude 3.5 Sonnet to evaluate Constitutional Classifiers for CBRN-related safety bypasses; 339 participants produced over 300,000 chat interactions, several teams earned bounties for successful jailbreaks, and the findings highlighted attack techniques and areas for improving model defenses.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.