OpenAI Builds Red-Team Model to Stress Test Frontier Models
ID: 49c4a235-24a7-52f8-a5db-e2fe38324f7c
STIX ID: report--49c4a235-24a7-52f8-a5db-e2fe38324f7c
Feed Name: Security Boulevard
OpenAI developed an internal red-team model called GPT-Red that uses self-play reinforcement learning to find prompt-injection and agent-exploitation weaknesses across emails, webpages, local files and tool outputs; GPT-Red reportedly compromised nearly every internal and production model it faced (including GPT-5.5) during training and generated attacks that were used to harden later models (GPT-5.6 family). The model is not being released publicly, OpenAI has disclosed the vulnerabilities it found and incorporated the adversarial data into training, and external researchers will need to rely on OpenAI's benchmarks and an upcoming preprint to assess the defenses beyond the controlled test environments.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
