logo

Automated LLM red teaming gets a learning layer

ID: 1ae52b76-cbc1-5d26-844f-ad21efd94ab0

STIX ID: report--1ae52b76-cbc1-5d26-844f-ad21efd94ab0

Feed Name: Help Net Security

Date Published: 2026-04-30

Date Updated: 2026-04-30

Author: Sinisa Markovic

...
...

Research summary: Capital One AI Foundations proposes Adaptive Instruction Composition, which replaces random combination of harmful queries and jailbreak tactics with a small contextual bandit using SBERT embeddings to learn promising attack combinations. In 10,000-trial simulations the approach more than doubled success rates versus WildTeaming on three open-weight target models, found broader classes of successful queries, and demonstrated transferability between models; the paper notes evaluator limitations, compute costs, and dual-use risks and recommends responsible disclosure.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.