Automated LLM red teaming gets a learning layer
ID: 1ae52b76-cbc1-5d26-844f-ad21efd94ab0
STIX ID: report--1ae52b76-cbc1-5d26-844f-ad21efd94ab0
Feed Name: Help Net Security
Research summary: Capital One AI Foundations proposes Adaptive Instruction Composition, which replaces random combination of harmful queries and jailbreak tactics with a small contextual bandit using SBERT embeddings to learn promising attack combinations. In 10,000-trial simulations the approach more than doubled success rates versus WildTeaming on three open-weight target models, found broader classes of successful queries, and demonstrated transferability between models; the paper notes evaluator limitations, compute costs, and dual-use risks and recommends responsible disclosure.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
