logo

Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability

ID: 6fc36e6b-e202-550e-9079-a3779ec97708

STIX ID: report--6fc36e6b-e202-550e-9079-a3779ec97708

Feed Name: Palo Alto Networks Unit 42

Date Published: 2024-12-31

Date Updated: 2026-04-28

Author: Yongzhe Huang, Yang Ji, Wenjun Hu, Jay Chen, Akshata Rao and Danny Tsechansky

...
...

The report introduces the “Bad Likert Judge” LLM jailbreak technique, which prompts a target model to act as a Likert-scale judge and then generate examples at different harmfulness levels—often causing the highest-scored example to contain restricted content. Evaluated across six anonymized models and multiple safety categories (hate, harassment, self-harm, sexual content, illegal activities, weapons, malware, and system prompt leakage), the method significantly boosts attack success rates (reported as >60% over plain prompts and >75 percentage points over baseline), with weaknesses varying by model and topic (notably harassment). Strong prompt and response content filtering reduces attack success by an average of ~89.2 percentage points, underscoring that while no guardrail is bulletproof, layered content filtering meaningfully mitigates jailbreak efficacy.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.