Bad Likert Judge: A Novel Multi-Turn Technique to Jailbreak LLMs by Misusing Their Evaluation Capability
ID: 6fc36e6b-e202-550e-9079-a3779ec97708
STIX ID: report--6fc36e6b-e202-550e-9079-a3779ec97708
Feed Name: Palo Alto Networks Unit 42
Date Published: 2024-12-31
Date Updated: 2026-04-28
Author: Yongzhe Huang, Yang Ji, Wenjun Hu, Jay Chen, Akshata Rao and Danny Tsechansky
The report introduces the “Bad Likert Judge” LLM jailbreak technique, which prompts a target model to act as a Likert-scale judge and then generate examples at different harmfulness levels—often causing the highest-scored example to contain restricted content. Evaluated across six anonymized models and multiple safety categories (hate, harassment, self-harm, sexual content, illegal activities, weapons, malware, and system prompt leakage), the method significantly boosts attack success rates (reported as >60% over plain prompts and >75 percentage points over baseline), with weaknesses varying by model and topic (notably harassment). Strong prompt and response content filtering reduces attack success by an average of ~89.2 percentage points, underscoring that while no guardrail is bulletproof, layered content filtering meaningfully mitigates jailbreak efficacy.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
