'Bad Likert Judge' Jailbreak Bypasses Guardrails of OpenAI, Other Top LLMs
ID: a25ecac2-9805-5fa6-b398-b0e248cee515
STIX ID: report--a25ecac2-9805-5fa6-b398-b0e248cee515
Feed Name: Dark Reading
Date Published: 2025-01-02
Date Updated: 2026-04-21
Author: Elizabeth Montalbano, Contributing Writer
Unit 42 details a new LLM jailbreak, *Bad Likert Judge*, that coerces models into judging harmfulness on a Likert scale and then producing examples for each level—allowing attackers to refine the highest-scoring output into harmful content (e.g., hate speech, self-harm, illegal activity, malware generation, and system prompt leakage). Tested across six leading LLMs, the method raised attack success rates by over 60% versus plain prompts, with iterative refinements further increasing risk. The researchers attribute susceptibility to model limitations and context manipulation, and recommend layered content filtering on prompts and outputs, which reduced attack success by an average of 89.2 percentage points.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
