logo

Researchers Discover Major Security Gaps in LLM Guardrails

ID: 5c018b05-6039-55ed-bc25-3c96b52963ea

STIX ID: report--5c018b05-6039-55ed-bc25-3c96b52963ea

Feed Name: Infosecurity Magazine (News)

Threat Score
70/100

Date Published: 2026-03-11

Date Updated: 2026-04-22

...
...

Unit 42 demonstrated that LLMs used as safety "AI Judges" can be manipulated by AdvJudge-Zero, an automated fuzzer that discovers low-perplexity token sequences which steer models to approve policy-violating content; the researchers reported a ~99% bypass success rate across common model architectures and recommend adversarial training to harden systems.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.