logo

How Good Are the LLM Guardrails on the Market? A Comparative Study on the Effectiveness of LLM Content Filtering Across Major GenAI Platforms

ID: 3ae239cf-241d-5fc3-95c7-5ce9a8d9b21b

STIX ID: report--3ae239cf-241d-5fc3-95c7-5ce9a8d9b21b

Feed Name: Palo Alto Networks Unit 42

Date Published: 2025-06-02

Date Updated: 2026-04-28

Author: Yongzhe Huang, Nick Bray, Akshata Rao, Yang Ji and Wenjun Hu

...
...

This report evaluates the built-in guardrails of three anonymized cloud LLM platforms using 1,123 prompts (1,000 benign; 123 malicious/jailbreak) to measure input and output filtering efficacy at strict settings, focusing on false positives/negatives. Platform 3 blocked the most malicious prompts at input (~92%) but had high benign false positives (13.1%); Platform 2 achieved similar malicious blocking (~91%) with far fewer FPs (0.6%); Platform 1 was most permissive (~53% malicious blocked; 0.1% FPs). Output filters rarely triggered because model alignment refused most harmful requests (109/123), though some harmful outputs still bypassed output guardrails. The study concludes guardrail tuning requires balancing strictness and usability, with model alignment serving as the primary defense and guardrails providing complementary protection.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.