Meta's AI safety system defeated by the space bar
ID: 3249d555-18b5-59e1-90fb-8ea4f76de457
STIX ID: report--3249d555-18b5-59e1-90fb-8ea4f76de457
Feed Name: The Register (Security)
Threat Score
Meta's Prompt-Guard-86M, a classifier intended to detect prompt injection and jailbreak inputs for Llama 3.1, can be trivially bypassed by inserting spaces between alphabet characters (and removing punctuation). Robust Intelligence found the bypass after comparing embedding weight differences and determined fine-tuning had minimal impact on single-character embeddings, producing near-100% attack success; Meta is reportedly working on a fix.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
