Tokenization Confusion
ID: 216d8973-49a0-52f9-ae45-93e37e6e665d
STIX ID: report--216d8973-49a0-52f9-ae45-93e37e6e665d
Feed Name: SpecterOps Blog
This research post examines bypass techniques for Meta’s Llama Prompt Guard 2 by exploiting differences between Unigram and BPE tokenization, showing that prompts can be crafted to appear safe to the guard (via token splitting and normalization quirks) while remaining actionable to back-end LLMs. It details practical methods (hyphenation, prefix insertion, vocabulary-driven subword confusion, and interpretability tooling) to suppress high-importance tokens like “previous” and “instruction,” and confirms effectiveness against models such as Qwen, while noting limitations and related work on adversarial tokenization.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
