Tokenization Confusion
ID: b0869a07-b9ef-58b2-b898-d4ed514d77d6
STIX ID: report--b0869a07-b9ef-58b2-b898-d4ed514d77d6
Feed Name: XPN InfoSec Blog
This post analyzes how differences in tokenizer architectures (Unigram used by Llama Prompt Guard 2 versus BPE used by many backend LLMs) can be abused to create inputs that the guard classifies as benign while backend models interpret them as malicious or jailbreak prompts. The author walks through tokenization internals, demonstrates automated methods to find confusing subwords, and provides concrete examples and tooling showing practical prompt-injection bypasses against small self-hosted LLM safety models.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
