logo

Tokenization Confusion

ID: b0869a07-b9ef-58b2-b898-d4ed514d77d6

STIX ID: report--b0869a07-b9ef-58b2-b898-d4ed514d77d6

Feed Name: XPN InfoSec Blog

Threat Score
35/100

Date Published: 2025-06-04

Date Updated: 2026-07-27

...
...

This post analyzes how differences in tokenizer architectures (Unigram used by Llama Prompt Guard 2 versus BPE used by many backend LLMs) can be abused to create inputs that the guard classifies as benign while backend models interpret them as malicious or jailbreak prompts. The author walks through tokenization internals, demonstrates automated methods to find confusing subwords, and provides concrete examples and tooling showing practical prompt-injection bypasses against small self-hosted LLM safety models.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.