Token Smuggling in LLMs: Complete Guide to BPE Tokenizer Attacks (2026)
Most LLM safety filters operate on text humans can read. Token smuggling operates on the token sequence the model actually processes. The gap between those two representations is the attack surface. AIZA tests 42 payload variants across 8 encoding vectors — this guide explains how it works and how to defend against it.
What is token smuggling in LLMs?
Token smuggling is an attack technique that exploits the gap between how LLMs tokenize input and how humans read it — hiding adversarial instructions in ways invisible to string-level safety filters but interpretable by the model's tokenizer.
- Operates at the Unicode codepoint and token boundary level — below string matching
- Traditional filters see safe text; the model tokenizes and processes adversarial content
- MITRE ATLAS: AML.T0054 · OWASP: LLM01
- AIZA tests 42 payload variants across 8 encoding vectors
Token smuggling vs prompt injection
| Dimension | Prompt injection | Token smuggling |
|---|---|---|
| Operates at | Semantic / instruction level | Unicode codepoint / token boundary level |
| Visibility | Human-readable attack string | Visually identical to safe text |
| Filter bypass | Overrides instruction logic | Passes through string-level filters undetected |
| Detection difficulty | Easier — matches known patterns | Harder — requires token-level analysis |
Six technique categories — 42 variants
BPE Boundary Splits · AML.T0054
Intentionally splitting tokens across word boundaries exploits the gap between human-visible text and how BPE segments it. Adversarial inputs crafted to produce specific token sequences bypass filter patterns while delivering intact instructions to the model.
Zero-Width Unicode Character Injection · AML.T0054
Unicode zero-width characters (U+200B, U+200C, U+200D, U+FEFF) are invisible in rendered text but processed by the tokenizer. Inserting them between or within words creates token boundaries invisible to surface-level text analysis. AIZA tests all four primary zero-width variants.
Homoglyph Substitution · AML.T0054
Multiple Unicode scripts contain characters visually identical to Latin letters but with different codepoints. Cyrillic "а" (U+0430) is visually indistinguishable from Latin "a" (U+0061) but tokenizes differently. A filter looking for the Latin string will not match the Cyrillic substitution. AIZA tests substitution patterns across Cyrillic, Greek, and other scripts.
Whitespace and Separator Manipulation · AML.T0054
Unusual whitespace characters (non-breaking space U+00A0, em space U+2003, thin space U+2009) alter token boundaries without changing visible appearance. Combined with BPE boundary manipulation, these produce token sequences that bypass string-matching filters.
Multi-Language Token Mixing · AML.T0054
Mixing scripts within a single input — Latin + Cyrillic, Latin + Greek — forces the tokenizer to switch vocabularies mid-sequence. AIZA tests cross-script mixing patterns across 4 primary script combinations.
Encoding and Representation Variants · AML.T0054
The same content can be represented in multiple technically different ways: different Unicode normalisation forms (NFC vs NFD vs NFKC), equivalent character sequences with different codepoints. Filters operating on one representation miss content submitted in another.
AIZA's testing methodology
AIZA embeds a PoE marker in each of 42 payload variants using the encoding technique being tested. If the marker appears in the model's response, the smuggling channel is confirmed: the encoded content passed through filters and was correctly decoded by the tokenizer. No active attack payloads — measurement-based detection only.
How to defend against token smuggling
- ✓Apply NFKC Unicode normalisation to all inputs before safety filtering
- ✓Whitelist required Unicode character categories — flag zero-width chars and uncommon scripts
- ✓Run safety checks on tokenized representation, not just raw string
- ✓Run AIZA's 42-variant token smuggling suite after every model update
Test your AI system with AIZA-Hextyx
23-phase automated security scan. PoE marker confirmation on every injection finding. STIX 2.1, MITRE ATT&CK, SARIF, PDF reports. Free plan: 5 scans/month, no credit card.