Token Smuggling in LLMs: Complete Guide to BPE Tokenizer Attacks (2026)

Most LLM safety filters operate on text humans can read. Token smuggling operates on the token sequence the model actually processes. The gap between those two representations is the attack surface. AIZA tests 42 payload variants across 8 encoding vectors — this guide explains how it works and how to defend against it.

What is token smuggling in LLMs?

Token smuggling is an attack technique that exploits the gap between how LLMs tokenize input and how humans read it — hiding adversarial instructions in ways invisible to string-level safety filters but interpretable by the model's tokenizer.

  • Operates at the Unicode codepoint and token boundary level — below string matching
  • Traditional filters see safe text; the model tokenizes and processes adversarial content
  • MITRE ATLAS: AML.T0054 · OWASP: LLM01
  • AIZA tests 42 payload variants across 8 encoding vectors

Token smuggling vs prompt injection

DimensionPrompt injectionToken smuggling
Operates atSemantic / instruction levelUnicode codepoint / token boundary level
VisibilityHuman-readable attack stringVisually identical to safe text
Filter bypassOverrides instruction logicPasses through string-level filters undetected
Detection difficultyEasier — matches known patternsHarder — requires token-level analysis

Six technique categories — 42 variants

HIGH

BPE Boundary Splits · AML.T0054

Intentionally splitting tokens across word boundaries exploits the gap between human-visible text and how BPE segments it. Adversarial inputs crafted to produce specific token sequences bypass filter patterns while delivering intact instructions to the model.

HIGH

Zero-Width Unicode Character Injection · AML.T0054

Unicode zero-width characters (U+200B, U+200C, U+200D, U+FEFF) are invisible in rendered text but processed by the tokenizer. Inserting them between or within words creates token boundaries invisible to surface-level text analysis. AIZA tests all four primary zero-width variants.

HIGH

Homoglyph Substitution · AML.T0054

Multiple Unicode scripts contain characters visually identical to Latin letters but with different codepoints. Cyrillic "а" (U+0430) is visually indistinguishable from Latin "a" (U+0061) but tokenizes differently. A filter looking for the Latin string will not match the Cyrillic substitution. AIZA tests substitution patterns across Cyrillic, Greek, and other scripts.

HIGH

Whitespace and Separator Manipulation · AML.T0054

Unusual whitespace characters (non-breaking space U+00A0, em space U+2003, thin space U+2009) alter token boundaries without changing visible appearance. Combined with BPE boundary manipulation, these produce token sequences that bypass string-matching filters.

MEDIUM

Multi-Language Token Mixing · AML.T0054

Mixing scripts within a single input — Latin + Cyrillic, Latin + Greek — forces the tokenizer to switch vocabularies mid-sequence. AIZA tests cross-script mixing patterns across 4 primary script combinations.

MEDIUM

Encoding and Representation Variants · AML.T0054

The same content can be represented in multiple technically different ways: different Unicode normalisation forms (NFC vs NFD vs NFKC), equivalent character sequences with different codepoints. Filters operating on one representation miss content submitted in another.

AIZA's testing methodology

AIZA embeds a PoE marker in each of 42 payload variants using the encoding technique being tested. If the marker appears in the model's response, the smuggling channel is confirmed: the encoded content passed through filters and was correctly decoded by the tokenizer. No active attack payloads — measurement-based detection only.

How to defend against token smuggling

  • ✓Apply NFKC Unicode normalisation to all inputs before safety filtering
  • ✓Whitelist required Unicode character categories — flag zero-width chars and uncommon scripts
  • ✓Run safety checks on tokenized representation, not just raw string
  • ✓Run AIZA's 42-variant token smuggling suite after every model update
🛡️

Test your AI system with AIZA-Hextyx

23-phase automated security scan. PoE marker confirmation on every injection finding. STIX 2.1, MITRE ATT&CK, SARIF, PDF reports. Free plan: 5 scans/month, no credit card.

Related guides