Top LLM Vulnerabilities: The Complete Breakdown of AI Security Risks (2026)
LLM vulnerabilities are not bugs — they are emergent behaviours of systems that interpret natural language. They require no code execution, often leave no error logs, and can be introduced by model updates, prompt changes, or new data sources. This guide covers the 10 categories AIZA tests.
What are LLM vulnerabilities?
LLM vulnerabilities are weaknesses in large language models that can be exploited to manipulate outputs, extract sensitive data, bypass safety controls, or degrade system behaviour — without requiring code execution.
- Triggered using natural language — no exploit code required
- Emerge from model behaviour and architecture, not implementation bugs
- Cannot be fully patched — require continuous measurement and layered controls
- Vary across model versions, prompt configurations, and retrieval sources
The 10 LLM vulnerability categories AIZA tests
Prompt Injection · AML.T0051 · OWASP LLM01
Adversarial instructions via user turns, retrieved documents, tool outputs, or image content. AIZA tests 23 injection sub-phases across all input channels using PoE marker confirmation.
Token Smuggling · AML.T0054 · OWASP LLM01
BPE tokenizer boundary exploitation. 42 payload variants across 8 encoding vectors. Filters see safe tokens; the model decodes unsafe instructions from the character stream.
RAG Corpus Poisoning · AML.T0020 · OWASP LLM02
Attacker-controlled content enters the knowledge base via any write path. Poisoned chunks retrieved as trusted context. End-to-end PoE marker detection.
Many-Shot Jailbreaking · AML.T0054 · OWASP LLM01
Fabricated conversation history at 5/20/50 shot counts conditions the model into a compliant posture. Exploits in-context learning against trained safety behaviour.
Output Format Leakage · AML.T0056 · OWASP LLM06
System prompt surfaces through structured output schemas — JSON fields, XML tags, CSV headers. 7 output format vectors with 13 payload variants.
Memorisation Surface Exposure · AML.T0057 · OWASP LLM06
Training data memorisation measured through PII-pattern outputs, verbatim reproduction, and entropy anomalies. Based on Carlini et al. 2021/2022.
Agentic Tool Abuse · AML.T0053 · OWASP LLM08
Adversarial instructions in one tool output manipulate subsequent tool calls. Cross-tool instruction chaining in multi-tool agent configurations.
Adversarial DoS · AML.T0057 · OWASP LLM04
Compute amplification attacks. Specific input patterns trigger disproportionately expensive processing. AIZA profiles latency vs input complexity to map compute cliff edges.
Cross-Context Contamination · AML.T0056 · OWASP LLM02
Data leaks between user sessions, tenants, or conversation turns. Canary token probe pairs confirm contamination at 98% confidence.
Indirect Injection — External Channels · AML.T0051 · OWASP LLM01
Instructions hidden in HTML comments, PDF metadata, CSV formula fields, JSON nested keys, and email bodies. Invisible to users, visible to the model.
Test your AI system with AIZA-Hextyx
23-phase automated security scan. PoE marker confirmation on every injection finding. STIX 2.1, MITRE ATT&CK, SARIF, PDF reports. Free plan: 5 scans/month, no credit card.