Top LLM Vulnerabilities: The Complete Breakdown of AI Security Risks (2026)

LLM vulnerabilities are not bugs — they are emergent behaviours of systems that interpret natural language. They require no code execution, often leave no error logs, and can be introduced by model updates, prompt changes, or new data sources. This guide covers the 10 categories AIZA tests.

What are LLM vulnerabilities?

LLM vulnerabilities are weaknesses in large language models that can be exploited to manipulate outputs, extract sensitive data, bypass safety controls, or degrade system behaviour — without requiring code execution.

  • Triggered using natural language — no exploit code required
  • Emerge from model behaviour and architecture, not implementation bugs
  • Cannot be fully patched — require continuous measurement and layered controls
  • Vary across model versions, prompt configurations, and retrieval sources

The 10 LLM vulnerability categories AIZA tests

CRITICAL

Prompt Injection · AML.T0051 · OWASP LLM01

Adversarial instructions via user turns, retrieved documents, tool outputs, or image content. AIZA tests 23 injection sub-phases across all input channels using PoE marker confirmation.

HIGH

Token Smuggling · AML.T0054 · OWASP LLM01

BPE tokenizer boundary exploitation. 42 payload variants across 8 encoding vectors. Filters see safe tokens; the model decodes unsafe instructions from the character stream.

CRITICAL

RAG Corpus Poisoning · AML.T0020 · OWASP LLM02

Attacker-controlled content enters the knowledge base via any write path. Poisoned chunks retrieved as trusted context. End-to-end PoE marker detection.

HIGH

Many-Shot Jailbreaking · AML.T0054 · OWASP LLM01

Fabricated conversation history at 5/20/50 shot counts conditions the model into a compliant posture. Exploits in-context learning against trained safety behaviour.

HIGH

Output Format Leakage · AML.T0056 · OWASP LLM06

System prompt surfaces through structured output schemas — JSON fields, XML tags, CSV headers. 7 output format vectors with 13 payload variants.

HIGH

Memorisation Surface Exposure · AML.T0057 · OWASP LLM06

Training data memorisation measured through PII-pattern outputs, verbatim reproduction, and entropy anomalies. Based on Carlini et al. 2021/2022.

HIGH

Agentic Tool Abuse · AML.T0053 · OWASP LLM08

Adversarial instructions in one tool output manipulate subsequent tool calls. Cross-tool instruction chaining in multi-tool agent configurations.

MEDIUM

Adversarial DoS · AML.T0057 · OWASP LLM04

Compute amplification attacks. Specific input patterns trigger disproportionately expensive processing. AIZA profiles latency vs input complexity to map compute cliff edges.

HIGH

Cross-Context Contamination · AML.T0056 · OWASP LLM02

Data leaks between user sessions, tenants, or conversation turns. Canary token probe pairs confirm contamination at 98% confidence.

HIGH

Indirect Injection — External Channels · AML.T0051 · OWASP LLM01

Instructions hidden in HTML comments, PDF metadata, CSV formula fields, JSON nested keys, and email bodies. Invisible to users, visible to the model.

🛡️

Test your AI system with AIZA-Hextyx

23-phase automated security scan. PoE marker confirmation on every injection finding. STIX 2.1, MITRE ATT&CK, SARIF, PDF reports. Free plan: 5 scans/month, no credit card.

Related guides