️ RAG Security · Architecture Guide · 2026

Private RAG vs Public AI: The Complete Enterprise Decision Framework

Should your organisation build a private RAG system or use public AI APIs? The answer depends on your data sensitivity, compliance obligations, threat model, and query volume. This guide covers all four dimensions — architecture, security, compliance, and cost — so you can make the right call before you build.

In This Guide
1. What each model actually is 2. Architecture comparison 3. Security threat models 4. Compliance obligations 5. Cost comparison 6. Decision matrix 7. Hybrid architectures 8. Securing whichever you choose 9. FAQ

What Each Deployment Model Actually Is

Before comparing them, it helps to be precise about what "private RAG" and "public AI" mean in enterprise context — because both terms get used loosely.

Private RAG

A private RAG system has three components, all under your control:

  1. Your knowledge base — internal documents, databases, wikis, or any corpus you want the AI to reference, stored in a vector database you own (Pinecone, Weaviate, Chroma, pgvector)
  2. Your embedding pipeline — a process that converts your documents into vector embeddings and indexes them, running on infrastructure you control
  3. Your LLM — either a self-hosted open-source model (Llama 3, Mistral, Mixtral) or a private deployment of a commercial model (Azure OpenAI Service, AWS Bedrock, Anthropic Claude via private API)

In a true private RAG deployment, queries, retrieved documents, and responses never leave your infrastructure. Data residency is guaranteed. Audit trails are complete.

Public AI API

A public AI API deployment sends your prompts — including any context or retrieved content — to a third-party server for processing. OpenAI's API, Anthropic's API, Google's Gemini API. Your data travels to their infrastructure, gets processed by their models, and results are returned. You control the prompt engineering; you don't control where processing happens or what happens to data in transit.

This is a spectrum. "Public AI" can mean: a simple API call with no RAG (highest data exposure), a public API with your retrieved documents injected into the prompt (moderate exposure), or a public model deployed in your cloud tenant via a managed service like Azure OpenAI (reduced exposure, often sufficient for moderate sensitivity data).

The often-missed middle ground: Azure OpenAI Service, AWS Bedrock, and Google Vertex AI are technically "managed public AI" — you use OpenAI/Anthropic/Google models but data stays within your cloud tenant, not on the vendor's shared infrastructure. For many compliance use cases this middle path — managed cloud AI + private RAG — is the practical answer rather than fully self-hosted vs fully public.

Architecture Comparison

Dimension Private RAG️ Public AI API
Data custody Stays in your environmentSent to vendor infrastructure
Data residency Your chosen region/jurisdictionVendor-determined (may vary)
LatencyHigher (retrieval + local inference) Lower (vendor-optimised inference)
Setup complexityHigh — vector DB, embedding pipeline, LLM hosting Low — API key + prompt engineering
Maintenance burdenHigh — infrastructure, model updates, scaling Vendor-managed
Model qualityDepends on chosen model (Llama 3, Mistral, etc.) Frontier models (GPT-4o, Claude 3.5)
Knowledge freshness Real-time (you control indexing)Depends on RAG implementation
Customisation Full control (fine-tuning, RLHF)Limited to prompt engineering
Audit logging Complete — you own all logsPartial — vendor controls server logs
Offline capability Yes (air-gapped possible)No — requires internet
Vendor lock-in Minimal — swap models freelyHigh — API format/pricing dependency
Time to first queryWeeks to months Hours to days

Security Threat Models — Where Each Is Exposed

This is where most architecture comparisons fall short. Both deployment models have real security risks — they're just different risks. The mistake is assuming one is inherently more secure than the other.

Private RAG — Unique Threats

Critical
RAG Poisoning
Adversarial content planted in indexed documents that influences model responses when retrieved. Attacker needs write access to any indexed source — email, shared drives, document uploads.
Critical
Indirect Prompt Injection via Internal Docs
Malicious instructions embedded in internal documents, emails, or wikis. When retrieved, the LLM executes them invisibly. Private RAG surfaces are often trusted unconditionally — making this more dangerous than in public deployments.
High
Vector Database Compromise
Direct attack on the embedding store — exfiltrating embeddings to reconstruct document content, injecting malicious vectors, or manipulating cosine similarity to control what gets retrieved.
High
Retrieval Permission Bypass
Crafted queries that extract documents beyond the requester's authorisation scope. Most RAG implementations lack fine-grained access control at the chunk level — a user in Sales can retrieve HR documents if the query is crafted correctly.
Medium
Embedding Model Attacks
Adversarial inputs crafted to manipulate embedding similarity scores — causing specific documents to be retrieved or suppressed regardless of semantic relevance.

️ Public AI API — Unique Threats

Critical
Data-in-Transit Exposure
Every prompt sent to the vendor API travels over the internet. TLS protects against interception, but the vendor processes your data on their infrastructure — a breach of their systems exposes your prompts and retrieved content.
Critical
Vendor Training Data Leakage
Some API tiers use your prompts to improve models unless explicitly opted out. Sensitive data submitted via API may become training data accessible to other users through model memorisation.
High
Cross-Tenant Data Exposure
On shared vendor infrastructure, inadequate tenant isolation can expose one organisation's data to another. More theoretical than practical with major vendors, but a real risk on lower-tier API plans.
High
Vendor Security Dependency
Your security posture is bounded by the vendor's. A breach of the vendor's infrastructure, a zero-day in their API, or a rogue employee at the vendor directly impacts your data — with zero control on your end.
Medium
API Key Compromise
A leaked API key gives an attacker unlimited access to send arbitrary prompts — potentially extracting data from your RAG context, running up your bill, or poisoning your application's responses from the outside.

Shared Threats — Both Models Face These

Regardless of deployment model, these attacks apply to any AI system with an external interface:

The counterintuitive finding: Private RAG systems are often MORE vulnerable to indirect prompt injection than public AI deployments — because internal documents are implicitly trusted. A public AI deployment sending every prompt through content filtering catches many injection attempts. A private RAG system that trusts its internal knowledge base completely has no equivalent gate for documents that enter through indirect channels (email, document uploads, shared folders).

️ Assess Your RAG Security Posture Free

The HexTyx AI Security Assessment evaluates your RAG pipeline across 5 dimensions — chunk integrity, retrieval controls, vector DB access, source attribution, and anomaly monitoring. Whether private or public.

Compliance Obligations by Framework

Compliance is often the deciding factor in the private vs public debate — but the requirements are more nuanced than "regulated industry = private RAG". What matters is what data enters the AI system and what controls are in place.

FrameworkPublic AI APIManaged Cloud AIPrivate RAGKey Requirement
GDPR (EU) Conditional Viable Best fit Art. 46 SCCs required for non-EU transfers; DPA with vendor; right to erasure in vector DB
HIPAA (US Healthcare) BAA required Azure/AWS BAA available Best fit BAA with all AI vendors touching PHI; audit logging; minimum necessary PHI in prompts
EU AI Act High-risk: extra controls High-risk: extra controls Easier to demonstrate High-risk systems need technical docs, human oversight, robustness testing — all deployment models
SOC 2 Type II Viable with vendor SOC 2 Viable Full control Vendor SOC 2 report must be reviewed; AI activity in scope for availability + confidentiality criteria
FedRAMP (US Government) Usually blocked FedRAMP-authorised only Required for CUI CUI and federal data requires FedRAMP-authorised systems; most public AI APIs are not authorised
PCI DSS v4 CHD must not enter prompts Scoped deployment only With network segmentation Cardholder data (CHD) must never enter AI prompts on any system unless explicitly scoped and assessed
ISO 27001 AI Annex With supplier controls With supplier controls Full control Supplier risk assessment required for all AI vendors; AI-specific controls in ISMS scope
NIST AI RMF Viable Viable Viable Framework-agnostic — applies to all deployment models; documentation and governance requirements

The Three Compliance Rules That Actually Drive the Decision

  1. If PHI, CHD, or CUI enters the system — you need a BAA, PCI scope control, or FedRAMP authorisation respectively. Private RAG makes this significantly easier to manage. Public API requires a vendor who offers the relevant agreement and a rigorous data minimisation approach.
  2. If your users are in the EU and personal data enters prompts — GDPR Article 46 requires transfer safeguards for non-EU processing. Standard Contractual Clauses with OpenAI/Anthropic/Google are available but add legal overhead. Private RAG with EU-hosted infrastructure eliminates the transfer question entirely.
  3. If you're deploying high-risk AI under the EU AI Act — the deployment model is less important than your documentation and controls. You need technical documentation, human oversight mechanisms, and adversarial robustness testing regardless of whether you use private RAG or a public API. The controls are harder to demonstrate convincingly to a regulator with a public API where you don't own the infrastructure.

Cost Comparison — The Real Numbers

Cost is the most frequently misunderstood dimension of this decision. Private RAG appears expensive upfront; public AI appears cheap. At scale the equation reverses — but the break-even point is often higher than people expect.

Private RAG — Monthly Costs

GPU compute (embedding + inference)$800–$5,000
Vector database hosting$200–$800
Storage (documents + embeddings)$100–$500
MLOps / monitoring$200–$600
Engineering maintenance (0.5 FTE)$6,000–$10,000
Total (small deployment)$2,000–$8,000
Total (with eng overhead)$8,000–$18,000

️ Public AI API — Monthly Costs

GPT-4o at 10M tokens/mo$75–$100
GPT-4o at 100M tokens/mo$750–$1,000
GPT-4o at 1B tokens/mo$7,500–$10,000
Claude Sonnet at 100M tokens/mo$900–$1,500
Vector DB (Pinecone/Weaviate)$70–$400
Total (moderate usage, 100M tokens)$1,000–$2,000
Total (high usage, 1B tokens)$8,000–$11,000

Break-even reality check: At moderate usage (100M tokens/month — roughly 100,000 daily queries), public AI API costs $1,000–$2,000/month vs private RAG's $2,000–$8,000/month infrastructure cost excluding engineering. Private RAG doesn't become cost-competitive until you're processing 500M+ tokens/month AND have the engineering capacity to maintain it. For most enterprises deploying AI for the first time, public AI API (ideally via managed cloud) is significantly cheaper until you reach scale.

Decision Matrix — Which Model Is Right for You

Apply this matrix to your specific situation. Each row is a signal — where multiple signals point the same direction, that's your answer.

Your data contains PHI, PII, financial records, or trade secrets
Data that cannot leave your security perimeter under any circumstances
Private RAG
You operate under HIPAA, FedRAMP, or handle CUI
Regulatory frameworks that restrict or complicate third-party data processing
Private RAG
GDPR applies and personal data enters AI prompts
EU users, personal data in queries or retrieved context
Private RAG
You process 500M+ tokens per month
Scale where infrastructure cost becomes cheaper than per-token API pricing
Private RAG
You need guaranteed data residency in a specific jurisdiction
Sovereignty requirements, data localisation laws
Private RAG
Air-gapped or offline operation is required
Military, critical infrastructure, classified environments
Private RAG
You need to deploy in weeks, not months
Speed to market is the primary constraint
Public API
Your data is non-sensitive and publicly available
Product documentation, public knowledge bases, general assistance
Public API
You process fewer than 100K queries per day
Moderate usage where per-token costs are manageable
Public API
You need frontier model quality (GPT-4o, Claude 3.5)
Tasks requiring highest reasoning, coding, or analysis quality
Public API
Moderate sensitivity data, EU users, cloud-hosted
Some PII in context but not regulated healthcare/financial data
Managed Cloud AI
SOC 2 compliance with moderate data sensitivity
B2B SaaS with standard enterprise data
Either with controls

Hybrid Architectures — The Practical Middle Ground

Most mature enterprise AI deployments are neither purely private RAG nor purely public API — they're layered architectures that apply the right model to the right data class.

Tiered Data Sensitivity Architecture

Route queries based on the sensitivity of the data they'll touch:

Public Model + Private Index

Use a frontier model (GPT-4o, Claude) via managed cloud API for inference quality, but keep your document index and embedding pipeline entirely private. Your documents never leave your environment — only the query and retrieved chunk summaries (stripped of raw sensitive content) travel to the model. This gives you 80% of the privacy benefit of private RAG at 20% of the infrastructure cost.

Private Embedding + Public Generation

Run your own embedding model (all-MiniLM, BGE, E5) on-premise for document indexing and query encoding, then send only the top-k retrieved chunks — transformed into structured, sanitised summaries — to a public model for generation. The model never sees your raw documents.

Securing Whichever You Choose

Security controls differ significantly between deployment models. Here's what each requires:

If You Choose Private RAG

If You Choose Public AI API

The security testing requirement applies to both: Whether you run private RAG or use a public API, you need continuous adversarial testing to validate your controls hold. Self-reported security posture — "we have guardrails" — is not the same as tested resilience. The HexTyx scanner tests both deployment models with the same 21-category adversarial suite.

️ Know Your RAG Security Score Before You Deploy

Get a structured security assessment across RAG pipeline security, prompt injection resilience, agent security, and compliance readiness. Free, 10 minutes, PDF report included.

Frequently Asked Questions

What is the difference between private RAG and public AI?
Private RAG keeps all data — documents, queries, responses — within your infrastructure using a private vector database and either a self-hosted or private-tenant LLM. Public AI sends your prompts and retrieved content to a third-party API for processing. The core difference is data custody: private RAG guarantees your data never leaves your environment; public AI sends it to the vendor's servers.
Is private RAG more secure than public AI APIs?
Not inherently. Private RAG eliminates data-in-transit exposure to third parties but introduces unique risks: RAG poisoning, indirect prompt injection through internal documents, vector database compromise, and retrieval permission bypass. Public AI eliminates the infrastructure attack surface but exposes data to the vendor's security posture. Both require different but equally rigorous security controls — neither is automatically safer.
Can HIPAA-regulated organisations use public AI APIs?
Yes, with conditions: a signed Business Associate Agreement (BAA) with the AI vendor, strict de-identification of PHI before it enters any prompt, full audit logging of AI interactions, and a vendor security assessment. OpenAI offers BAAs for ChatGPT Enterprise; Anthropic offers them for Claude Enterprise; Azure OpenAI Service includes BAA coverage under the standard Microsoft Enterprise Agreement. Many healthcare organisations choose private RAG to eliminate BAA complexity and maintain complete PHI control.
What are the main security risks specific to private RAG?
Five risks are unique to private RAG: (1) RAG poisoning — adversarial content in indexed documents; (2) indirect prompt injection — malicious instructions in internal documents executed during retrieval; (3) vector database compromise — attacks on the embedding store; (4) retrieval permission bypass — queries extracting documents beyond the user's authorisation scope; (5) embedding model attacks — adversarial inputs manipulating similarity scores. All five require dedicated controls beyond standard application security.
When does private RAG become cost-effective vs public AI APIs?
The infrastructure break-even (excluding engineering overhead) is typically 500M–1B tokens/month — roughly 500,000–1,000,000 daily queries. Including engineering maintenance (0.5 FTE), break-even moves to 1B+ tokens/month. For most organisations below this scale, public AI APIs or managed cloud AI (Azure OpenAI, AWS Bedrock) are significantly cheaper. Private RAG is driven by compliance and data sensitivity requirements more often than pure cost optimisation.
What is the best hybrid architecture for regulated enterprises?
The most practical hybrid: managed cloud AI (Azure OpenAI or AWS Bedrock) for inference — data stays within your cloud tenant, not on shared vendor infrastructure — combined with a private vector database (self-hosted Weaviate, pgvector, or a managed private instance) for your document index. This gives you data residency control over your knowledge base, frontier model quality for generation, and compliance with GDPR, HIPAA (with BAA), and SOC 2 — at significantly lower infrastructure cost than fully self-hosted RAG.
Does the EU AI Act require private RAG?
No. The EU AI Act does not mandate a specific deployment model. It mandates controls: technical documentation, human oversight, audit logging, and adversarial robustness testing for high-risk AI systems — regardless of whether you use private RAG or a public API. However, GDPR Article 46 transfer requirements for personal data processing outside the EU make private RAG significantly simpler to manage for European personal data, since it eliminates the international transfer question entirely.

Related Resources