Attackers don't need to breach your infrastructure to steal your model — they just need API access. How model extraction works, why traditional security tools are blind to it, and the defenses that actually make it expensive.
For most AI-powered products, the model itself is the most valuable intellectual property the company owns — months of R&D, proprietary training data, fine-tuning expertise, and domain-specific knowledge all compressed into a set of weights. The uncomfortable reality is that an attacker doesn't need to breach any infrastructure to steal that asset. They just need API access. Model theft — cataloged in MITRE ATLAS under the Exfiltration tactic as model extraction — is an inference-based attack: the adversary queries a public API thousands of times, collects the responses, and trains a functionally equivalent replica. No firewall logs it as exfiltration. No DLP tool flags it. The model leaves one prediction at a time.
MITRE ATLAS identifies three distinct sub-techniques under exfiltration via inference API. Inferring training data membership determines whether a specific data point was present in the original training set. Model inversion reconstructs portions of the underlying training data directly from model outputs. Model extraction — the focus here — steals the model itself by systematically querying the API and using the responses to train a replica, using the production API purely as a labeling oracle without ever touching the actual weights, infrastructure, or training data directly.
Reconnaissance starts with probing the target API using benign inputs to understand what it accepts, what format its outputs take, what rate limits and authentication exist, and critically, whether it returns full probability distributions or just a single top-ranked prediction — an API that returns softmax probabilities across every possible class leaks far more extractable signal than one returning only a single label. Resource development follows, with the attacker preparing automation scripts, allocating compute, and typically choosing a smaller, efficient target architecture to train as the surrogate. Access frequently requires nothing more than legitimate use of the service — signing up, obtaining a normal API key, and staying technically within rate limits, since model theft often involves no unauthorized access in the traditional sense at all. The actual extraction then proceeds by systematically querying across the input space, collecting full probability outputs where available, and using those as soft labels to train a surrogate via knowledge distillation; for language models specifically, this looks like sending prompts designed to elicit characteristic behavior and fine-tuning an open-source base model on the collected response set.
A stolen model can simply be the product — sold or deployed directly without the R&D investment the original required. It can serve as a local white-box copy for developing more precise adversarial attacks against the production API it was extracted from. It enables direct cost arbitrage, offering equivalent inference at a lower price point by undercutting the original's pricing with none of the development cost behind it. It hands a competitor genuine competitive intelligence about which features the model actually prioritizes in its decisions. And for AI-powered security products specifically — fraud detection, content moderation — an extracted local copy becomes the ideal testbed for developing evasion techniques against the live production system.
Firewalls and WAFs see traffic that looks like completely legitimate API usage — no malicious payloads, no injection patterns, just queries and responses shaped exactly like normal customer traffic. DLP tools have nothing to flag because no files are being exfiltrated; the model leaks through inference results, not a file transfer. Standard SIEM rules miss it unless AI-specific behavioral baselines have been deliberately built, since query volume anomalies otherwise blend into normal traffic patterns. Basic per-minute rate limiting fails against a patient attacker who simply spreads the same query volume across weeks instead of hours. And authentication provides no defense at all when the attacker has entirely valid credentials — they're a paying customer, or a compromised but legitimate account, not an intruder in any conventional sense. Model theft is fundamentally an abuse-of-function attack rather than a breach: the attacker uses the API exactly as designed, just for an unintended purpose at scale.
A SaaS company offering AI-powered document classification trained on a large proprietary legal corpus can have a competitor sign up for the service, collect tens of thousands of query-response pairs over several months, and train a local replica — then launch a directly competing product at a fraction of the price, having paid nothing for either the training data or the underlying model development. A financial institution running an AI fraud detection model can have that model extracted and probed until an attacker discovers it flags transactions above a specific dollar threshold combined with certain merchant codes — at which point fraudulent transactions get deliberately structured to stay just under that threshold, materially increasing the success rate of fraud the bank's own model was supposed to catch.
Limiting how many queries any single user, IP, or API key can make directly increases the cost and time required for extraction. Implementing daily and monthly caps, not just per-minute rate limits, closes the "patient attacker spreading queries over weeks" gap that basic rate limiting misses entirely. Progressive throttling that slows responses after a threshold is crossed, combined with requiring business justification for any elevated quota request, makes large-scale extraction meaningfully impractical rather than merely inconvenient — since extraction genuinely requires thousands to millions of queries, even a modest daily cap changes the economics significantly.
Enforce strict authentication on every inference endpoint — no public, unauthenticated access to anything that returns model predictions. Apply role-based access control with least privilege as the default. Require identity verification specifically for accounts requesting high query volume. And segment internal versus external API access with separate rate limits and separate monitoring, since the threat model for each is genuinely different. If an attacker can't reach the API at all, extraction is impossible by definition; if they can only reach a limited subset of functionality, any extraction that does occur stays meaningfully incomplete.
Log every query, response metadata, user identity, timestamp, and source IP as a baseline. Monitor specifically for systematic patterns — even coverage across the input space, small deliberate input perturbations, or unusually diverse input types from a single account. Flag accounts whose query diversity is unusually broad, since legitimate users typically query within a fairly narrow domain relevant to their actual use case. Extraction attacks have a distinct behavioral signature once you're looking for it: a single account submitting ten thousand carefully varied images over a single weekend is not how a normal customer uses an image classification API.
Output perturbation adds calibrated noise to confidence scores and probability distributions, degrading the signal available for distillation while preserving enough utility for legitimate users — the attacker gets less information per query, which means exponentially more queries are needed for comparable extraction quality. Prediction truncation returns only top-k labels instead of full probability distributions, or truncates and summarizes long language model outputs, reducing information per query the same way. Behavioral monitoring watches for the characteristic signature of extraction specifically: systematic coverage across the input space, unusually high diversity from a single account, bulk activity concentrated in off-hours, and queries with no conversational continuity between them. Model watermarking embeds detectable patterns in the model's own behavior that survive extraction — this doesn't prevent theft, but if a stolen model later surfaces on a public model hub, in a competitor's product, or in an underground forum, the watermark provides forensic evidence supporting legal action.
For organizations running both frameworks side by side, model theft maps directly onto OWASP's unbounded consumption risk category — OWASP frames the same underlying behavior as a resource exhaustion and cost-control problem, but the attack vector is identical: an adversary abusing an inference API at scale. The MITRE ATLAS framing helps a SOC build detection rules and threat intelligence; the OWASP framing helps developers implement the actual rate limits and access controls at design time. Used together, they cover both the detection and the prevention side of the same underlying risk.