ThePromptBuddy logoThePromptBuddy
All Insights
AnthropicGoogle

Prompt Injection Is Worse Than SQL Injection

Siddhi Thoke
 Editorial cartoon split scene: a parameterized-query shield blocks a small SQLi villain on the left while a many-armed Prompt Injection creature strolls past four flimsy paper fences labeled Filter, Classifier, System Prompt, Guardrail on the right.

Prompt injection is the supply-chain failure of the LLM era. It sits at the top of the OWASP Top 10 for LLM Applications (LLM01:2025), and the UK National Cyber Security Centre put a finer point on it in December 2025: "Prompt injection is not SQL injection. It may be worse."

The thesis here is direct. The SQL injection analogy is useful for naming the problem and dangerous for bounding it. SQLi has a deterministic fix (parameterized queries) that, when used correctly, is 100% effective. Prompt injection has no architectural fix, only layered mitigations, each with non-trivial residual failure rates. If your mental model is "we just need to sanitize inputs," you have already lost.

This piece argues the case in five moves: how the analogy got coined, what the production incident catalogue now looks like, why parameterized queries cannot exist for LLMs, what actually reduces risk in numbers, and what to ship this quarter.

The analogy that almost works

Simon Willison coined "prompt injection" in September 2022, drawing the explicit parallel to SQLi. The structural similarity is real. Both are input-trust failures. Both cross a boundary between code and data. Both took the security industry years to name and even longer to take seriously. Both have followed the same lifecycle from research curiosity to weaponized exploit with assigned CVEs.

That is where the analogy stops being useful.

In 2023, Willison was already worried about the gap. "When parameterized SQL queries are used correctly, systems are 100% protected against SQL injection attacks. It's not unreasonable to want a security fix that, when applied correctly, works 100% of the time." Two and a half years later, no equivalent for prompt injection exists.

The NCSC's December 2025 post puts the structural reason plainly: "Under the hood of an LLM, there's no distinction made between 'data' or 'instructions'. There is only ever 'next token'." An LLM does not parse its input. It predicts. Every byte that lands in context, whether from the system prompt, the user, a retrieved document, a tool result, or an image, gets averaged into the same probability distribution.

Five incidents that prove this is production-grade now

The 2024 to 2026 incident catalogue makes the abstraction concrete. Five examples, picked because each broke a different defense.

Slack AI (August 2024). PromptArmor showed that Slack AI would surface API keys from private channels when fed instructions hidden in any public message. Slack patched after disclosure. The exploit needed nothing exotic: a public message that the AI summarizer ingested.

spAIware (September 2024). A persistent injection planted instructions in ChatGPT's long-term memory, surviving across sessions and silently exfiltrating data on subsequent conversations. The first credible "prompt injection as malware" demo.

EchoLeak, CVE-2025-32711 (June 2025). The first known zero-click prompt-injection-to-exfiltration in production. CVSS 9.3. Aim Security's exploit chained reference-style markdown with a Teams proxy on the Microsoft 365 Copilot CSP allowlist, bypassing Microsoft's own XPIA prompt-injection classifier. The attacker sent an email. That was the entire user interaction.

ShadowLeak (September 2025). Radware's disclosure against ChatGPT Deep Research integrating with Gmail. The first server-side leak: data exfiltrated from OpenAI infrastructure, not the user's client. OpenAI patched in early August. A January 2026 variant called ZombieAgent bypassed the fix.

LangGrinch, CVE-2025-68664 (December 2025). Critical vulnerability in LangChain Core, CVSS 9.3, allowing secrets exposure through a serialization-flavored injection. Affects every team that imports LangChain into a production agent stack. Most have not patched.

Five vectors, five defenses bypassed: hidden in retrieved documents, persistent memory, allowlisted tools, server-side agent integrations, framework-level deserialization. None used "ignore previous instructions." The 2022 demo is now the least sophisticated thing in the field.

If the kind of agent that finds these bugs is interesting, our piece on Anthropic's Claude Mythos Preview as a zero-day machine is the adjacent read. The same capability that lets an AI agent find vulnerabilities at scale is what lets adversaries weaponize them.

Vertical timeline of 12 prompt injection incidents from February 2023 through December 2025 with CVE-tagged events highlighted, illustrating the shift from research demos to production exploits.

Why isn't prompt injection just SQL injection with extra steps?

Because there is no parameterized-query equivalent. SQLi is fixed by an architectural separation between query template and parameter. The database engine parses one and binds the other. LLMs have no parser. Every token, regardless of source, is concatenated into a single context window and rolled through the same forward pass.

Three structural facts make this worse than SQLi. First, the data channel is unbounded. An LLM consumes text, retrieved documents, image pixels, audio transcripts, OCR'd PDFs, calendar invites, Slack messages, web pages, tool outputs. Every one of those is a potential injection vector. SQL only cares about the query input.

Second, the model is non-deterministic. The same hostile input can succeed on Monday and fail on Tuesday. Pen-testing for prompt injection is a probabilistic exercise; pen-testing for SQLi is binary.

Third, agent contexts multiply blast radius. Once an LLM has tools, every tool is a potential exfiltration channel. An injected prompt that gets the agent to send an email, write to a database, or commit code is a privilege escalation, not a leak. The Replit AI agent that dropped a customer's production database after ignoring its "do not touch production" guardrail is the canonical agent failure mode, not an outlier.

In April 2026, Google's "AI threats in the wild" telemetry reported a +32% relative increase in malicious indirect-injection content on the public web between November 2025 and February 2026. The supply is not slowing.

The lethal trifecta (and how to break it)

Willison's most useful contribution since coining the term is the "lethal trifecta" framework, published June 2025. Any agent system that combines all three of:

  1. Access to private data
  2. Exposure to untrusted content
  3. Ability to communicate externally

is exploitable. Remove any one leg and the trifecta breaks. The corollary, sometimes called the privilege-drop principle, is that when an LLM processes data from a party, its privileges should drop to that party's level.

Every incident above hits the trifecta. EchoLeak: Copilot had M365 data, ingested attacker email, could resolve external markdown URLs. ShadowLeak: Deep Research had Gmail access, ingested attacker email, could call out via tool. The Aonan Guan GitHub Actions hijacks against Anthropic, Google, and Microsoft? Same shape. (Notably, those bounties paid $100, $500, and undisclosed respectively, and no CVEs were assigned, which means real-world prevalence is materially under-reported.)

The trifecta is the single most actionable mental model in the literature. If you cannot remove a leg, you should not build the agent.

Multimodal makes it worse

If text-only injection sounds bad, the multimodal numbers are worse. A 2025 academic paper measured attack success rates of around 82% for hidden image instructions and 86%+ for audio injection (WhisperInject), with physical-room audio attacks holding 87 to 88% across 5-meter distances. A vision-language agent that scans a screenshot, a contract PDF, or a meeting recording is now in scope. Text-only filters are blind to all of it.

Tools like Cursor 3, repositioned from IDE into an agent control room, are the exact attack surface this matters for: agents reading repo files, viewing screenshots, calling APIs. Each capability widens the trifecta.

Which defenses actually move the numbers, and by how much?

Four defenses have shipped with public effectiveness numbers. The numbers are good. None claim 100%. All four authors publicly acknowledge residual risk is unavoidable.

DefenseMechanismReductionSource
Spotlighting (Microsoft)Encodes untrusted content (delimit, datamark, encode) so the model can distinguish itASR above 50% down to under 2%arXiv 2403.14720
Constitutional Classifiers (Anthropic)Input/output classifier plus RL training to refuse injected instructionsJailbreak rate 86% to 4.4%Anthropic, Feb 2025
Instruction Hierarchy (OpenAI)Trains the model to assign privilege: system, developer, user, tool output+63% robustness on main evalsarXiv 2404.13208
CaMeL (Google DeepMind)System-level: privileged LLM plans, quarantined LLM parses untrusted data and cannot call tools77% task completion with provable security (vs 84% undefended)arXiv 2503.18813

 Horizontal bar chart comparing four prompt injection defenses showing Spotlighting cuts ASR from 50% to under 2%, Constitutional Classifiers cut jailbreaks from 86% to 4.4%, Instruction Hierarchy adds 63% robustness, and CaMeL achieves 77% provable-security task completion.

The pattern is informative. Three of the four are model-level (Microsoft's Spotlighting, Anthropic's classifiers, OpenAI's hierarchy). One is system-level (CaMeL). All three model-level defenses leave a non-trivial residual ASR. The system-level defense is the only one that gets close to provable guarantees, and it does so by enforcing capability limits around the LLM, not on it.

EchoLeak is instructive here too: Microsoft's own XPIA classifier, the production version of Spotlighting-class defenses in M365 Copilot, was bypassed in the disclosed exploit. Vendor classifiers reduce ASR. They do not eliminate it.

What is theatre

Some defenses sound like solutions and aren't. Three to be skeptical of:

Generic "guardrail" SaaS that quotes a 95% block rate without methodology. In an agent context, the residual 5% means catastrophic compromise on average within 20 hostile prompts. Willison and the NCSC have both pushed back on this metric publicly. If a vendor cannot show their attack corpus, treat the number as marketing.

System-prompt hardening alone. "You must never ignore instructions" has been bypassed by every published exploit since 2023. It is necessary for hygiene, not sufficient for defense.

Pure post-hoc output filters. Filters that scan model output for PII or sensitive tokens lose to encoding tricks (base64, character substitution, multilingual rephrasing) that injected instructions trivially produce. They catch sloppy attackers and miss capable ones.

What the evidence supports as actually working is defense in depth: Spotlighting plus classifier plus instruction hierarchy plus capability sandboxing plus human-in-the-loop on sensitive actions plus removing at least one leg of the trifecta. No single layer suffices.

The numbers in the bug-bounty market

If you needed a market signal that prompt injection is real, here it is. HackerOne's 9th Hacker-Powered Security Report, published October 2025, recorded:

  • +540% YoY rise in valid prompt-injection reports
  • +210% YoY in overall AI vulnerability reports
  • +270% YoY in customer programs with AI in scope (now 1,121)
  • $81M in total bug bounties paid in FY2025, up 13% YoY

40% of organizations have already experienced prompt injection, jailbreaks, or guardrail bypasses, while fewer than 50% test continuously. Google's 2025 Vulnerability Reward Program paid $17.1M (+40%) but explicitly excludes direct prompt injection and jailbreaks from scope, which suggests the major vendors still treat this as an alignment problem rather than a security one.

Column chart of HackerOne 2025 AI vulnerability growth showing prompt injection reports up 540 percent year over year, overall AI vulnerabilities up 210 percent, AI-scope customer programs up 270 percent to 1,121, and total bounties paid 81 million dollars up 13 percent.

Compare to 2010s SQLi disclosure curves and the shape is identical. We are roughly two years into the public-incident phase of an attack class that took two decades to discipline.

What should AI builders ship this quarter to reduce risk?

Eight items. Cheap to expensive, ranked by attack-surface reduction per hour of engineering work.

  1. Audit every agent for the lethal trifecta. If all three legs are present, document the business justification. Most agents do not need them.
  2. Apply the privilege-drop principle. When the agent processes content from party X, restrict its tool permissions to what X is authorized to do.
  3. Allowlist outbound destinations. URLs the agent can resolve, domains it can email, packages it can install. Default-deny.
  4. Move sensitive actions behind human-in-the-loop confirmation. Database writes, financial transfers, code merges to main, public posts.
  5. Adopt structured outputs. Force the agent to return JSON conformant to a schema you control, and reject everything else.
  6. Add Spotlighting-style markers around any untrusted content in the prompt. The numbers (from above 50% down to under 2%) are real even at the prompt-engineering layer.
  7. Continuously red-team using ATLAS-mapped techniques. MITRE ATLAS added a Command-and-Control tactic in November 2025 specifically for AI agents. Its corpus is what mature security teams test against.
  8. Treat your framework dependency like any other supply chain. The LangGrinch CVE is a reminder that an injected instruction inside a shared utility hits every downstream agent. Pin versions. Read changelogs.

For teams building agentic dev tools (the Codex on legacy code migrations class of integration, plus terminal-native coding agents generally), the trifecta cleanup is the highest-leverage item this quarter. These tools have all three legs by default.

The strongest counter-arguments, addressed

"Models are getting better at refusing injected instructions." True, and irrelevant. Constitutional Classifiers reduce base jailbreak rate from 86% to 4.4%. That is a 95% reduction and still catastrophic at agent scale. Better is not enough.

"Most apps don't expose tool-calling, so the trifecta does not apply." Most apps in 2024 didn't. Most apps in 2026 do. Tool-calling is now standard in the agent template every framework ships. The default has flipped.

"This is alignment work, not security work." OpenAI, Anthropic, and Google DeepMind currently agree with this framing. The NCSC, OWASP, NIST, and MITRE ATLAS do not. Every credible standards body has moved prompt injection into the security-vulnerability category in 2025. The vendor framing will follow.

The bottom line

Prompt injection is the SQL injection of the agent decade, but worse on three axes (no parser, unbounded data channel, agent blast radius) and unlikely to get an architectural fix. The defenses that ship with numbers are real. None are sufficient alone. The framework that actually maps to defensible engineering is system-level capability limits (CaMeL-style sandboxing) plus the trifecta discipline plus continuous red-teaming.

If you are an AI builder, the work this quarter is not waiting for a parameterized-prompt API. That API is not coming. The work is breaking trifectas, moving sensitive actions behind human gates, and building enough defense layers that the residual ASR is a risk you can underwrite.

FAQ

Is prompt injection actually a security vulnerability or just an alignment problem?

OWASP, NIST, NCSC, and MITRE all classify it as a security vulnerability. Some vendors (Google's VRP) still exclude it from bug-bounty scope, framing it as alignment. The real-world incidents (EchoLeak CVSS 9.3, LangGrinch CVSS 9.3, ShadowLeak server-side leak) make the alignment-only framing harder to defend.

Can I "fix" prompt injection with a better system prompt?

No. Every published exploit since 2023 bypassed system-prompt hardening. It is a hygiene layer, not a defense. Spotlighting markers and classifiers do better but still leave residual attack-success rates above zero.

What is the single most important thing to fix in an agent today?

Audit for the lethal trifecta (private data, untrusted content, external comms). If all three are present, remove a leg or accept the breach risk. Most production agents do not need all three.

Is multimodal prompt injection real or theoretical?

Real. Published 2025 numbers: 82% attack success rate for hidden image instructions, 86%+ for audio (WhisperInject), 87 to 88% for physical-room audio at 5-meter distance. Vision-language agents are exploitable today.

How is this different from SQL injection?

SQL injection has a parameterized-query fix that is 100% effective when applied correctly. Prompt injection has no equivalent because the LLM has no parser separating data from instructions. The NCSC summary: "there is only ever 'next token'."

Does using Anthropic Claude or OpenAI GPT make my app safer than open-weights models?

Marginally. Constitutional Classifiers (Claude) and Instruction Hierarchy (GPT) reduce ASR by 95% and 63% respectively. Both leave residual risk. The system-level defenses (CaMeL-style sandboxing, capability allowlists) outperform any model-level defense and apply equally to open-weights deployments.