Prompt injection and agent security

Prompt injection: text the model reads (an email, web page, tool description, document) contains instructions that override or redirect the agent. It is ranked LLM01 in the OWASP Top 10 for LLM Applications (2025 edition). Because agents read untrusted content and act with tools, the impact grows with their permissions.

Frameworks

  • OWASP Top 10 for LLM Applications 2025: Prompt Injection, Sensitive Information Disclosure, Supply Chain, Data and Model Poisoning, Improper Output Handling, Excessive Agency, System Prompt Leakage, Vector and Embedding Weaknesses, Misinformation, Unbounded Consumption.
  • OWASP Top 10 for Agentic Applications 2026, published 2025-12-09 (title and date confirmed again 2026-10-05; the ten item names are in the downloadable PDF, not on the web page).

The “lethal trifecta”

Simon Willison (2025-06-16): an agent that combines access to private data, exposure to untrusted content and the ability to communicate externally can be tricked into exfiltrating data. Remove one leg. His list of past incidents includes Microsoft 365 Copilot, GitHub’s MCP server, GitLab Duo, Slack, Amazon Q and ChatGPT Operator, mostly fixed after disclosure; users who combine their own tools are still exposed.

Real incidents (sourced)

  • EchoLeak, CVE-2025-32711 (Microsoft 365 Copilot): a zero-click attack where one crafted email led to data exfiltration; researchers reported bypassing Microsoft’s XPIA classifier, link redaction and a CSP-allowed Teams proxy. Paper at AAAI Fall Symposium 2025.
  • MCP tool poisoning (Invariant Labs, 2025-04-01): hidden instructions in tool descriptions, “shadowing” of trusted tools and post-approval “rug pulls”.
  • Browser agents: Anthropic (2025-11-24) says prompt injection remains unsolved; Claude Opus 4.5 reached about 1% attack success against its internal adaptive attacker, using RL training, classifiers and red teaming.

2026 updates (sourced)

  • Prompt injection to code execution: Microsoft’s security blog (2026-05-07) reports prompt-injection-to-RCE flaws in Semantic Kernel (CVE-2026-26030, CVE-2026-25592, fixed in Python 1.39.4+ and .NET 1.71.0+) and says that once a model is wired to tools, prompt injection becomes “a code execution primitive”; it advises treating all AI-controlled tool parameters as attacker-controlled.
  • OWASP (via Help Net Security, 2026-06-11): prompt injection maps to six of the ten categories in the Top 10 for Agentic Applications, and the 2026 edition documents CVEs, advisories and breaches rather than theory. It also cites Meta’s “Agents Rule of Two”: an unapproved agent should have at most two of the three trifecta properties, and all three need human oversight.
  • Supply chain: Google’s threat report (2026-05-11) describes compromise of AI gateway packages such as LiteLLM and malicious agent skill packages; see use-of-ai-in-hacking.

Mitigations

  • Least privilege and scoped credentials; human confirmation for sensitive actions (MCP spec: humans SHOULD be able to deny tool calls).
  • Break the trifecta: no outbound channel when reading untrusted content; sandbox.
  • Treat tool descriptions and annotations as untrusted; pin or hash tool definitions; isolate servers from each other.
  • Defence in depth from EchoLeak authors: prompt partitioning, filtering, provenance-based access control, strict CSP.
  • Instructions in CLAUDE.md are guidance rather than enforcement (author’s view); use hooks and permissions (claude-md-and-agent-instructions).

In this vault

onlyclaws (Claw-style assistants run with broad access, so these risks apply directly), ai-browsers, tool-description-design, model-context-protocol, ai-safety-and-governance, use-of-ai-in-hacking, data-privacy-in-ai.

Open items

  • Reports of further 2026 MCP-server CVEs and a first malicious MCP server in the wild came from secondary sources and were removed as unverified.

Sources