All writing
Strategy4 min read

Your AI policy is a PDF. Your agents read tool descriptions.

The diligent move is to vet the tool descriptions before you connect. That is also exactly where the measured attacks land — in the metadata, at registration, before a single call is made.

The policy took four months. Legal, security, the AI working group, a board paper. Forty-odd pages of acceptable use, data classification, human-in-the-loop. It went on the intranet the same week an engineer in that company connected a new server in their IDE. The policy was not in the prompt. A tool description was.

An agent does not read your acceptable-use policy. It reads tool descriptions, and then it acts on them. That sentence is usually where the conversation stops, as if the fix were obvious: write the policy into the tools. Review the servers. Vet the descriptions. Approve the list.

That is what diligence looks like. It is also the trap.

You do not get owned when the tool runs

Tool poisoning, named by Invariant Labs in 2025, hides instructions inside a tool’s description at registration. Those descriptions are appended to the agent’s base prompt automatically. Many hosts never show them to the user. The engineer approved a connector. The agent received a paragraph of English that was never on a slide.

The attack does not wait for the tool to run. It is already in the description you just approved — or in the one that replaced it after you approved it.

MCPTox — the first systematic benchmark, later at AAAI — tested 45 live MCP servers, 353 authentic tools, and 1,312 malicious cases across 20 agents. Average attack success: 36.5 percent.

AgentAttack success rate
o1-mini72.8%
Phi-470.2%
GPT-4o-mini61.8%
Gemini-2.5-flash59.7%
Claude-3.7-Sonnet34.3%
Average attack success rate by agent, selected from the 20 evaluated. MCPTox, arXiv:2508.14925.

Read the ordering, not the magnitudes. The more capable models were the more vulnerable, because the attack exploits the thing they are best at: following instructions. They almost never objected. Claude-3.7-Sonnet’s refusal rate, the highest in the study, was under 3 percent.

This inverts the reflex everyone in that steering committee still has. You do not buy your way out with the frontier tier. Capability is the vulnerability. The model’s own judgement is not a control. Waiting two releases is not a control. Inspecting every tool response is not a control either — that is a different failure, and a later one. MCPTox measures metadata at registration. You were compromised during the handshake, before a single call was made.

The doorway showed up by accident

For two years there was nowhere to stand. Agent integrations were bespoke. Each team wired its own tools. You cannot put a control on a thousand one-off calls you cannot name, which is why governance lived in frameworks — NIST’s four functions, ISO/IEC 42001, the EU AI Act — while the thing that actually acted never crossed a boundary those rules could touch.

Then the Model Context Protocol arrived, sold as plumbing: JSON-RPC, a host, a client per server, tools and resources and prompts. USB-C for agents. True. Mostly uninteresting.

The consequential property is a side effect. Every autonomous action now walks through one door. A door you can name is a door you can put a gateway on: inspect the payload, strip regulated data before it reaches the context window, log which tool was invoked with which parameters, and — this is the part the poisoning numbers make non-optional — treat the description itself as untrusted input. Pin what you approved. Diff it on reconnect. Alert when it changes. A server that was safe at review is free to update itself later.

That door is available only once. Claim it before the agent fleet exists and governance is infrastructure you build one time. Arrive afterwards and you are chasing unofficial servers the way the last decade chased unsanctioned buckets — except this untracked asset has write access to production.

The inventory you cannot produce

This is where the auditor arrives, and the story has not changed. NIST’s Map function asks for a live inventory of AI systems and their data flows — exactly what untracked IDE integrations destroy. ISO/IEC 42001 is certifiable; auditors want logs proving something inspected the traffic, not a PDF asserting that it should. The EU AI Act wants data provenance for high-risk systems and genuine human oversight, and the prohibited-practice tier reaches €35 million or 7 percent of global turnover.

There is a diagnostic you can run this week. List every MCP server reachable from your engineers’ IDEs. If you cannot produce that list today, you will not produce the log a year from now. The auditor grades the log. The attacker does not wait for it.

Stand in the door

The lesson is not a better PDF. It is being in the path before the fleet is, and refusing to treat the description as configuration.

  1. 01Put a gateway in the path before you scale. One control plane doing inspection, redaction, and logging beats five teams each inventing their own least privilege.
  2. 02Treat tool metadata as untrusted input. Pin the descriptions you approved, diff them on reconnect, alert on change. This is the control the poisoning numbers actually ask for.
  3. 03Ban token passthrough. Forwarding a user’s token downstream destroys the audit trail and gives every service the same blast radius. Exchange it for a scoped token at each hop.
  4. 04Fix consent on the server. Approved client IDs, exact redirect URIs, validated state — so a confused deputy cannot borrow your user’s authority.
  5. 05Show the untruncated command. If a human is the last line of defence, do not hide what they are approving behind an ellipsis.

None of this is technically hard. It is organizationally hard: it asks a security function to own a protocol boundary it did not choose and cannot see on any architecture diagram.

MCP did not create this problem. It relocated it somewhere you can finally reach. Governance that lives in a document is a statement of intent. Governance that lives in the protocol is a fact about your system. Only one of those survives an agent with a write tool and a poisoned description.

Sources

  1. 01Chen et al.. MCPTox: A Benchmark for Tool Poisoning Attack on Real-World MCP Servers. arXiv:2508.14925; AAAI 2026, 2025.
  2. 02Luca Beurer-Kellner and Marc Fischer. MCP Security Notification: Tool Poisoning Attacks. Invariant Labs, 2025.
  3. 03Model Context Protocol specification — Authorization. Model Context Protocol.
  4. 04Securing the Model Context Protocol (MCP): Risks, Controls, and Practices. arXiv:2511.20920, 2025.
  5. 05MCP Server Governance: Best Practices for AI Security. Tyk.
  6. 06MCP Permission Models: Designing Secure Interactions. Aembit.
  7. 07Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence. Official Journal of the European Union, 2024.