You’ve run the scan. SAST came back clean, SCA found no known-vulnerable components, and the config audit didn’t flag anything. So the MCP server you just wired into your agent must be fine, right?
Conventional AppSec tooling inspects static, inspectable artefacts: source code, dependency manifests, configuration. It hunts for memory-safety flaws, injection and known-vulnerable components.
MCP adds a different attack surface: the vulnerability lives in runtime tool metadata that steers the model’s reasoning, not in executable logic. A traditional vulnerability is a defect in code or a dependency. An MCP-specific vulnerability is malicious instruction embedded in a tool’s description. The first attacks memory safety and control flow; the second attacks the model’s judgement inside its context window. Nothing executable changes, so a scanner that only reads code has nothing to detect.
We’ll walk four attack classes conventional scanners miss: tool poisoning, rug pulls, cross-server tool shadowing and state handle hijacking. Up front, one distinction: prompt injection is the delivery mechanism; the vulnerability itself is the unverified tool metadata it carries, per Check Point’s Black Hat USA 2026 briefing. This is one piece of the full picture of MCP security.
How do MCP-specific vulnerabilities like tool poisoning differ from traditional software vulnerabilities?
Traditional vulnerabilities live in static artefacts a SAST or SCA scanner can parse: code, configuration, dependencies. MCP tool poisoning is the opposite. It embeds malicious instructions in a tool’s natural-language description; no memory-safety or control-flow defect exists. The attack sits in the context window as trusted text. In the first large-scale study, a send_reply tool’s description told the model to change the recipient to an attacker’s number; the same study found eight MCP vulnerabilities, only three overlapping with traditional software vulnerabilities.
Tool poisoning is prompt injection delivered at the integration layer, before any conversation begins.
Where do MCP-native scanners and taxonomies live? The OWASP MCP Top 10 and MITRE ATLAS cover the taxonomy; Invariant Labs’ mcp-scan does the scanning, and the MCPTox benchmark evaluates coverage. CWE and CVE still lack MCP-specific categories. Those gaps are one strand of the wider MCP security landscape.
The blind spot cuts both ways. MCP-native scanners miss conventional memory-safety and dependency flaws; general-purpose static analysis can’t see poisoned tool semantics. A hybrid pipeline, SonarQube alongside mcp-scan, covers both, and it’s what to demand when assessing an MCP vendor.
What is a rug pull attack, and why can’t MCP detect a redefined tool?
A rug pull is trust established at approval time and then revoked: a tool is approved while benign, then redefined afterwards. MCP reflects tool descriptions dynamically with no integrity check or re-verification step, so the protocol can’t notice the tool it approved is no longer the tool it runs.
The name and schema stay the same; the description or handler changes. MCP picks the change up on the fly, and without a re-approval trigger the poisoned version goes live with no extra review. Your approval was a point-in-time check against a definition that drifts later, which is what rug pulls exploit.
Microsoft’s example is an invoice-enrichment service hiding an order to grab unpaid invoices; in the wild, postmark-mcp added a BCC field to send_email, copying outgoing mail to an attacker.
The protocol can’t detect it because reflection shows the current definition but offers no integrity baseline to compare against. The fix is yours: pin and sign tool definitions, fingerprint the canonical definition, and re-trigger authorisation on drift. The Enhanced Tool Definition Interface formalises this but isn’t in the core spec. Approval only helps with that monitoring, which is why redefinition defeats the consent model the protocol assumed.
What is cross-server tool shadowing, and how does one compromised MCP server escalate to trusted ones?
Cross-server tool shadowing happens because an MCP client aggregates tools from every server into one flat list, so a compromised server can register a tool whose name collides with, or mimics, a trusted server’s tool. Calls meant for the legitimate tool get intercepted or altered, and one bad server escalates instead of staying contained.
Multiple servers can offer the same tool name, and the client might round-robin them or pick the alphabetically first. An attacker can also hide behind invisible characters in the description. A malicious tool description in one server shapes how the agent builds parameters for a separate, legitimate tool, without ever invoking it directly.
Invariant Labs showed this in April 2025: a malicious trivia-game server targeted a legitimate WhatsApp server in the same session, extracting message history through it (source). The compromised server wasn’t isolated; it became a launch point into everything else.
That escalation spills into credentials. A shadowing tool that intercepts a call also receives the tokens or consent it carried, a confused-deputy failure and why token passthrough is forbidden in the MCP authorisation spec. We cover the consequences in the confused deputies article.
What is state handle hijacking, and why does MCP’s stateless design enable it?
State handle hijacking is the replay of a server-issued identifier, a cart ID or workflow ID, by someone who never minted it. MCP is stateless, so every request is independent and no protocol-level session boundary binds the handle to its owner.
A server that needs state mints a handle and returns it to the agent, and the agent echoes it back as an ordinary tool argument on each call. The spec is explicit: possession of a handle is not authentication. Anyone who obtains or guesses it can replay it to read or modify another party’s state.
MCP removed the session concept. The 2025-11-25 versions used server-assigned session IDs, where an attacker who guessed one could impersonate the original client. The newer stateless spec closes that hole but shifts trust onto the handle.
Containment is server-side: bind every handle to the authenticated principal that minted it, use non-deterministic identifiers with expiry, and verify each inbound request. That’s the boundary discipline the protocol’s session and consent model assumed, and once a handle is replayed, gateway containment stops the damage spreading.
The four classes are one pattern: MCP trusts runtime metadata and treats trust decisions as static, showing up as poisoning, post-approval drift, cross-server escalation and stateless replay.
What changes in practice? Demand signed and pinned tool definitions with drift monitoring, and treat the aggregated tool list and stateless handles as attack surface to be contained and bound server-side. That discipline — from adopting a server to containing its blast radius — is what securing the Model Context Protocol covers.
The scanner stayed silent because it was pointed at the wrong layer. Before asking “is this server’s code clean?”, ask what its tools can do to your agent’s judgement, and to every other server it trusts, after you’ve approved it.
Frequently Asked Questions
Is prompt injection the same thing as tool poisoning?
Not quite. Prompt injection is the delivery mechanism, while tool poisoning is the vulnerability it exploits. Tool poisoning plants malicious instructions in a tool’s description or metadata, and prompt injection is how that text reaches the model’s context window as trusted input. Treating them as the same thing hides the real fix: the problem is unverified tool metadata, not the injection channel itself.
Is a rug pull attack just another name for tool poisoning, or is it a separate attack class?
A rug pull is a sub-technique of tool poisoning, not a separate class. The distinction is timing. Standard tool poisoning ships malicious metadata from the first install, while a rug pull is approved in a benign state and then silently redefined afterwards. That post-approval mutation is what defeats point-in-time review, because the agent keeps calling a tool it already trusted.
Does human-in-the-loop approval stop these attacks?
No, and this is the trap. Human-in-the-loop approval is a point-in-time check, and a rug pull mutates the tool after that check has passed. The reviewer approved a benign definition, then the server changed it. Approval only helps if it is paired with tool-definition hashing, pinning and drift monitoring that re-trigger authorisation whenever a definition changes.
What does tool description drift actually look like, and how do I monitor for it?
Description drift is any change to a tool’s natural-language metadata after it was approved: altered instructions, a new recipient, or shifted parameters while the tool name and schema stay the same. Monitoring means fingerprinting every approved definition, storing the hash, and alerting on any mismatch between the stored baseline and the definition the server currently reflects at runtime.
Can a compromised MCP server do damage if my agent never calls its tools?
Yes, if its tools are exposed in the aggregated list. Shadowing does not need the attacker’s tool to be called: colliding with or mimicking a trusted tool is enough to intercept calls meant for that tool. The mere presence of a hostile definition in the shared tool list can redirect traffic, so exposure, not invocation, is the risk.
What happens if a shadowing tool receives credentials meant for a trusted server?
It can capture or replay them. When a shadowing tool intercepts a call intended for a trusted server, it also receives the tokens, identifiers or consent that call carried. That is a confused-deputy failure: a request or credential is trusted by a downstream service because it appears to come from a legitimate tool, which is exactly the credential exposure we cover in the confused deputies article.
How is state handle hijacking different from session hijacking?
They exploit the same trust gap at different layers. Session hijacking relied on guessing server-assigned session IDs in the 2025-11-25 protocol versions, where a session concept still existed. State handle hijacking is the stateless equivalent: MCP now has no session at all, so a cart ID or workflow ID becomes an ordinary tool argument that any party can replay without a protocol boundary binding it to its owner.
Why are MCP-specific vulnerabilities still missing from CVE and CWE?
Because both taxonomies catalogue defects in discrete software artefacts, and MCP attacks are semantic rather than syntactic. There is no vulnerable function or dependency to assign a CVE to when the flaw is a tool description that steers the model. Until MCP-specific categories exist, OWASP MCP Top 10 and MITRE ATLAS are the closest authoritative references.
Where can I find MCP-specific vulnerability scanners and threat taxonomies?
Start with the OWASP MCP Top 10 and MITRE ATLAS for taxonomy, then add a purpose-built scanner such as Invariant Labs’ mcp-scan for runtime tool-metadata checks. The MCPTox benchmark is useful for evaluating coverage. Note that CWE and CVE still lack MCP-specific categories, so these MCP-native resources are your practical baseline today.
MCP-native scanning versus general-purpose static analysis: which class of vulnerability does each miss?
Each covers what the other cannot. MCP-native scanners catch poisoned tool semantics but miss conventional memory-safety flaws and known-vulnerable dependencies. General-purpose SAST and SCA tools read code and manifests but are blind to runtime tool metadata. A hybrid pipeline, for example SonarQube alongside mcp-scan, is what covers both classes.
Do I need to worry about these attacks if I only use MCP servers from reputable vendors?
Yes. Reputation is not an integrity guarantee, because rug pulls and shadowing happen after the vendor relationship is established. A trusted server can be compromised, silently redefined, or shadowed by another server in the same aggregated list. Provenance evidence and continuous re-verification matter more than brand, which is what to demand from a vendor during assessment.
How do I contain the blast radius if a state handle is replayed?
Containment is server-side, not protocol-side. Bind every handle to the authenticated principal that minted it, use non-deterministic identifiers with expiry, and verify each inbound request rather than trusting the handle as proof of access. Once a handle is replayed, gateway-level controls that limit what that session can reach are what stop the damage spreading.