Plugging in an MCP server means handing a non-deterministic model tools that reach into production systems and asking it to act on your behalf. Before granting access, ask what evidence you have actually seen.
Vendor confidence outruns independent verification here. Independent scans find only about 8.5% of public MCP servers use OAuth, while most lean on static keys or non-expiring tokens. That gap between what the spec promises and what a server actually does is the MCP Authentication Gap. Assume a server does not meet the authorisation model until the publisher proves otherwise.
This article works the assessment in three tiers: pre-adoption posture, the evidence bar for regulated data, and a read on the ecosystem. Each feeds an MCP governance programme that sets your acceptance thresholds, not the vendor’s sales deck. For the wider picture, see the MCP security landscape.
What should I look for when evaluating an MCP server’s security posture before adopting it?
Check four signals before adopting any server: provenance, authentication posture, maintenance, and vulnerability history. Who published it and where, whether it implements OAuth 2.1 or falls back to static keys or non-expiring tokens, how actively it is maintained, and what its CVE and NVD record says.
Start with provenance. An MCP Registry listing tells you the server exists; it is not a security endorsement. Registries exist to solve tool discovery and do not host, authorise, or execute calls; many apply only basic review. Publisher identity, ownership history, and release track record are the opening signals. Signed releases and verified publishers should be required.
Next, authentication posture: does the server implement spec-mandated OAuth 2.1, or fall back to static keys or non-expiring tokens? One scan of 5,205 open-source servers found 53% rely on static keys or non-expiring tokens, the MCP Authentication Gap in practice.
Then maintenance. A dormant repository is itself a risk signal, so look at release cadence, issue responsiveness, and how quickly fixes land.
Then vulnerability history. Check CVE and NVD records before and after adoption. Researchers filed more than 30 CVEs in 60 days in early 2026, one scoring 9.6 and downloaded 437,000 times. Conventional scanners miss tool poisoning, so add MCP-specific dynamic scanning for it (more on this below). Fold your findings into the Build vs Buy Evaluation and let the governance programme set the thresholds.
What evidence should a vendor provide before an agent touches production financial or regulated data?
Demand audience-bound OAuth 2.1, scope minimisation, audit-trail guarantees, incident history, and independent verification. Demand proof of what the agent cannot do, and a commitment to what the vendor will show you after an incident.
Under Zero Trust for Agentic AI, no AI agent should be trusted by default. An agent is a non-human identity acting at machine speed, so deny access until evidence is on the table.
Authentication is the floor. Audience-bound OAuth 2.1 means short-lived, scoped tokens with PKCE, a clean split between authorisation server and resource server, and no token passthrough. Passthrough creates confused deputies that act with authority they were never granted, which is why the spec forbids it.
When agents act across many MCP servers on behalf of your business, Enterprise-Managed Authorization puts your identity provider in the decision seat, governing cross-app access centrally with no per-user consent screens. OAuth 2.1 sets the floor for single-server, per-user access; reach for Enterprise-Managed Authorization (EMA) as your footprint widens.
Then demand scope minimisation matched to the task, short token lifetimes, and separate rules for AI versus human access. Demand immutable MCP-layer logging of every tool call: agent identity, user identity, server, tool, sanitised arguments, approval status, and outcome. And demand incident history, third-party audits, penetration tests, and SBOM attestations. Source these demands from your governance programme, not per vendor, so the bar stays consistent.
The same scrutiny scales one level further when you stop asking about one server or one vendor and ask about the whole ecosystem.
Is MCP’s security posture actually worse than PyPI and NPM, or just newer?
MCP’s posture is newer, and the difference comes down to maturity. PyPI and NPM earned their controls over a decade of publicised incidents, and MCP is being asked to compress that curve into months.
Those ecosystems reached registry signing, maintainer verification, and malware scanning only after well-publicised incidents, from typosquatting through to the Shai-Hulud worm that compromised over 500 npm packages. Because MCP rides the same packaging and registry machinery, SBOM and AIBOM generation, SLSA provenance, code signing, and secret scanning all carry over.
What MCP lacks is mature ecosystem-scale review and revocation. And it adds something new: reflective tool descriptions are a model-trusted input channel that PyPI and NPM never had to defend. Tool poisoning hides instructions in metadata the model reads but you never see, so it stays invisible to the scanners those ecosystems rely on.
So treat it as a maturity gap to manage. CVE and NVD coverage is thinner for MCP-specific threats, but that is a measurement problem rather than proof of safety. MCP starts behind on maturity and ahead on novel risk. The answer: it sits earlier on the same maturity curve — a position best read against the wider MCP security picture.
None of this resolves into a verdict. The correct response to a maturity deficit is to make the evidence bar rise with the stakes. Your pre-adoption posture signals are the floor, the regulated-data evidence bar is the ceiling for high-stakes access, and the PyPI and NPM comparison tells you the gap is manageable rather than disqualifying.
Grant agent access to financial or regulated data on evidence, never on vendor confidence. At bottom, this is a trust decision, and a calibrated Build vs Buy Evaluation sourced from governance is what turns “newer but riskier” into a defensible adoption posture. That posture comes from assessing MCP risk in context, not from a verdict on the protocol.
Frequently Asked Questions
What is the MCP Authentication Gap, and why should I assume it exists by default?
View the MCP Authentication Gap as a design assumption, not a defect to excuse. The protocol’s authorisation model presumes human consent, yet many servers fall back to static API keys or skip audience-bound tokens, so the spec’s guarantees cannot be taken on faith. Assume a server does not meet the model until the publisher demonstrates otherwise, and treat that proof as the entry condition for your assessment.
Does listing on an MCP Registry mean a server has been security-reviewed?
No. A registry listing is discovery, not endorsement. MCP Registries generally apply limited review, and publication tells you only that a server exists and who claims to maintain it. Treat the listing as the starting point for provenance checks, then verify publisher identity, ownership history, and release track record yourself before granting any agent access.
What is tool poisoning, and why do conventional scanners miss it?
Tool poisoning hides malicious instructions inside tool descriptions and metadata, which a model then reads and trusts as legitimate. Conventional SAST and SCA tools inspect code and dependencies, not the reflective metadata that steers agent behaviour, so these attacks slip straight through. That is why you need MCP-specific dynamic scanning of tool definitions before and after adoption, not just a conventional pipeline.
Should I use OAuth 2.1 or extension flows like Enterprise-Managed Authorization for cross-app AI access?
Use OAuth 2.1 for single-server, per-user access, where short-lived, scoped tokens and consent screens fit naturally. Reach for Enterprise-Managed Authorization when agents act across many MCP servers on behalf of an organisation, letting an identity provider govern cross-app access centrally without per-user consent friction. The two are complementary: OAuth 2.1 sets the floor, and EMA extends control as your agent footprint widens.
What happens if a vendor cannot provide an audit trail or incident history?
Deny access. Immutable MCP-layer logging of every tool call is the non-negotiable that makes an incident investigable, capturing agent identity, user identity, server, tool, sanitised arguments, approval status, and outcome. A vendor that cannot disclose incident history or show independent verification has not earned regulated-data access. Escalate the gap through your governance programme rather than accepting a verbal assurance.
Do I need OAuth 2.1 for every MCP server, even read-only internal ones?
Apply the standard in proportion to the stakes, but never waive it silently. A read-only internal server touching non-sensitive data may tolerate lighter controls, yet it still needs an identified publisher, least privilege, and a logging story. The moment a server can reach financial or regulated data, audience-bound OAuth 2.1 with short-lived scoped tokens becomes the floor. Let governance, not convenience, set that line.
What is token passthrough, and why is it a problem?
Token passthrough is when an MCP server accepts a token from one resource and forwards it to another without validating that the token was issued for it. That creates confused-deputy conditions, letting a server act with authority it was never granted. Demand audience-bound tokens and a clean separation between authorisation server and resource server, so a token for one service cannot be replayed against another.
How often should I re-assess an MCP server after it is in production?
Treat the first assessment as a baseline, not a one-off. Re-check on every material release, when ownership or registry listing changes, and on a fixed cadence tied to your governance programme. Add continuous tool-definition scanning, because rug pulls and tool poisoning can appear after adoption with no version bump. A server that passed last quarter may not pass this one.
What are SBOM, AIBOM, and SLSA provenance, and why do they matter here?
An SBOM lists the software components inside a server, an AIBOM extends that inventory to models and AI assets, and SLSA provenance records how and from what the artifact was built. Together they let you verify a server’s supply chain rather than trust a summary. Ask for them as procurement evidence, then check that the components and build steps match what you actually deployed.
Can I reduce an MCP server’s blast radius without abandoning it?
Often, yes. A gateway layer can sit between agents and servers to enforce least privilege, inspect tool calls, and constrain what a compromised server can reach. That containment approach lets you keep a useful server in place while capping the damage if it is subverted. Blast-radius control complements your evidence bar; it does not replace the authentication and logging guarantees you still need.
If MCP-specific threats are under-reported in CVE and NVD, how do I measure risk?
This is a measurement problem, not proof of safety. CVE and NVD coverage remains thinner for MCP-specific threats, so absence of records tells you little. Broaden your inputs: MCP-specific dynamic scanning, publisher disclosure history, independent audits, and telemetry from your own audit trail. Combine those sources and revisit them regularly, because the ecosystem’s visibility is still maturing quickly.
Do I need a full governance programme if I am only piloting one server?
Start the discipline now, even if the programme is thin. A pilot is exactly how ad-hoc judgement becomes permanent, so define acceptance thresholds, logging requirements, and re-assessment triggers before the server touches anything sensitive. The programme does not need to be heavy to be useful: it needs to be the source of your evidence demands, so the bar stays consistent as you scale beyond one server.