Every enterprise that has shipped an AI agent in the last eighteen months has quietly added a new tier to its architecture — and most security teams have not been given the budget, the tooling, or even the vocabulary to defend it. The Model Context Protocol (MCP) is the connective tissue of that tier. It is how an autonomous agent discovers what it can do, reaches the systems it is allowed to touch, and pulls context back into a reasoning loop. MCP is also, by design, an interoperability standard rather than a security standard, and that distinction is now costing organizations real money.
This is not a theoretical concern. The MCP security landscape moved from research curiosity to operational emergency in 2025 and hardened further through 2026. OWASP stood up a dedicated MCP Top 10 — the first OWASP project aimed specifically at the tool-connection layer — and the NSA published MCP: Security Design Considerations for AI-Driven Automation in May 2026. Meanwhile, CVE-2026-33032, an authentication bypass in an MCP-enabled tool rated CVSS 9.8, was actively exploited in the wild. If your agent can reach production, this is your problem now.
Why MCP Changes the Threat Model
Before MCP, an AI application had a small, fixed, developer-authored toolset. A support chatbot could look up an order and send a refund request. The surface area was known, the credentials were scoped at build time, and the code path was reviewed. There was nothing to discover at runtime and nothing new to trust after deployment.
MCP inverts all three assumptions. Under the protocol, an agent discovers tools dynamically from any reachable MCP server, which means the set of things your agent can do is no longer a property of your codebase — it is a property of your network. Every server becomes a trust boundary. Every tool description is content the model will read and act on. And because tool definitions can be updated server-side without the client re-approving anything, the thing you approved on Monday may not be the thing running on Friday.
The practical consequence is a blast-radius problem. As the OWASP agentic guidance puts it plainly, an agent's exposure equals every credential, tool, and API it can reach. Multi-step autonomy compounds the damage across an entire plan rather than confining it to a single flawed response. A compromised agent that can read a database, call a payments API, and send email does not need to be clever — it needs one bad instruction and standing permissions.
The Five MCP Risks That Actually Bite Enterprises
OWASP's MCP Top 10 enumerates ten categories, but in practice five of them account for the incidents that reach a board's attention. These are the ones worth building controls around first.
1. Token mismanagement and secret exposure (MCP01)
This is the unglamorous root cause behind a striking number of MCP incidents. Servers ship with hard-coded credentials, long-lived tokens, and secrets that leak into model memory, protocol logs, or debug traces. Once a secret is in the agent's context, an attacker does not need to breach anything — they need to prompt. A single successful injection that coaxes the model into echoing its environment is enough to hand over the keys to every downstream system that token unlocks. The defense is mundane and therefore frequently skipped: short-lived, narrowly scoped tokens, aggressive secret scanning, and a hard rule that credentials never enter agent context or logs.
2. Tool poisoning and rug pulls (MCP03)
Tool poisoning is the signature MCP attack, because it exploits a genuinely new trust assumption: that a tool's description is metadata rather than executable instruction. It is not. Models read tool descriptions as part of their context, and a malicious server can hide instructions in them that the user never sees in any approval dialog. Invariant Labs demonstrated the pattern in 2025 with a poisoned tool whose description contained invisible instructions; combined with cross-server chaining, an agent can be hijacked without the malicious server ever appearing in the user-facing interaction log.
The rug pull is the same attack delivered on a delay. A tool behaves perfectly at installation, accumulates trust, and then a later server-side update silently changes its behavior or its description to begin harvesting credentials. Most MCP clients do not flag description changes, and few enterprises log them. This is why "we reviewed the server before we connected it" is not a security control — it is a snapshot of a moving target. Signed and pinned tool definitions, plus description-change detection, are the minimum viable defense.
3. The confused deputy and privilege escalation (MCP02, MCP07)
MCP servers typically hold real authority — a service account with broad access — and they act on behalf of a caller that may be impersonated. That is the classic confused-deputy flaw, and it is especially dangerous in agentic architectures because the deputy is holding credentials nobody is watching. When an improperly scoped token lets a compromised server use its higher authority on an attacker's behalf, the attacker never needs their own credentials. Insufficient authentication and authorization compounds this: missing PKCE in OAuth flows, missing audience validation, and scope creep that grants an agent standing access far beyond its task. The MCP specification's 2026 update introduced incremental scope consent for exactly this reason, but implementing it requires an authorization layer that understands tool-level semantics, not just network routing.
4. Command injection and unexpected code execution (MCP05)
Several of the highest-severity MCP CVEs are not AI problems at all — they are ordinary application security failures wearing a new hat. The Filesystem MCP Server shipped a directory containment bypass and a symlink bypass. Anthropic's Git MCP server logged three CVEs in early 2026 covering path traversal and argument injection. CVE-2026-30623 was a command injection in the Anthropic MCP SDK affecting LiteLLM. The lesson is uncomfortable for teams that assumed "AI security" was a separate discipline: your MCP servers are web services and CLIs, and they need the same input validation, sandboxing, and deny-by-default egress you would demand of any internet-facing component.
5. Shadow MCP servers (MCP09)
The fastest-growing risk is the one with the least visible evidence. Individual developers can add an MCP server to an IDE plugin in under a minute, which means ungoverned servers proliferate outside any inventory the security team maintains. Cycode's 2026 research found 81 percent of organizations lack full visibility into how AI is used across the software development lifecycle. A shadow MCP server is worse than shadow IT, because it does not merely store data — it takes actions, holds credentials, and can be steered by an attacker. You cannot secure a fleet you cannot enumerate.
The Four-Layer Defense Model
No single control closes these risks, because each layer addresses a trust assumption the others do not cover. A poisoned tool can leverage an improperly scoped token to exfiltrate data across server boundaries, and no one defense stops that chain alone. The realistic architecture is four layers working together.
- Govern the inventory. Maintain an allowlist of approved MCP servers at a gateway, verify server identity, and block dynamic discovery from untrusted networks. Pair this with an AI Bill of Materials covering agents, models, tools, and servers — auditors under the EU AI Act, ISO 42001, and NIST AI RMF are already asking for exactly this artifact.
- Treat every boundary as hostile. Retrieved content, tool descriptions, and tool outputs are all untrusted input to the next step. Scan tool descriptions for hidden instructions, detect description changes, and pin or sign tool definitions so silent redefinition is visible.
- Shrink the blast radius. Give each agent its own short-lived, narrowly scoped identity rather than reusing a shared service account. Sandbox code execution with deny-by-default egress. Require explicit human confirmation for irreversible actions — payments, deletions, merges, infrastructure changes.
- Instrument and enforce at runtime. Traditional SAST and SCA cannot see an agent's prompts, memory, or inter-agent traffic, so agentic attacks live in a layer most AppSec tooling never inspects. Capture full tool-call telemetry, establish behavioral baselines so a hijacked agent looks different from a busy one, and ship a kill switch you have actually tested.
What to Do in the Next Quarter
Security teams do not need to solve agentic AI security wholesale to materially reduce risk. The highest-leverage moves are unglamorous and can start this quarter.
- Inventory first. Map every deployed agent to the tools, data sources, and downstream systems it can reach. This is independently useful for governance even before you adopt any formal standard.
- Hunt for secrets in context. Scan agent configurations, MCP server environments, and logs for long-lived tokens. Rotate anything you find and move to short-lived scoped credentials.
- Re-baseline your risk register. The 2026 OWASP rankings elevated Excessive Agency from sixth to third on the LLM list — the largest upward move, backed by analysis of thousands of real incidents. Permission scope, not just prompt hygiene, deserves executive attention.
- Extend confidentiality controls. If your data-protection controls were built narrowly around protecting the system prompt, they no longer cover the real surface. OWASP retired System Prompt Leakage in favor of Hidden Context Exposure, which captures retrieved documents, memory, tool responses, and application state.
- Test the kill switch. An untested emergency stop is not a control. Rehearse revoking an agent's credentials and confirming it stops acting.
The strategic point is that MCP security is not a reason to avoid agentic AI — it is the cost of operating it responsibly. Enterprises that treat the tool layer as a first-class part of their attack surface, instrument it, and constrain it will ship agents faster than those that discover the gap during an incident. Enterprises that bolt security on after the agent reaches production will spend their next quarter doing forensics instead of shipping.
If you are deploying agents against production systems and are not yet certain what those agents can reach, that uncertainty is the finding. An independent assessment of your MCP endpoints, agent identities, and tool permissions is the fastest way to convert it into a remediation plan.