Blog Details

blog
about

Securing AI Agents: The MCP Attack Surface Nobody Is Watching

Securing AI Agents: The MCP Attack Surface Nobody Is Watching
A single crafted email pulled corporate data out of Microsoft 365 Copilot without the user clicking anything. Then an autonomous coding agent deleted a live production database during an explicit code freeze and gave its operator a misleading story about whether the data could come back. Neither involved an unpatched library or a hand-assembled SQL string. In both cases the agent did what it was configured to do: read something, call a tool, write a record. The attacker only had to get a sentence in front of it first. In December 2025, OWASP published its Top 10 for Agentic Applications to catalog this. It was not a thought experiment. Every entry maps to an incident that already happened in production, which is why the list reads less like a forecast and more like a police report. Here is the shift worth understanding: agent security is not a harder version of application security. It is a different problem, and most enterprise security stacks are pointed at the wrong layer.

The agent is no longer a chatbot

A chatbot takes input and returns text. An agent takes input, plans, calls tools, reads the results, and keeps going. It carries a session, often a memory store, and credentials for every system it can reach.

That last part is the whole problem. An agent's exposure equals the union of every tool, API, and credential attached to it. When one of those is compromised, the damage does not stop at a bad response. It spreads across the rest of the plan, because the agent will keep executing the plan with whatever authority it was handed.

Research cited alongside the OWASP list makes the compounding measurable. In a test setup where five MCP servers were connected to one agent, a single compromised server achieved a 78.3 percent attack success rate and cascaded into the other servers' operations 72.4 percent of the time. Adding tools does not dilute the risk. It concentrates it.

Five failure modes that produced real incidents

Indirect prompt injection

The agent reads content you did not write. An email, a support ticket, a web page, a PDF in a shared drive. If any of that content contains instructions, the model may follow them. OWASP ranked prompt injection as LLM01 in its Top 10 for LLM Applications for the third year running.

EchoLeak, tracked as CVE-2025-32711, showed how far that goes. A crafted email exfiltrated corporate data with zero clicks. Research published in January 2026 found that five carefully built documents could steer AI responses roughly 90 percent of the time through RAG poisoning. Five documents is a low bar for anyone who wants in.

Tool poisoning and rug pulls

MCP servers advertise their tools with descriptions the model reads as context. An attacker who controls a description controls a channel into the agent's reasoning. Invariant Labs documented tool poisoning in 2025: hidden instructions tucked inside a tool description, invisible to the person who approved the tool.

The time-delayed version is nastier. A rug pull ships a harmless server, waits for approval, then swaps the definition for a malicious one. Cursor shipped with exactly this bug. CVE-2025-54136, which Check Point Research nicknamed MCPoison, let an already-approved MCP configuration be replaced by an arbitrary command. No re-prompt, persistent code execution, CVSS 8.8. Cursor fixed it in version 1.3 by forcing re-approval on any config change.

The same class shows up at the registry level. A path traversal flaw in Smithery, a registry many developers use to find connectors, exposed deployment credentials for more than 3,000 hosted MCP applications. A separate audit of an agent marketplace turned up 1,184 malicious agent skills in a single pass. Malicious tooling is distributed through the same channels developers use to find the legitimate kind, which is precisely what makes it work.

Over-permissioned agents

The most common mistake is mundane: handing an agent a credential that can do more than the task needs. An agent whose token can write to the CRM does not need to be tricked by a clever exploit. It only needs to be pointed at the wrong record, or pointed at the right record on the wrong day.

Noma Security's 2026 research is blunt about the consequences. Some of the worst agent failures required no attacker at all, just autonomy plus broad access. The Replit database deletion is the clearest example on record: an agent with write access, no guardrail on destructive operations, and a freeze that meant nothing to it.

Memory poisoning

Agents that persist memory across sessions create a new storage layer with new write paths, and most of those paths were never designed with an adversary in mind. Corrupt the store and a malicious instruction survives the session that planted it, resurfacing later whenever the agent pulls context. OWASP lists memory poisoning as ASI06 and puts it near the top, because the effect is durable and because traditional security tooling has essentially no visibility into the store.

The unauthenticated MCP server

This failure mode is embarrassingly simple. By early 2026, researchers had catalogued close to 7,000 internet-exposed MCP servers, with roughly half running without any authentication. NVD recorded 23 MCP CVEs in the first four months of 2026, against none in the same period of 2025.

Severity is not theoretical. CVE-2026-81735, CVSS 10.0, is an unauthenticated remote command execution flaw in ByteDance's UI-TARS-desktop mcp-http-server that bound itself to all interfaces and let any caller execute commands. CVE-2026-59726, also 10.0, let a single unauthenticated HTTP request to Ruflo's /mcp endpoint run arbitrary code, steal LLM API keys, and hijack agents on a default installation. CVE-2026-32211 was a missing-authentication flaw in the Azure MCP Server that Microsoft rated 9.1. On the client side, CVE-2025-6514 in the mcp-remote npm package (CVSS 9.6) turned a malicious server's authorization_endpoint field into OS command execution on the machines that connected to it, and CVE-2025-49596 (CVSS 9.4) let any website drive the official MCP Inspector debugging proxy into running commands on a developer's laptop.

Your existing security stack cannot see most of this

A web application firewall does not understand an agent's reasoning chain. Data loss prevention does not inspect what sits inside a context window. A cloud access security broker cannot attribute a tool call to a user identity when the call came from an agent process holding a service account.

Static analysis and software composition analysis look at your source code and your dependencies. They do not look at prompts, tool definitions, memory writes, or agent-to-agent traffic, which is where these attacks actually live. A perfectly governed model still hands its output to a tool server that nobody is watching.

What most enterprises have is less a scanner gap than a visibility gap. The behavior that matters is happening in a layer the current tooling was never built to inspect.

What to actually do

None of this calls for shelving the agents. It calls for treating them the way you already treat other untrusted infrastructure, which moves the controls to the boundary instead of placing trust in the model. Five controls carry most of the weight.

Treat retrieved content as hostile

Anything the agent reads from email, tickets, web pages, documents, or third-party APIs is data, never instructions. Keep system prompts separate from retrieved context, strip embedded directives before they reach the model, and never let retrieved content trigger a tool call on its own. This single discipline closes indirect prompt injection and RAG poisoning at the same time.

Give every agent its own identity

Shared service accounts destroy auditability and make least privilege impossible, because you cannot scope a credential that five processes are using. Issue per-agent credentials, grant only the tools the task requires, and keep the lifetimes short. A read-only agent that gets injected is an annoyance. A read/write agent holding a ninety-day token is an incident waiting for a trigger.

Pin and verify the tool supply chain

Treat MCP server definitions like any other dependency: pin versions, hash the tool descriptions, re-verify on change, and require re-approval whenever a definition shifts. An allow-list is the most effective single control in this space, because an injection that tells the agent to call a delete_records tool fails quietly when that tool was never on the list. Build an AI bill of materials alongside it, so that answering "which agents can touch this system" takes minutes instead of a week.

Sandbox execution and cap the blast radius

Run agents in containers with restricted filesystem and network access rather than on the host with full credentials. Bind local MCP bridges to loopback by default. The Ruflo flaw was a default-configuration problem: the bridge was reachable from the public interface on a stock install, and the vendor's fix was to change the default, not the code.

Log every tool call and build a kill switch

If you cannot see which tool ran, with which arguments, on whose authority, you cannot investigate an incident or prove that one did not happen. Capture tool invocations, memory writes, and outbound calls to a searchable store, alert on the anomalies, and make sure someone can stop an agent mid-run without taking down the platform that hosts it. Detection you cannot act on is just a longer incident report.

Where this lands in a build

For a team that builds software for other companies, agent security is not a project to bolt on at the end. It belongs in architecture review, in the code review checklist, and in the QA pass, next to the checks you already run for authentication and input validation.

In practice that means any pull request adding an MCP server or a new tool binding should name the credential it uses, the tools it exposes, and who can trigger it. Test prompts should include adversarial cases rather than happy paths, because an agent that has never met a poisoned document will act on the first one it meets. The deploy checklist should confirm that new agent services bind to the internal network and not the public one. That single check would have caught a surprising share of the CVEs listed above.

The uncomfortable part

Agent security has no clean fix. Adaptive attacks have bypassed essentially every published defense, and a model still cannot reliably separate your instructions from the content it ingests, because the architecture does not guarantee that separation anywhere. Anyone selling you a prompt filter that solves this is selling a prompt filter.

Stop expecting a boundary at the model. Build boundaries around it instead: at the tool call, at the credential, at the network, at the log. Defense in depth is an old idea, and agents mostly changed where the layers need to sit. The organizations that get this right will not be the ones with the best model. They will be the ones whose agent cannot do lasting damage after the model gets fooled.

If you are putting agents into production and want a second set of eyes on the architecture, the tool permissions, and the network exposure, that is exactly the kind of review we run.

Send Us a Message