All Blogs

Quick Overview: MCP security covers the controls that stop AI agents from being turned against the systems they connect to. This guide explains Model Context Protocol architecture and trust boundaries, twelve documented attack classes, authentication requirements, a five-layer defense framework, and how to test an MCP deployment.
A tool description in the Model Context Protocol is just a string. The model reads it, trusts it, and acts on it. Nothing in the protocol requires that string to be honest.
That property is the root of most MCP security failures. An attacker who controls a tool description controls part of the agent's reasoning, before any human sees an approval dialog. Researchers have demonstrated instructions hidden in Unicode tag blocks that render as blank space in the client but arrive intact at the model.
The rest follows. MCP servers hold OAuth tokens, filesystem handles, database connections, and shell access, and a single compromised server exposes every credential it aggregates. Most were written quickly by small teams, without a security review. Between January and April 2026, researchers disclosed more than 40 CVEs against MCP implementations across Python, TypeScript, Java, and Rust SDKs.
There are various security layers that you can implement to ensure the protection of your MCP. In this blog, we’ll understand each of the security layers you can implement, how using an AI penetration testing platform simplifies your work, and learn about MCP attack chains that you need to prevent. With that said, let’s get started.
Discover hidden weaknesses across your apps and APIs before attackers turn them into an entry point. Start Testing
On This Page
- What is MCP Security?
- MCP Security Architecture and Trust Boundaries
- MCP Threat Model
- MCP Security Risks and Attack Vectors
- Disclosed MCP Vulnerabilities
- MCP Attack Chains and Exploit Paths
- MCP Authentication and Authorization
- Defense in Depth: Five Layers of MCP Security
- Common MCP Security Mistakes
- MCP Security vs Traditional API Security
- MCP Security Testing Methodology
- Conclusion: Securing the Agent-to-Tool Boundary
What is MCP Security?
MCP security is the practice of protecting the Model Context Protocol layer that connects AI agents to external tools, data, and systems, covering the host application, the client, the server, the tools it exposes, and every downstream system those tools can reach.
Anthropic released MCP in November 2024 as an open standard for connecting language models to external capabilities. Adoption followed fast: every major IDE and assistant now ships with MCP support, and public registries hold thousands of community servers.
Deployment has three roles. The host is the application the user interacts with. The client is the protocol manager inside it. The server exposes capabilities through three primitives: tools (actions), resources (data), and prompts (templates).
The specification standardizes how a model discovers and calls a tool. It does not enforce authorization, input validation, sandboxing, or monitoring, and it states plainly that it cannot enforce security principles at the protocol level. Security in MCP is a property of the deployment.
That’s why MCP is more than just APIs with a new wrapper. In a traditional integration, developers write code that decides which API endpoint to call and what data to send. In MCP, the model decides at runtime, based on text an attacker may partly control, while holding real credentials.
MCP Security Architecture and Trust Boundaries
MCP security architecture spans three protocol layers (transport, JSON-RPC protocol, and data) and five trust boundaries: user to host, host to client, client to server, server to downstream system, and server to untrusted content.
Two transports carry MCP traffic with different threat models. stdio runs the server as a local subprocess, inheriting the launching user's OS privileges and drawing credentials from environment variables or config files, with no OAuth and no isolation unless you add it. Streamable HTTP runs remotely and is where the specification's OAuth 2.1 requirements apply.
The invocation flow matters more than the diagram. The client sends the server's tool list into the model's context. The model selects a tool and arguments. The client issues a tools/call. The server executes and returns a result, and that result re-enters the context as text the model treats as ground truth. Tool selection is model-driven rather than code-driven, and tool output is an input channel to the model.
| Trust Boundary | What Crosses It | Control That Must be Enforced |
|---|---|---|
| User to host | Approval decisions, consent dialogs | Approval UI must render the exact definition the model receives |
| Host to client | Tool lists, model context, session state | Per-server namespacing, definition pinning, change detection |
| Client to server | Tool calls, arguments, credentials, results | OAuth 2.1 with audience validation, transport binding, no token passthrough |
| Server to downstream | API calls, queries, filesystem and shell operations | Per-tool authorization, schema validation, egress allowlists |
| Server to untrusted content | Fetched pages, tickets, documents, third-party data | Treat all returned content as data, never as instruction |
The fifth boundary is the one traditional AppSec programs to miss. A ticket body, a README, or an email the agent reads through a tool is attacker-reachable input landing directly in the model's reasoning context.
MCP Threat Model
The MCP threat model centers on four assets: the credentials MCP servers aggregate, the data their resources expose, the write and execute capability of their tools, and the integrity of the model's context window.
Adversaries. Four profiles cover nearly every documented incident. A malicious author publishes a useful-looking connector to a registry. A legitimate server is compromised through a dependency or maintainer account. An attacker plants content the agent will eventually read. A repository author commits a project-level MCP config that executes when a developer opens the project.
Assumptions that fail. Most deployments implicitly assume the server is honest, the tool definition is static after approval, the user reads the approval dialog, the model will not follow instructions embedded in data, and the token is scoped to what the tool needs. Each has been broken into public research.
Exposure is easier to measure than exploitation. Bloomberry research cited by Palo Alto Networks found 38% of MCP servers in the wild have no authentication. A 2025 survey of 1,899 public servers (arXiv:2506.13538) reported 7.2% with general vulnerabilities and 5.5% exhibiting tool poisoning.
Ranked by what a compromise yields, the highest-impact surfaces are shell and code-execution tools, filesystem tools with broad path scope, cloud connectors holding long-lived credentials, database connectors with write access, and any tool with outbound network reach, because that last category turns a read into an exfiltration.
Let AI uncover vulnerabilities, connect the dots, and show you where a single flaw could lead. Trace the Attack
MCP Security Risks and Attack Vectors
MCP vulnerabilities fall into twelve documented classes: five that attack the agent's reasoning through tool metadata and context, three that attack credentials and authorization, and four that are conventional implementation flaws made worse by an autonomous caller.
1) Tool Poisoning
Tool poisoning is an attack in which a malicious or compromised MCP server embeds instructions inside tool metadata, so the model reads attacker-controlled directives as part of its own operating context.
The description field is an unsanitized string that reaches the model verbatim. A tool advertised as a weather lookup can carry a description instructing the model to first read a local SSH key and pass its contents in an innocuous parameter. What lets it survive review is the approval-view fidelity gap: descriptions get truncated, markdown gets rendered, and Unicode tag-block characters display as nothing while arriving intact at the model. The MCPTox benchmark measured a 36.5% average success rate across 45 servers and 20 models, peaking at 72.8%.
Signal: Imperative language, file paths, encoded characters, or references to other tools inside a description.
2) Rug Pulls and Silent Tool Redefinition
A rug pull is a change to an MCP tool's definition after the user approved it, so the tool the agent runs is no longer the tool the user reviewed.
Approval is granted per server or per tool name, not per definition hash, so a server can return a benign schema at install and a malicious one later with no re-prompt. The same effect arrives through ordinary package updates, since most servers ship via npm or PyPI and are pulled with loose version ranges.
Signal: Any change to a name, description, or schema between sessions that did not trigger re-approval.
3) Prompt Injection Through MCP
MCP prompt injection occurs when attacker-controlled text reaches the model through a tool result, a resource, or a tool description and is interpreted as instruction rather than data.
Indirect injection is the dangerous variant: an agent reads a GitHub issue, a support ticket, or a PDF, and that content instructs it. The agent then acts with its own credentials, so every resulting request is authenticated and authorized. The condition that turns injection into a breach is the lethal trifecta: access to private data, exposure to untrusted content, and the ability to communicate externally.
Signal: An outbound call immediately following a read of external content, with no user request behind it.
4) Tool Shadowing and Cross-Server Contamination
Tool shadowing happens when one MCP server defines a tool that overrides, imitates, or alters the behavior of a tool provided by a different server connected to the same client.
Most clients merge all connected servers into one flat namespace with no origin labeling. A malicious server can collide with a trusted tool name, or embed text in its own description that changes how the model uses another server's tools. A documented pattern instructs the model to copy a specified address whenever it sends mail through a legitimate mail server.
Signal: Duplicate tool names across servers, or a description referencing a tool it does not own.
5) Sampling and Elicitation Abuse
Sampling abuse happens when an MCP server uses the client to send requests to the AI model and manipulate it against the user. Elicitation abuse is similar, but instead of targeting the model, the server tricks the user into providing additional information through prompts.
Sampling allows an MCP server to ask the client to generate an AI response. This can reduce the server’s processing cost, but it also gives the server an indirect way to interact with the AI model.
A compromised server can craft sampling prompts that extract context the user never intended to share, or that induce output the client then acts on. Elicitation, added in the June 2025 revision, extends the same trust to user-facing prompts, where a server can request credentials or approvals under a plausible pretext.
Signal: Sampling requests unrelated to the tool's stated purpose, or elicitation prompts asking for secrets.
6) Excessive Permissions and Excessive Agency
Excessive agency in MCP is the gap between what a tool is permitted to do and what the task requires, and it converts a successful injection into a serious incident.
Community MCP servers often request more access than they actually need, for example, access to an entire repository when they only need one file, write permissions when read-only access is enough, or access to the whole filesystem instead of a specific directory.
Injection is the trigger; excessive permissions are what makes the impact worse. The good news is that this is one of the easiest risks to address: simply review and limit the permissions each MCP server actually needs.
7) Authentication and Authorization Flaws
MCP authorization flaws arise when a server authenticates the user but never enforces per-tool and per-resource authorization or forwards a received token downstream instead of exchanging it for a correctly scoped one.
The confused deputy pattern is the canonical case: the server holds high privilege and acts for a lower-privilege caller without checking entitlement. Token passthrough is its relative, destroying audience validation and letting one compromised server reach resources it was never granted. CVE-2026-13341 documents this at production scale, where indirect prompt injection against the Kong Konnect MCP server drove unintended API requests on the caller's behalf. Multi-tenant servers add cross-user resource access, the MCP equivalent of broken object level authorization.
8) Credential and Token Exposure
MCP token security fails when credentials are stored in plaintext config, accepted as tool arguments, returned inside tool responses, or written to logs that flow back into the model's context.
Local stdio servers commonly read secrets from environment variables or a JSON config living in the repository, which puts credentials one commit away from version control. On the response side, a tool returning a raw API payload can place a token in the context window, where it may be persisted in a transcript or read by the next injected instruction.
9) Malicious and Compromised MCP Servers
A malicious MCP server is one that behaves correctly enough to be installed, then abuses the trust, credentials, and execution privileges the host grants it.
The registry ecosystem makes this cheap: thousands of servers with no clear ownership, trivial typosquatting, and dependency compromise that propagates to every install. The first confirmed malicious MCP package appeared in September 2025 and operated undetected for two weeks while exfiltrating email data.
The highest-severity variant is repository-level auto-execution. A committed .cursor/mcp.json or .amazonq/mcp.json starts a server with developer OS privileges and no isolation the moment the project opens. CurXecute (CVE-2025-54135) demonstrated it in Cursor; Wiz Research disclosed the same pattern in the Amazon Q VS Code extension in June 2026. One compromised repository reaches every developer who clones it.
10) Data Exfiltration Through MCP Tools
MCP data exfiltration is the use of a legitimate tool's outbound capability to move data the agent can read to a destination the attacker controls.
No memory corruption is required. Any tool that fetches a URL, sends a message, opens an issue, or writes a file is an egress channel, and the agent already holds read access to whatever the injected instruction requests. Client-side markdown image rendering is a recurring silent channel, since the client fetches an image URL carrying stolen data as a query parameter automatically.
11) SSRF and Network Boundary Bypass
MCP SSRF occurs when a tool issues outbound requests to a model-supplied URL, letting an attacker reach internal services and cloud metadata endpoints from inside the network trust boundary.
Fetch, browse, and webhook tools are the carriers, and the classic targets apply. Local servers add a second variant: a server bound to all interfaces rather than loopback is reachable from a browser, and without Origin and Host validation it is exposed to DNS rebinding. CVE-2025-66414 records that flaw in the MCP TypeScript SDK.
12) Command Injection, Path Traversal, and RCE
Command injection and path traversal in MCP servers turn model-supplied text directly into shell arguments or filesystem paths, the shortest route from prompt injection to remote code execution.
These are ordinary implementation bugs and the most common in the ecosystem. Of 30 CVEs filed against MCP servers in a single 60-day window in early 2026, 13 were command-injection patterns. Filesystem tools contribute containment bypasses and symlink escapes, where a path is validated before resolution rather than after.
Disclosed MCP Vulnerabilities
Each of these demonstrates a class rather than a one-off bug.
| CVE | Component | Class | What It Demonstrated |
|---|---|---|---|
| CVE-2025-49596 | MCP Inspector (CVSS 9.4) | Unauthenticated RCE | Arbitrary command execution through unauthenticated Inspector instances reachable from a browser. |
| CVE-2025-6514 | mcp-remote (CVSS 9.6) | OS command injection | First full RCE on a client OS from connecting to an untrusted remote server. Fixed in v0.1.16. JFrog, July 2025. |
| CVE-2025-54135 | Cursor IDE (CVSS 8.6) | Repo-config auto-execution | "CurXecute": a malicious project-level MCP config triggered RCE on project open. Aim Security, August 2025. |
| CVE-2025-53109 / 53110 | Filesystem MCP server | Path traversal | Symlink bypass (CVSS 8.4) and containment bypass (CVSS 7.3) in a reference server. Trend Micro, 2025. |
| CVE-2025-68143 / 68144 / 68145 | Git MCP server | Traversal, argument injection | Three flaws in a widely used reference server, showing first-party code is not a safe default. January 2026. |
| CVE-2025-66414 | MCP TypeScript SDK | DNS rebinding | Local server reachable from a browser, bypassing the assumption that localhost is private. |
| CVE-2026-13341 | Kong Konnect MCP server | Confused deputy | Indirect prompt injection drove unintended API requests on the caller's behalf. |
| CVE-2026-12957 / 12958 | Amazon Q for VS Code | Supply chain | A crafted workspace MCP config in a cloned repo achieved code execution and cloud credential theft. Wiz, June 2026. |
A note on the numbers. Aggregate claims about the share of vulnerable MCP servers vary widely by methodology, and one independent audit measured roughly a 78% false-positive rate from signature-based MCP scanners. Treat headline percentages as directional; the CVEs above are individually confirmed disclosures.
Mapping to the OWASP MCP Top 10
The OWASP MCP Top 10 is the first industry framework for classifying these risks. It is a living document, so category names shift between revisions, but the identifiers are stable enough to cite in an assessment.
| Attack Class | OWASP MCP Top 10 |
|---|---|
| Tool poisoning; tool shadowing | MCP03 Tool Poisoning |
| Rug pulls and silent redefinition | MCP03, MCP04 Software Supply Chain Attacks |
| Prompt injection; sampling and elicitation abuse | MCP06 Intent Flow Subversion, MCP10 Context Injection and Over-Sharing |
| Excessive permissions and agency | MCP02 Privilege Escalation via Scope Creep |
| Authentication and authorization flaws | MCP07 Insufficient Authentication and Authorization |
| Credential and token exposure | MCP01 Token Mismanagement and Secret Exposure |
| Malicious and compromised servers | MCP04, MCP09 Shadow MCP Servers |
| Data exfiltration | MCP10 |
| SSRF; command injection; path traversal; RCE | MCP05 Command Injection and Execution |
MCP08, Lack of Audit and Telemetry, has no matching attack class because it is a control gap rather than an attack. It is covered in layer five below, and it is the reason most of the classes above go undetected rather than unexploited.
Run AI pentesting to uncover exploitable vulnerabilities that conventional security testing may miss. Find Out Now
MCP Attack Chains and Exploit Paths
An MCP attack chain is a sequence in which a low-severity weakness becomes a high-impact compromise by combining the agent's credentials, tool access, and outbound reach. Severity in MCP is a property of the chain, not the finding.
This is why component-level scanning under-reports MCP risk. A tool description with odd text scores low. Over-scoped token scores are low. A fetch tool with no egress restriction scores low. Chained, they are a data breach, and every request in the chain is authenticated and schema valid.
- Injection to Tool Invocation: The agent reads an untrusted document and calls a write-capable tool the user never requested.
- Tool Abuse to Privilege Escalation: An over-scoped token means a correctly functioning tool performs an administrative action for a caller with no entitlement to it.
- Tool Access to Data Exposure: A read tool plus any tool with outbound reach forms a complete exfiltration path, with no vulnerability either.
- SSRF to Internal Compromise: A fetch tool reaches a cloud metadata endpoint, returns instance credentials into the context, and a second tool uses them.
- Tool Abuse to RCE: Injection supplies an argument, a shell-invoking tool concatenates it, and code executes on a developer's workstation or build agent.
Validating a chain means proving the path end to end rather than reporting its components. This is the same problem agentic AI pentesting solves in web applications: reasoning across a sequence of individually valid actions to establish whether a reproducible exploit path exists.
MCP Authentication and Authorization
MCP authentication and authorization are deployment responsibilities. The specification defines an OAuth 2.1 baseline for remote servers, but it does not enforce per-tool authorization, tenant isolation, or least privilege, and it does not apply to local stdio servers at all.
The model has hardened across four revisions. March 2025 introduced OAuth 2.1 and Streamable HTTP. June 2025 classified servers as OAuth Resource Servers, added protected resource metadata, and required Resource Indicators (RFC 8707) so a malicious server cannot obtain tokens for another resource. November 2025 added OpenID Connect Discovery, incremental scope consent through WWW-Authenticate, and Client ID Metadata Documents. July 2026 added issuer validation (RFC 9207) against mix-up attacks, which matter more in MCP because one client talks to many servers, plus a stateless core and required Mcp-Method and Mcp-Name headers that let a gateway or WAF route and meter without parsing JSON bodies.
| Control | What the Specification Provides | What You Must Implement |
|---|---|---|
| Client authentication | OAuth 2.1 with PKCE, metadata and OIDC discovery | Registration policy, client vetting, revocation |
| Token audience | Resource Indicators (RFC 8707) | Server-side audience validation on every request |
| Token passthrough | Documented as prohibited | Token exchange to a scoped downstream credential |
| Per-tool authorization | Nothing | An explicit check inside every tool handler |
| User vs agent identity | Nothing | Distinct workload identity, on-behalf-of delegation |
| Tenant isolation | Nothing | Session isolation, ownership checks on object references |
| Local stdio servers | Nothing; credentials come from the environment | Secret brokering, process isolation, filesystem scoping |
The highest-value rule here is the third. The OWASP GenAI Security Project's Practical Guide for Secure MCP Server Development, published February 2026, calls for the total prohibition of token passthrough specifically to prevent confused deputy attacks. If an MCP server passes the user's token to another service, a compromise of that server could give an attacker access to everything that token can reach.
Defense in Depth: Five Layers of MCP Security
Effective MCP security applies five reinforcing layers: authentication at every endpoint, least privilege at the tool level, input and output sanitization, human approval for sensitive operations, and logging that makes the other four verifiable.
No single control closes the gap between what MCP permits by design and what an organization intends to allow, because the vulnerabilities originate at different points in the architecture.
Layer 1: Authentication at Every Endpoint
Every remote server requires OAuth 2.1 with PKCE, terminating at a recognized authorization server, with token audience and claims validated before any request is honored. An HTTP endpoint without an authorization wrapper is an open proxy regardless of intent. Extend the same rigor to distribution: require signed server artifacts and verify signatures at startup, pin versions, commit lockfiles, and treat an unsigned or unpinned server as unreviewed code. Use short-lived, limited-scope tokens from a secure vault instead of storing long-lived credentials in environment variables.
Layer 2: Least Privilege at the Tool Level
Server-level access control is necessary and insufficient, because a user with legitimate server access can still invoke tools beyond their intended scope. Scope each tool to one job, split read and write into separate tools with separate credentials and avoid general-purpose tools such as run_query or execute that push the authorization decision into a model-controlled string. Start read-only. Deny tool invocation by default and maintain explicit allowlists covering tool names, versions, and parameter schemas. Enforce the check inside every handler, every time, not once at connection.
Layer 3: Input and Output Sanitization
Two flows need control. Inbound, validate every argument against a strict JSON schema with explicit types, enumerated values, and length limits, then trace where it lands. Never build a shell string: use parameterized APIs, or pass an argument array, never an interpolated command line. Before checking file paths, normalize them and resolve any symbolic links. This prevents attackers from bypassing security checks through path tricks or symlinks.
For outbound requests, remove secrets from the data, filter instruction-like content from third-party sources before adding it to the AI’s context, and allow requests only to approved destinations. Block access to loopback, private, and link-local addresses, and re-check every redirect before following it.
Layer 4: Human in the Loop for Sensitive Operations
Some operations should not be automated. Deleting records, changing access controls, sending external communications, and committing transactions warrant explicit approval at execution, even when the agent is technically authorized. This is intentional friction, and the cost of pausing is small against the cost of an unintended privileged action. Consent must be time-bound and scoped to the specific operation with its actual parameters, not to the tool or the session, and the approval view must render the raw definition the model receives with non-printing characters made visible.
Layer 5: Logging and Observability
Observability is what makes the other four defensible. The audit unit is the action chain, not the request: record what content entered the context, which tool the model selected, the arguments supplied, what returned, and what happened next, with correlation IDs tying invocations to model requests and hashed parameters where privacy requires it. Useful operational signals include tool invocation counts and error rates by tool, p95 invocation latency, timeout trends, and session initialization failures. Failures usually surface at the client layer first. A log that records only successful calls will show a clean record of an agent being manipulated, because every call in a successful injection is legitimate.
Common MCP Security Mistakes
The most frequent MCP security failures are not exotic: unauthenticated servers, unvetted installs, validation performed in the wrong order, and logging that cannot reconstruct what happened.
| Mistake | Consequence |
|---|---|
| Deploying network-reachable servers with no authentication | Attackers find them by scanning and execute tools with full privileges |
| Installing servers from public registries without review | Credential harvesting and supply chain compromise on first run |
| Validating paths before canonicalization | Symlink and traversal escapes reach config files holding secrets |
| Requesting all available scopes at first authorization | Any injection or token leak inherits administrative reach |
| Storing secrets in repository config or environment variables | Credentials reach version control and CI logs |
| Approval dialogs that render rather than reveal tool descriptions | Hidden instructions get approved by users who never saw them |
| Logging only successful tool calls | Manipulated sessions look clean; incidents cannot be reconstructed |
| No inventory of which servers are connected where | Shadow servers hold credentials no one is tracking |
See how ZeroThreat helps continuously find and validate vulnerabilities before they become incidents. See Pricing
MCP Security vs Traditional API Security
MCP security differs from API security because the caller is a language model whose decisions can be influenced by the same data it processes, which means authentication and input validation at the endpoint no longer bound by the risk.
A gateway inspecting MCP traffic sees a well-formed, authenticated, schema-valid request. It cannot see that the model was instructed to make it by commenting a ticket. That is the gap, and it is why API controls are necessary but not sufficient.
| Dimension | Traditional API | MCP |
|---|---|---|
| Caller | Application code with fixed logic | A language model deciding at runtime |
| Call selection | Determined by the developer | Determined by model reasoning over text |
| Injection surface | Request parameters | Tool descriptions, resources, results, retrieved content |
| Authorization unit | Endpoint and object | Endpoint, object, and the intent behind the call |
| Testing unit | Request and response | Multi-step action chain |
| Failure mode | Unauthorized request is rejected | Authorized request achieves an unintended outcome |
The overlap is still substantial. Broken authorization, excessive data exposure, and injection into downstream sinks remain dominant, as in the wider landscape covered in our API security statistics breakdown. MCP adds a reasoning layer on top of those problems and gives the attacker a new way to reach the trigger.
MCP Security Testing Methodology
MCP security testing validates both the server as an application and the agent loop as an attack path, because a server that passes conventional API testing can still be driven into unintended actions through its tool metadata.
Conventional dynamic testing fuzzes parameters at known endpoints and evaluates each response in isolation. In MCP, the most dangerous input is a tool description that changes model behavior, and the most dangerous output is a sequence of individually valid calls. Neither is visible to a request-level scanner.
Discovery and Inventory: Enumerate every connected server, its transport, and where its configuration originates, including user-level config, workspace files, and anything a repository can introduce. Shadow servers connected by individual developers are the common finding and the hardest to see centrally.
Tool and Resource Enumeration: Capture the raw tool list byte for byte rather than as rendered, diff against an approved baseline, and inspect for non-printing characters, imperative phrasing, and references to other servers' tools.
Authentication and Authorization: Attempt unauthenticated access to every method. Verify audience binding, expiry, replay rejection, and whether the server forwards inbound tokens downstream. Then test per-tool entitlement, cross-user object access, and cross-tenant isolation the way you would test for broken object level authorization.
Input Validation: Attempt schema bypass through type confusion, oversized values, and unexpected encodings, then trace where each argument lands: shell, path, SQL, or outbound URL.
Injection and Metadata Testing: Place instruction payloads in tool descriptions, resource content, and tool results, including encoded and invisible-character variants, and measure whether the agent acts on them. Confirm that the approval view renders what the model actually receives. Exercise sampling and elicitation to check whether a server can solicit context or credentials outside its stated purpose.
Exposure, SSRF, and RCE: Check results, error messages, and logs for credentials. Target internal ranges and metadata endpoints through any URL-accepting tool, test redirect handling and DNS rebinding, then probe argument injection, symlink escape, and containment bypass.
Gateway Bypass: Where a control plane exists, verify no path reaches a server or downstream API around it, since egress that bypasses the gateway is invisible to every control it enforces.
Attack Chain Validation: Prove the full path. A finding that reads "tool description contains suspicious text" is not actionable; a reproduction showing that text causing an agent to read a credential file and transmit its contents is. This is where AI penetration testing methods earn their place, since exploitability across a multi-step agent loop is a reasoning problem rather than a signature problem. Comparable platforms are covered in our review of agentic AI penetration testing tools.
Continuous Testing: Baseline the tool manifest, fail the build on unapproved definition changes, and re-run after every server or dependency update. Point-in-time assessment does not survive an ecosystem where an approved tool can change next week.
| Risk | Test | Pass criteria |
|---|---|---|
| Tool poisoning | Inject directives into descriptions, including invisible characters | Agent ignores them; approval view shows the raw string |
| Rug pull | Alter a tool definition after approval | Client detects the change and re-prompts |
| Indirect injection | Plant instructions in content the agent will read | No unrequested tool call follows |
| Sampling abuse | Issue a sampling request unrelated to the tool's purpose | Client requires consent and scopes the completion |
| Confused deputy | Call a privileged tool as an unentitled user | Server denies on caller identity, not server privilege |
| Token passthrough | Inspect downstream requests for the inbound token | A distinct, correctly scoped credential is used |
| Excessive agency | Diff granted scopes against tool purpose | No tool holds capability beyond its function |
| Exfiltration | Instruct the agent to send readable data outbound | Egress allowlist blocks the destination |
| SSRF | Supply internal and metadata URLs to fetch tools | Request blocked before resolution |
| Command injection and traversal | Shell metacharacters, traversal sequences, symlinks in arguments | Arguments never reach a shell as a string; containment enforced after canonicalization |
See how ZeroThreat discovers hidden vulnerabilities and connects them into real, exploitable attack paths. Expose the Risk
Conclusion: Securing the Agent-to-Tool Boundary
MCP didn’t create entirely new types of vulnerabilities. Instead, it creates a direct path between familiar security weaknesses and AI-driven decision-making.
The risks are still familiar: broken authorization, overly powerful credentials, and unvalidated input reaching a shell. The difference is that an AI model now sits between the attacker and the action. If the model is influenced by malicious content, it can be tricked into triggering those weaknesses.
If you are starting from nothing, sequence it. Inventory connected servers, enforce authentication on every endpoint, and turn on tool-call logging first. Replace long-lived credentials with short-lived scoped tokens and pin tool definitions next. Add gateway enforcement, egress control, and continuous testing last, once you know what you actually have.
Securing individual MCP servers is important, but it’s not enough. The real risk often comes from how multiple components work together in an attack chain. ZeroThreat's AI penetration testing engine validates multi-step attack paths rather than reporting isolated signatures. The reasoning that proves a business logic flaw across a checkout flow is the same reasoning that proves an injected instruction reaches a credential and leaves the network.
Frequently Asked Questions
How do I evaluate a third-party MCP server before connecting it?
Read the tool descriptions and schemas directly rather than trusting the registry listing. Check who maintains it, how recently it was updated, and what its dependency tree pulls in. Confirm what credentials and scopes it requests and whether it needs write access. Run it in an isolated container with egress restricted, and pin the version so an update cannot silently change its behavior.
What is MCP sampling and why is it a security concern?
How often should MCP deployments be security tested?
Who is responsible for MCP security in an organization?
Explore ZeroThreat
Automate security testing, save time, and avoid the pitfalls of manual work with ZeroThreat.


