Award ZeroThreat Wins Bronze Stevie® Award in Tech Startup of the Year Read more
leftArrow

All Blogs

Agentic AI

An Introduction Guide to MCP in Cybersecurity

Published Date: Oct 6, 2026
Model Context Protocol Security Guide

Quick Overview: MCP security covers the controls that stop AI agents from being turned against the systems they connect to. This guide explains Model Context Protocol architecture and trust boundaries, twelve documented attack classes, authentication requirements, a five-layer defense framework, and how to test an MCP deployment.

A tool description in the Model Context Protocol is just a string. The model reads it, trusts it, and acts on it. Nothing in the protocol requires that string to be honest.

That property is the root of most MCP security failures. An attacker who controls a tool description controls part of the agent's reasoning, before any human sees an approval dialog. Researchers have demonstrated instructions hidden in Unicode tag blocks that render as blank space in the client but arrive intact at the model.

The rest follows. MCP servers hold OAuth tokens, filesystem handles, database connections, and shell access, and a single compromised server exposes every credential it aggregates. Most were written quickly by small teams, without a security review. Between January and April 2026, researchers disclosed more than 40 CVEs against MCP implementations across Python, TypeScript, Java, and Rust SDKs.

There are various security layers that you can implement to ensure the protection of your MCP. In this blog, we’ll understand each of the security layers you can implement, how using an AI penetration testing platform simplifies your work, and learn about MCP attack chains that you need to prevent. With that said, let’s get started.

Discover hidden weaknesses across your apps and APIs before attackers turn them into an entry point. Start Testing

On This Page
  1. What is MCP Security?
  2. MCP Security Architecture and Trust Boundaries
  3. MCP Threat Model
  4. MCP Security Risks and Attack Vectors
  5. Disclosed MCP Vulnerabilities
  6. MCP Attack Chains and Exploit Paths
  7. MCP Authentication and Authorization
  8. Defense in Depth: Five Layers of MCP Security
  9. Common MCP Security Mistakes
  10. MCP Security vs Traditional API Security
  11. MCP Security Testing Methodology
  12. Conclusion: Securing the Agent-to-Tool Boundary

What is MCP Security?

MCP security is the practice of protecting the Model Context Protocol layer that connects AI agents to external tools, data, and systems, covering the host application, the client, the server, the tools it exposes, and every downstream system those tools can reach.

Anthropic released MCP in November 2024 as an open standard for connecting language models to external capabilities. Adoption followed fast: every major IDE and assistant now ships with MCP support, and public registries hold thousands of community servers.

Deployment has three roles. The host is the application the user interacts with. The client is the protocol manager inside it. The server exposes capabilities through three primitives: tools (actions), resources (data), and prompts (templates).

The specification standardizes how a model discovers and calls a tool. It does not enforce authorization, input validation, sandboxing, or monitoring, and it states plainly that it cannot enforce security principles at the protocol level. Security in MCP is a property of the deployment.

That’s why MCP is more than just APIs with a new wrapper. In a traditional integration, developers write code that decides which API endpoint to call and what data to send. In MCP, the model decides at runtime, based on text an attacker may partly control, while holding real credentials.

MCP Security Architecture and Trust Boundaries

MCP security architecture spans three protocol layers (transport, JSON-RPC protocol, and data) and five trust boundaries: user to host, host to client, client to server, server to downstream system, and server to untrusted content.

Two transports carry MCP traffic with different threat models. stdio runs the server as a local subprocess, inheriting the launching user's OS privileges and drawing credentials from environment variables or config files, with no OAuth and no isolation unless you add it. Streamable HTTP runs remotely and is where the specification's OAuth 2.1 requirements apply.

The invocation flow matters more than the diagram. The client sends the server's tool list into the model's context. The model selects a tool and arguments. The client issues a tools/call. The server executes and returns a result, and that result re-enters the context as text the model treats as ground truth. Tool selection is model-driven rather than code-driven, and tool output is an input channel to the model.

Trust BoundaryWhat Crosses ItControl That Must be Enforced
User to hostApproval decisions, consent dialogsApproval UI must render the exact definition the model receives
Host to clientTool lists, model context, session statePer-server namespacing, definition pinning, change detection
Client to serverTool calls, arguments, credentials, resultsOAuth 2.1 with audience validation, transport binding, no token passthrough
Server to downstreamAPI calls, queries, filesystem and shell operationsPer-tool authorization, schema validation, egress allowlists
Server to untrusted contentFetched pages, tickets, documents, third-party dataTreat all returned content as data, never as instruction

The fifth boundary is the one traditional AppSec programs to miss. A ticket body, a README, or an email the agent reads through a tool is attacker-reachable input landing directly in the model's reasoning context.

MCP Threat Model

The MCP threat model centers on four assets: the credentials MCP servers aggregate, the data their resources expose, the write and execute capability of their tools, and the integrity of the model's context window.

Adversaries. Four profiles cover nearly every documented incident. A malicious author publishes a useful-looking connector to a registry. A legitimate server is compromised through a dependency or maintainer account. An attacker plants content the agent will eventually read. A repository author commits a project-level MCP config that executes when a developer opens the project.

Assumptions that fail. Most deployments implicitly assume the server is honest, the tool definition is static after approval, the user reads the approval dialog, the model will not follow instructions embedded in data, and the token is scoped to what the tool needs. Each has been broken into public research.

Exposure is easier to measure than exploitation. Bloomberry research cited by Palo Alto Networks found 38% of MCP servers in the wild have no authentication. A 2025 survey of 1,899 public servers (arXiv:2506.13538) reported 7.2% with general vulnerabilities and 5.5% exhibiting tool poisoning.

Ranked by what a compromise yields, the highest-impact surfaces are shell and code-execution tools, filesystem tools with broad path scope, cloud connectors holding long-lived credentials, database connectors with write access, and any tool with outbound network reach, because that last category turns a read into an exfiltration.

Let AI uncover vulnerabilities, connect the dots, and show you where a single flaw could lead. Trace the Attack

MCP Security Risks and Attack Vectors

MCP vulnerabilities fall into twelve documented classes: five that attack the agent's reasoning through tool metadata and context, three that attack credentials and authorization, and four that are conventional implementation flaws made worse by an autonomous caller.

1) Tool Poisoning

Tool poisoning is an attack in which a malicious or compromised MCP server embeds instructions inside tool metadata, so the model reads attacker-controlled directives as part of its own operating context.

The description field is an unsanitized string that reaches the model verbatim. A tool advertised as a weather lookup can carry a description instructing the model to first read a local SSH key and pass its contents in an innocuous parameter. What lets it survive review is the approval-view fidelity gap: descriptions get truncated, markdown gets rendered, and Unicode tag-block characters display as nothing while arriving intact at the model. The MCPTox benchmark measured a 36.5% average success rate across 45 servers and 20 models, peaking at 72.8%.

Signal: Imperative language, file paths, encoded characters, or references to other tools inside a description.

2) Rug Pulls and Silent Tool Redefinition

A rug pull is a change to an MCP tool's definition after the user approved it, so the tool the agent runs is no longer the tool the user reviewed.

Approval is granted per server or per tool name, not per definition hash, so a server can return a benign schema at install and a malicious one later with no re-prompt. The same effect arrives through ordinary package updates, since most servers ship via npm or PyPI and are pulled with loose version ranges.

Signal: Any change to a name, description, or schema between sessions that did not trigger re-approval.

3) Prompt Injection Through MCP

MCP prompt injection occurs when attacker-controlled text reaches the model through a tool result, a resource, or a tool description and is interpreted as instruction rather than data.

Indirect injection is the dangerous variant: an agent reads a GitHub issue, a support ticket, or a PDF, and that content instructs it. The agent then acts with its own credentials, so every resulting request is authenticated and authorized. The condition that turns injection into a breach is the lethal trifecta: access to private data, exposure to untrusted content, and the ability to communicate externally.

Signal: An outbound call immediately following a read of external content, with no user request behind it.

4) Tool Shadowing and Cross-Server Contamination

Tool shadowing happens when one MCP server defines a tool that overrides, imitates, or alters the behavior of a tool provided by a different server connected to the same client.

Most clients merge all connected servers into one flat namespace with no origin labeling. A malicious server can collide with a trusted tool name, or embed text in its own description that changes how the model uses another server's tools. A documented pattern instructs the model to copy a specified address whenever it sends mail through a legitimate mail server.

Signal: Duplicate tool names across servers, or a description referencing a tool it does not own.

5) Sampling and Elicitation Abuse

Sampling abuse happens when an MCP server uses the client to send requests to the AI model and manipulate it against the user. Elicitation abuse is similar, but instead of targeting the model, the server tricks the user into providing additional information through prompts.

Sampling allows an MCP server to ask the client to generate an AI response. This can reduce the server’s processing cost, but it also gives the server an indirect way to interact with the AI model.

A compromised server can craft sampling prompts that extract context the user never intended to share, or that induce output the client then acts on. Elicitation, added in the June 2025 revision, extends the same trust to user-facing prompts, where a server can request credentials or approvals under a plausible pretext.

Signal: Sampling requests unrelated to the tool's stated purpose, or elicitation prompts asking for secrets.

6) Excessive Permissions and Excessive Agency

Excessive agency in MCP is the gap between what a tool is permitted to do and what the task requires, and it converts a successful injection into a serious incident.

Community MCP servers often request more access than they actually need, for example, access to an entire repository when they only need one file, write permissions when read-only access is enough, or access to the whole filesystem instead of a specific directory.

Injection is the trigger; excessive permissions are what makes the impact worse. The good news is that this is one of the easiest risks to address: simply review and limit the permissions each MCP server actually needs.

7) Authentication and Authorization Flaws

MCP authorization flaws arise when a server authenticates the user but never enforces per-tool and per-resource authorization or forwards a received token downstream instead of exchanging it for a correctly scoped one.

The confused deputy pattern is the canonical case: the server holds high privilege and acts for a lower-privilege caller without checking entitlement. Token passthrough is its relative, destroying audience validation and letting one compromised server reach resources it was never granted. CVE-2026-13341 documents this at production scale, where indirect prompt injection against the Kong Konnect MCP server drove unintended API requests on the caller's behalf. Multi-tenant servers add cross-user resource access, the MCP equivalent of broken object level authorization.

8) Credential and Token Exposure

MCP token security fails when credentials are stored in plaintext config, accepted as tool arguments, returned inside tool responses, or written to logs that flow back into the model's context.

Local stdio servers commonly read secrets from environment variables or a JSON config living in the repository, which puts credentials one commit away from version control. On the response side, a tool returning a raw API payload can place a token in the context window, where it may be persisted in a transcript or read by the next injected instruction.

9) Malicious and Compromised MCP Servers

A malicious MCP server is one that behaves correctly enough to be installed, then abuses the trust, credentials, and execution privileges the host grants it.

The registry ecosystem makes this cheap: thousands of servers with no clear ownership, trivial typosquatting, and dependency compromise that propagates to every install. The first confirmed malicious MCP package appeared in September 2025 and operated undetected for two weeks while exfiltrating email data.

The highest-severity variant is repository-level auto-execution. A committed .cursor/mcp.json or .amazonq/mcp.json starts a server with developer OS privileges and no isolation the moment the project opens. CurXecute (CVE-2025-54135) demonstrated it in Cursor; Wiz Research disclosed the same pattern in the Amazon Q VS Code extension in June 2026. One compromised repository reaches every developer who clones it.

10) Data Exfiltration Through MCP Tools

MCP data exfiltration is the use of a legitimate tool's outbound capability to move data the agent can read to a destination the attacker controls.

No memory corruption is required. Any tool that fetches a URL, sends a message, opens an issue, or writes a file is an egress channel, and the agent already holds read access to whatever the injected instruction requests. Client-side markdown image rendering is a recurring silent channel, since the client fetches an image URL carrying stolen data as a query parameter automatically.

11) SSRF and Network Boundary Bypass

MCP SSRF occurs when a tool issues outbound requests to a model-supplied URL, letting an attacker reach internal services and cloud metadata endpoints from inside the network trust boundary.

Fetch, browse, and webhook tools are the carriers, and the classic targets apply. Local servers add a second variant: a server bound to all interfaces rather than loopback is reachable from a browser, and without Origin and Host validation it is exposed to DNS rebinding. CVE-2025-66414 records that flaw in the MCP TypeScript SDK.

12) Command Injection, Path Traversal, and RCE

Command injection and path traversal in MCP servers turn model-supplied text directly into shell arguments or filesystem paths, the shortest route from prompt injection to remote code execution.

These are ordinary implementation bugs and the most common in the ecosystem. Of 30 CVEs filed against MCP servers in a single 60-day window in early 2026, 13 were command-injection patterns. Filesystem tools contribute containment bypasses and symlink escapes, where a path is validated before resolution rather than after.

Disclosed MCP Vulnerabilities

Each of these demonstrates a class rather than a one-off bug.

CVEComponentClassWhat It Demonstrated
CVE-2025-49596MCP Inspector (CVSS 9.4)Unauthenticated RCEArbitrary command execution through unauthenticated Inspector instances reachable from a browser.
CVE-2025-6514mcp-remote (CVSS 9.6)OS command injectionFirst full RCE on a client OS from connecting to an untrusted remote server. Fixed in v0.1.16. JFrog, July 2025.
CVE-2025-54135Cursor IDE (CVSS 8.6)Repo-config auto-execution"CurXecute": a malicious project-level MCP config triggered RCE on project open. Aim Security, August 2025.
CVE-2025-53109 / 53110Filesystem MCP serverPath traversalSymlink bypass (CVSS 8.4) and containment bypass (CVSS 7.3) in a reference server. Trend Micro, 2025.
CVE-2025-68143 / 68144 / 68145Git MCP serverTraversal, argument injectionThree flaws in a widely used reference server, showing first-party code is not a safe default. January 2026.
CVE-2025-66414MCP TypeScript SDKDNS rebindingLocal server reachable from a browser, bypassing the assumption that localhost is private.
CVE-2026-13341Kong Konnect MCP serverConfused deputyIndirect prompt injection drove unintended API requests on the caller's behalf.
CVE-2026-12957 / 12958Amazon Q for VS CodeSupply chainA crafted workspace MCP config in a cloned repo achieved code execution and cloud credential theft. Wiz, June 2026.

A note on the numbers. Aggregate claims about the share of vulnerable MCP servers vary widely by methodology, and one independent audit measured roughly a 78% false-positive rate from signature-based MCP scanners. Treat headline percentages as directional; the CVEs above are individually confirmed disclosures.

Mapping to the OWASP MCP Top 10

The OWASP MCP Top 10 is the first industry framework for classifying these risks. It is a living document, so category names shift between revisions, but the identifiers are stable enough to cite in an assessment.

Attack ClassOWASP MCP Top 10
Tool poisoning; tool shadowingMCP03 Tool Poisoning
Rug pulls and silent redefinitionMCP03, MCP04 Software Supply Chain Attacks
Prompt injection; sampling and elicitation abuseMCP06 Intent Flow Subversion, MCP10 Context Injection and Over-Sharing
Excessive permissions and agencyMCP02 Privilege Escalation via Scope Creep
Authentication and authorization flawsMCP07 Insufficient Authentication and Authorization
Credential and token exposureMCP01 Token Mismanagement and Secret Exposure
Malicious and compromised serversMCP04, MCP09 Shadow MCP Servers
Data exfiltrationMCP10
SSRF; command injection; path traversal; RCEMCP05 Command Injection and Execution

MCP08, Lack of Audit and Telemetry, has no matching attack class because it is a control gap rather than an attack. It is covered in layer five below, and it is the reason most of the classes above go undetected rather than unexploited.

Run AI pentesting to uncover exploitable vulnerabilities that conventional security testing may miss. Find Out Now

MCP Attack Chains and Exploit Paths

An MCP attack chain is a sequence in which a low-severity weakness becomes a high-impact compromise by combining the agent's credentials, tool access, and outbound reach. Severity in MCP is a property of the chain, not the finding.

This is why component-level scanning under-reports MCP risk. A tool description with odd text scores low. Over-scoped token scores are low. A fetch tool with no egress restriction scores low. Chained, they are a data breach, and every request in the chain is authenticated and schema valid.

  • Injection to Tool Invocation: The agent reads an untrusted document and calls a write-capable tool the user never requested.
  • Tool Abuse to Privilege Escalation: An over-scoped token means a correctly functioning tool performs an administrative action for a caller with no entitlement to it.
  • Tool Access to Data Exposure: A read tool plus any tool with outbound reach forms a complete exfiltration path, with no vulnerability either.
  • SSRF to Internal Compromise: A fetch tool reaches a cloud metadata endpoint, returns instance credentials into the context, and a second tool uses them.
  • Tool Abuse to RCE: Injection supplies an argument, a shell-invoking tool concatenates it, and code executes on a developer's workstation or build agent.

Validating a chain means proving the path end to end rather than reporting its components. This is the same problem agentic AI pentesting solves in web applications: reasoning across a sequence of individually valid actions to establish whether a reproducible exploit path exists.

MCP Authentication and Authorization

MCP authentication and authorization are deployment responsibilities. The specification defines an OAuth 2.1 baseline for remote servers, but it does not enforce per-tool authorization, tenant isolation, or least privilege, and it does not apply to local stdio servers at all.

The model has hardened across four revisions. March 2025 introduced OAuth 2.1 and Streamable HTTP. June 2025 classified servers as OAuth Resource Servers, added protected resource metadata, and required Resource Indicators (RFC 8707) so a malicious server cannot obtain tokens for another resource. November 2025 added OpenID Connect Discovery, incremental scope consent through WWW-Authenticate, and Client ID Metadata Documents. July 2026 added issuer validation (RFC 9207) against mix-up attacks, which matter more in MCP because one client talks to many servers, plus a stateless core and required Mcp-Method and Mcp-Name headers that let a gateway or WAF route and meter without parsing JSON bodies.

ControlWhat the Specification ProvidesWhat You Must Implement
Client authenticationOAuth 2.1 with PKCE, metadata and OIDC discoveryRegistration policy, client vetting, revocation
Token audienceResource Indicators (RFC 8707)Server-side audience validation on every request
Token passthroughDocumented as prohibitedToken exchange to a scoped downstream credential
Per-tool authorizationNothingAn explicit check inside every tool handler
User vs agent identityNothingDistinct workload identity, on-behalf-of delegation
Tenant isolationNothingSession isolation, ownership checks on object references
Local stdio serversNothing; credentials come from the environmentSecret brokering, process isolation, filesystem scoping

The highest-value rule here is the third. The OWASP GenAI Security Project's Practical Guide for Secure MCP Server Development, published February 2026, calls for the total prohibition of token passthrough specifically to prevent confused deputy attacks. If an MCP server passes the user's token to another service, a compromise of that server could give an attacker access to everything that token can reach.

Defense in Depth: Five Layers of MCP Security

Effective MCP security applies five reinforcing layers: authentication at every endpoint, least privilege at the tool level, input and output sanitization, human approval for sensitive operations, and logging that makes the other four verifiable.

No single control closes the gap between what MCP permits by design and what an organization intends to allow, because the vulnerabilities originate at different points in the architecture.

Layer 1: Authentication at Every Endpoint

Every remote server requires OAuth 2.1 with PKCE, terminating at a recognized authorization server, with token audience and claims validated before any request is honored. An HTTP endpoint without an authorization wrapper is an open proxy regardless of intent. Extend the same rigor to distribution: require signed server artifacts and verify signatures at startup, pin versions, commit lockfiles, and treat an unsigned or unpinned server as unreviewed code. Use short-lived, limited-scope tokens from a secure vault instead of storing long-lived credentials in environment variables.

Layer 2: Least Privilege at the Tool Level

Server-level access control is necessary and insufficient, because a user with legitimate server access can still invoke tools beyond their intended scope. Scope each tool to one job, split read and write into separate tools with separate credentials and avoid general-purpose tools such as run_query or execute that push the authorization decision into a model-controlled string. Start read-only. Deny tool invocation by default and maintain explicit allowlists covering tool names, versions, and parameter schemas. Enforce the check inside every handler, every time, not once at connection.

Layer 3: Input and Output Sanitization

Two flows need control. Inbound, validate every argument against a strict JSON schema with explicit types, enumerated values, and length limits, then trace where it lands. Never build a shell string: use parameterized APIs, or pass an argument array, never an interpolated command line. Before checking file paths, normalize them and resolve any symbolic links. This prevents attackers from bypassing security checks through path tricks or symlinks.

For outbound requests, remove secrets from the data, filter instruction-like content from third-party sources before adding it to the AI’s context, and allow requests only to approved destinations. Block access to loopback, private, and link-local addresses, and re-check every redirect before following it.

Layer 4: Human in the Loop for Sensitive Operations

Some operations should not be automated. Deleting records, changing access controls, sending external communications, and committing transactions warrant explicit approval at execution, even when the agent is technically authorized. This is intentional friction, and the cost of pausing is small against the cost of an unintended privileged action. Consent must be time-bound and scoped to the specific operation with its actual parameters, not to the tool or the session, and the approval view must render the raw definition the model receives with non-printing characters made visible.

Layer 5: Logging and Observability

Observability is what makes the other four defensible. The audit unit is the action chain, not the request: record what content entered the context, which tool the model selected, the arguments supplied, what returned, and what happened next, with correlation IDs tying invocations to model requests and hashed parameters where privacy requires it. Useful operational signals include tool invocation counts and error rates by tool, p95 invocation latency, timeout trends, and session initialization failures. Failures usually surface at the client layer first. A log that records only successful calls will show a clean record of an agent being manipulated, because every call in a successful injection is legitimate.

Common MCP Security Mistakes

The most frequent MCP security failures are not exotic: unauthenticated servers, unvetted installs, validation performed in the wrong order, and logging that cannot reconstruct what happened.

MistakeConsequence
Deploying network-reachable servers with no authenticationAttackers find them by scanning and execute tools with full privileges
Installing servers from public registries without reviewCredential harvesting and supply chain compromise on first run
Validating paths before canonicalizationSymlink and traversal escapes reach config files holding secrets
Requesting all available scopes at first authorizationAny injection or token leak inherits administrative reach
Storing secrets in repository config or environment variablesCredentials reach version control and CI logs
Approval dialogs that render rather than reveal tool descriptionsHidden instructions get approved by users who never saw them
Logging only successful tool callsManipulated sessions look clean; incidents cannot be reconstructed
No inventory of which servers are connected whereShadow servers hold credentials no one is tracking

See how ZeroThreat helps continuously find and validate vulnerabilities before they become incidents. See Pricing

MCP Security vs Traditional API Security

MCP security differs from API security because the caller is a language model whose decisions can be influenced by the same data it processes, which means authentication and input validation at the endpoint no longer bound by the risk.

A gateway inspecting MCP traffic sees a well-formed, authenticated, schema-valid request. It cannot see that the model was instructed to make it by commenting a ticket. That is the gap, and it is why API controls are necessary but not sufficient.

DimensionTraditional APIMCP
CallerApplication code with fixed logicA language model deciding at runtime
Call selectionDetermined by the developerDetermined by model reasoning over text
Injection surfaceRequest parametersTool descriptions, resources, results, retrieved content
Authorization unitEndpoint and objectEndpoint, object, and the intent behind the call
Testing unitRequest and responseMulti-step action chain
Failure modeUnauthorized request is rejectedAuthorized request achieves an unintended outcome

The overlap is still substantial. Broken authorization, excessive data exposure, and injection into downstream sinks remain dominant, as in the wider landscape covered in our API security statistics breakdown. MCP adds a reasoning layer on top of those problems and gives the attacker a new way to reach the trigger.

MCP Security Testing Methodology

MCP security testing validates both the server as an application and the agent loop as an attack path, because a server that passes conventional API testing can still be driven into unintended actions through its tool metadata.

Conventional dynamic testing fuzzes parameters at known endpoints and evaluates each response in isolation. In MCP, the most dangerous input is a tool description that changes model behavior, and the most dangerous output is a sequence of individually valid calls. Neither is visible to a request-level scanner.

Discovery and Inventory: Enumerate every connected server, its transport, and where its configuration originates, including user-level config, workspace files, and anything a repository can introduce. Shadow servers connected by individual developers are the common finding and the hardest to see centrally.

Tool and Resource Enumeration: Capture the raw tool list byte for byte rather than as rendered, diff against an approved baseline, and inspect for non-printing characters, imperative phrasing, and references to other servers' tools.

Authentication and Authorization: Attempt unauthenticated access to every method. Verify audience binding, expiry, replay rejection, and whether the server forwards inbound tokens downstream. Then test per-tool entitlement, cross-user object access, and cross-tenant isolation the way you would test for broken object level authorization.

Input Validation: Attempt schema bypass through type confusion, oversized values, and unexpected encodings, then trace where each argument lands: shell, path, SQL, or outbound URL.

Injection and Metadata Testing: Place instruction payloads in tool descriptions, resource content, and tool results, including encoded and invisible-character variants, and measure whether the agent acts on them. Confirm that the approval view renders what the model actually receives. Exercise sampling and elicitation to check whether a server can solicit context or credentials outside its stated purpose.

Exposure, SSRF, and RCE: Check results, error messages, and logs for credentials. Target internal ranges and metadata endpoints through any URL-accepting tool, test redirect handling and DNS rebinding, then probe argument injection, symlink escape, and containment bypass.

Gateway Bypass: Where a control plane exists, verify no path reaches a server or downstream API around it, since egress that bypasses the gateway is invisible to every control it enforces.

Attack Chain Validation: Prove the full path. A finding that reads "tool description contains suspicious text" is not actionable; a reproduction showing that text causing an agent to read a credential file and transmit its contents is. This is where AI penetration testing methods earn their place, since exploitability across a multi-step agent loop is a reasoning problem rather than a signature problem. Comparable platforms are covered in our review of agentic AI penetration testing tools.

Continuous Testing: Baseline the tool manifest, fail the build on unapproved definition changes, and re-run after every server or dependency update. Point-in-time assessment does not survive an ecosystem where an approved tool can change next week.

RiskTestPass criteria
Tool poisoningInject directives into descriptions, including invisible charactersAgent ignores them; approval view shows the raw string
Rug pullAlter a tool definition after approvalClient detects the change and re-prompts
Indirect injectionPlant instructions in content the agent will readNo unrequested tool call follows
Sampling abuseIssue a sampling request unrelated to the tool's purposeClient requires consent and scopes the completion
Confused deputyCall a privileged tool as an unentitled userServer denies on caller identity, not server privilege
Token passthroughInspect downstream requests for the inbound tokenA distinct, correctly scoped credential is used
Excessive agencyDiff granted scopes against tool purposeNo tool holds capability beyond its function
ExfiltrationInstruct the agent to send readable data outboundEgress allowlist blocks the destination
SSRFSupply internal and metadata URLs to fetch toolsRequest blocked before resolution
Command injection and traversalShell metacharacters, traversal sequences, symlinks in argumentsArguments never reach a shell as a string; containment enforced after canonicalization

See how ZeroThreat discovers hidden vulnerabilities and connects them into real, exploitable attack paths. Expose the Risk

Conclusion: Securing the Agent-to-Tool Boundary

MCP didn’t create entirely new types of vulnerabilities. Instead, it creates a direct path between familiar security weaknesses and AI-driven decision-making.

The risks are still familiar: broken authorization, overly powerful credentials, and unvalidated input reaching a shell. The difference is that an AI model now sits between the attacker and the action. If the model is influenced by malicious content, it can be tricked into triggering those weaknesses.

If you are starting from nothing, sequence it. Inventory connected servers, enforce authentication on every endpoint, and turn on tool-call logging first. Replace long-lived credentials with short-lived scoped tokens and pin tool definitions next. Add gateway enforcement, egress control, and continuous testing last, once you know what you actually have.

Securing individual MCP servers is important, but it’s not enough. The real risk often comes from how multiple components work together in an attack chain. ZeroThreat's AI penetration testing engine validates multi-step attack paths rather than reporting isolated signatures. The reasoning that proves a business logic flaw across a checkout flow is the same reasoning that proves an injected instruction reaches a credential and leaves the network.

Frequently Asked Questions

How do I evaluate a third-party MCP server before connecting it?

Read the tool descriptions and schemas directly rather than trusting the registry listing. Check who maintains it, how recently it was updated, and what its dependency tree pulls in. Confirm what credentials and scopes it requests and whether it needs write access. Run it in an isolated container with egress restricted, and pin the version so an update cannot silently change its behavior.

What is MCP sampling and why is it a security concern?

How often should MCP deployments be security tested?

Who is responsible for MCP security in an organization?

Explore ZeroThreat

Automate security testing, save time, and avoid the pitfalls of manual work with ZeroThreat.