InfoQ 中文 · 10/9/2026, 09:17:52
Securing MCP in Production: A Four-Layer Defense-in-Depth Architecture Beyond Gateways
Proposes a four-layer defense-in-depth architecture for securing Model Context Protocol (MCP) in production, moving beyond simple gateway authentication. The framework addresses critical vulnerabilities like SSRF and schema drift by enforcing controls across secure tool execution, isolated management planes, restricted outbound trust, and semantic integrity via manifest pinning.
SOURCE COVERAGEOriginal coverage
Contents20 sections
MCP security must be built in layers, covering four control planes: execution, infrastructure management, outbound trust, and semantic integrity. Each layer requires enforcement at its corresponding trust point.
Key practices include pinning tool manifests at registration, using scoped tokens for outbound requests, and embedding diff-based reviews in CI. Gateways cannot replace semantic-layer protections; CVE-2026-26118 demonstrates that relying solely on inbound authentication creates severe blind spots.
This article is intended for SREs, platform engineers, and AI infrastructure architects.

Key Takeaways
- Treat MCP security as four distinct control layers rather than a single protocol feature. Execution, infrastructure management, outbound trust, and semantic integrity each require specific enforcement points.
- Use gateways for authentication, authorization, and auditing, but do not expect them to detect semantic abuse or protect the management plane.
- Pin tool manifests at registration time to prevent schema drift and "rug-pull" behavior after approval. Adopt diff-based review as an operational model, rather than simple allow/deny gating.
- Implement egress controls and scoped tokens for outbound traffic. The SSRF vulnerability in Azure MCP Server (CVE-2026-26118) proves that inbound authentication alone is insufficient.
- Prioritize operability: CI gates, isolated tools, diff-based manifest reviews, and behavioral baselines. Standards will catch up later, but production cannot wait.
Why This Article, and Why Now
In late 2025, our team decided to adopt MCP (Model Context Protocol) as the integration layer for a production-grade multi-agent platform. We asked only one question: What architectural controls are required to trust it before large-scale automated (tool-based) execution?
The concise answer—and the thesis of this article—is that four layers are needed: secure tool execution, an isolated management plane, restricted outbound trust boundaries, and semantic integrity. Each should be enforced at its corresponding trust point, rather than relying solely on a gateway.
These layers were not immediately obvious, but the subsequent six months of vulnerability disclosures have outlined them clearly. In the first sixty days of 2026, over thirty CVE reports targeted MCP deployments.
March data from Adversa AI indicated that when scanning over five hundred MCP servers, 38% of critical endpoints lacked authentication, and 43% contained command execution vulnerabilities. Microsoft released a patch on March 10 fixing an SSRF vulnerability in Azure MCP Server (CVE-2026-26118, CVSS 8.8), which could lead to managed identity token leakage.
On March 9, lead maintainer David Soria Parra published the 2026 MCP Roadmap. The document listed "Enterprise Readiness" as a priority area but noted that among the four directions, it might be the least defined.
At the MCP Developer Summit on April 2–3, 2026, AWS and Uber shared their production architectures for MCP Gateway and Registry. Pinterest's engineering team also announced its domain-specific MCP server ecosystem. Enterprise adoption of MCP in production is outpacing the maturation of security standards.
Taken together, these are not isolated defects. They cluster around a few key boundaries, implying that architectural responses are needed rather than sporadic patches. This article describes the architectural control patterns I believe every production MCP deployment needs today. Each pattern is compatible with the current MCP specification and requires no protocol changes.
MCP Security Is a Control Plane Problem
In early implementations, the initial instinct was to place a gateway in front of MCP traffic. This is the correct instinct. A gateway is where centralized authentication, authorization, auditing, and policy evaluation occur. InfoQ recently published a detailed implementation guide for a least-privilege AI Agent Gateway using MCP, OPA, and ephemeral runners. That design comes very close to ideal.
However, a gateway is merely an enforcement point, not a complete control plane. A gateway cannot ensure that tool handlers execute parameters securely, nor can it isolate inspectors, testing tools, and management consoles surrounding MCP. It also cannot prevent MCP servers from making insecure outbound calls using overly broad credentials. Furthermore, a gateway cannot detect when a tool definition approved by the team last week has changed.
I ask the same question for every MCP failure mode: Where is the earliest trustworthy enforcement point?
For command injection, the answer is the tool handler and the CI pipeline. For exposed inspectors, it is the management plane and network boundary. For credential leakage via outbound requests, it is the egress policy and token scoping. For tool definition drift, it is the registration boundary, where manifests should be pinned, diffed, and reviewed.
From a structural perspective, this is a control plane rather than a single gateway. Distributing enforcement across these points also distributes responsibility among the teams and vendors building, hosting, and running MCP servers. CoSAI’s Shared Responsibility Framework clarifies how security responsibilities are divided within the AI stack and who is accountable when an agent fails. These enforcement points will be formalized into four control layers in the next section.
The Four Control Layers
These four layers do not constitute a rigorous taxonomy. Each layer has inherent limitations when considered in isolation; for instance, hardening the execution environment does nothing to mitigate vulnerable outbound paths, and pinning manifests cannot verify the legitimacy of an inspector’s identity. Each layer has its earliest triggering enforcement point and is typically owned by different teams. It is precisely these differences that establish them as independent security boundaries, rather than four perspectives on the same problem.
This framework is becoming increasingly common in both industry and academia. Acharya and Gupta (2026) proposed MCPShield, a security framework that categorizes threats into four attack surfaces. Rostamzadeh et al. (2026) introduced a defense layout taxonomy, noting that existing mitigations are overly concentrated at the tool layer, while host orchestration, transport, and supply chain layers remain under-protected. The current consensus is that MCP defense is a layered problem.

Table 1. Four-Layer Control Matrix
An overview of each layer follows:
- Layer 1 is Execution: the code that runs when tools are invoked.
- Layer 2 is Infrastructure Management: including inspectors, harnesses, registration interfaces, and management consoles.
- Layer 3 is Outbound Trust Boundaries: targets reachable by the server itself.
- Layer 4 is Semantic Integrity: whether tool definitions retain consistent meaning over time.
Each layer has a distinct earliest trusted enforcement point and requires different primary controls. Gateways participate in only two of these layers, and even then, only partially.
After establishing the model, the article will proceed layer by layer, providing corresponding controls and implementation code.
Layer 1: Secure Tool Execution
Layer 1 focuses on one core principle: tool handlers must treat their parameters as data, not instructions.
Among the 30 CVEs recorded in the first sixty days of 2026, 13 exhibited the same pattern: unvalidated user-controlled input reaching a shell or dynamic interpreter. For example, CVE-2026-2130 (mcp-maigret), CVE-2026-2178 (xcode-mcp-server), and CVE-2026-2131 (HarmonyOS-mcp-server) all executed tool parameters via exec(). CVE-2026-1977 used eval() on chart specification parameters in Python. CVE-2026-27203 manipulated environment variables through unvalidated newlines, causing payloads to trigger upon the next server restart.
Command injection is not a new concept. Its architectural significance in MCP stems primarily from the trust model. Developers often view tool parameters as typed JSON, accompanied by a JSON schema that describes the input structure but offers no guarantees regarding the safety of values within a shell execution environment.
Code 1. Vulnerable vs. Fixed Shell Execution Patterns
const { username } = toolCall.params;
exec(`docker run maigret ${username}`);
const { username } = toolCall.params;
execFile('docker', ['run', 'maigret', username]); Copy code
The architectural principle is that tool handlers are not convenient wrappers for tooling, but controlled input boundaries.
CI rules are more effective than code review checklists. The following Semgrep rule blocks commits passing tool handler parameters into shell interpreters during build time, covering the two languages used by most MCP server implementations.
Code 2. Semgrep Rule: Blocking Unsafe Execution in MCP Handlers
rules:
- id: mcp-unsafe-exec-js
languages: [javascript, typescript]
severity: ERROR
message: >
MCP tool handler passes a parameter into a shell interpreter.
Use execFile or spawn with array arguments instead.
pattern-either:
- pattern: exec(`...${$PARAM}...`)
- pattern: exec($CMD + $PARAM)
- pattern: eval($PARAM)
- id: mcp-unsafe-exec-py
languages: [python]
severity: ERROR
pattern-either:
- pattern: subprocess.run(..., shell=True, ...)
- pattern: os.system($CMD)
- pattern: eval($PARAM)Copy code
Note that these rules are starting points and do not cover the full production landscape. They capture common sinks, but production rule sets should extend to other interpreters, indirect sinks, and language-specific evasion techniques.
Huang et al. constructed a dataset of 114 malicious MCP servers, demonstrating that multi-component attack chains are often more effective than single-component attacks. This finding supports a defense-in-depth approach, as no single control mechanism should be expected to catch all issues. Each trust boundary requires its own safeguards.
Layer 2: Securing the Control Plane
When evaluating MCP, we found that the test harness was listening on an internal network without authentication. Closing it took only twenty minutes, yet it was open by default. Teams rapidly deploying MCP rarely notice such exposure before someone else does.
This test harness falls within the scope of the control plane: the interface for granting trust permissions and registering tools. Once compromised, attackers gain far more privileges than just a single tool invocation.
Of the 30 CVEs, six targeted MCP's development and runtime infrastructure rather than the protocol itself. For example, CVE-2026-23744 (MCPJam Inspector) exposed an unauthenticated endpoint listening on 0.0.0.0 by default, allowing arbitrary MCP server installation; CVE-2026-23523 (Dive MCP Host) exploited crafted deep links to install malicious configurations within user client applications.
This layer is critical because development environments typically offer more access than production environments, including source code, secrets, build systems, and deployment credentials. Vulnerabilities in the control plane expose the environment where trust permissions are granted and tools are registered, not just individual tool invocations.
My expectations for the MCP control plane align with those for CI/CD control planes:
- Prohibit anonymous access
- Do not expose broad internal network access by default
- Minimize filesystem reachability
- Use short-lived credentials whenever possible
- Management endpoints must require authentication and logging
Treat inspectors, harnesses, and registration interfaces as "production-adjacent" components, because they effectively are.
Layer 3: Egress Trust, Egress Restrictions, and Token Scoping
Layer 3 focuses on the targets a server can autonomously access at runtime.
The SSRF vulnerability in Azure MCP Server (CVE-2026-26118, CVSS 8.8) illustrates why inbound authentication alone is insufficient. Attackers replaced Azure resource identifiers with malicious URLs; the server used its managed identity token to make outbound calls to these URLs, allowing attackers to control which Azure resources the server could access.
We need to adopt three controls in parallel:
- Enforce authentication on every inbound endpoint.
- Implement controlled egress access (i.e., an "egress allow list") so servers can only access services required for their operation.
- Ensure downstream credentials have minimal privileges, matching the token's blast radius to the tool's impact scope.
Code 3. NetworkPolicy Restricting MCP Server Egress Traffic
# Allow egress only to internal services it needs and to DNSapiVersion: networking.k8s.io/v1kind: NetworkPolicymetadata: name: mcp-server-egressspec: podSelector: matchLabels: app: mcp-server policyTypes: - Egress egress: - to: - ipBlock: cidr: 10.0.0.0/8 ports: - port: 443 - to: - namespaceSelector: {} ports: - port: 53 protocol: UDPCopy code
In this scenario, two assumptions are worth noting. First, the cluster's CNI must actually enforce NetworkPolicy (Calico, Cilium, etc., can do this; some default installations may not). Second, listing Egress under policyTypes for selected pods denies all traffic not explicitly allowed, but pairing this with a namespace-scoped default-deny policy is clearer and avoids side effects from this choice.
Automated file search tools should not hold access tokens capable of controlling company cloud accounts. If a tool needs access to a new domain or internal service, it should be an explicit change, not implicit behavior.
This might not be the most interesting work, but it is the most important. Many agent-based systems secure the "front door" but forget what happens after entering the process.
Layer 4: Semantic Integrity and Manifest Pinning
Layer 4 focuses on semantics rather than syntax—whether tools still retain the meaning they had when trusted.
I believe the most architecturally significant attacks are not input validation vulnerabilities, but semantic attacks. Requests may be well-formed, schema-valid, and authenticated, yet still dangerous because the tool's meaning has drifted since trust was granted.
Pillar Security identified 11 categories of MCP-related vulnerabilities in March 2026 (The New AI Attack Surface: 3 AI Security Predictions for 2026), including supply chain typosquatting and cross-server context abuse. Solo.io proposed "rug-pull" attacks, where servers change behavior after registration.
These attacks cannot be prevented by input validation. The inputs are legitimate, authentication does not block them because the caller is authenticated, and gateway policies do not stop them because the requests conform to the schema.
Manifest pinning prevents this class of attack. It is important to clarify that this is a pattern implemented at the gateway level, rather than a feature of the MCP specification itself. This control is analogous to Subresource Integrity (SRI) on the Web. During registration, the gateway canonicalizes the server’s tool Manifest, hashes the names, descriptions, and parameter schemas, and stores the hash as a signature baseline. Upon reconnection or update, the gateway re-checks the hash: if it matches, the update is allowed; if it differs, the update is held pending review by a diff classifier, which distinguishes between cosmetic changes and substantive schema modifications.
Code 4. Pinning the Manifest During MCP Server Registration
import { createHash } from 'node:crypto';interface ToolManifest { name: string; description: string; parameters: Record;}const registry = new Map();const operatorQueue: Array = [];function canonicalJson(value: unknown): string { if (value === null || typeof value !== 'object') { return JSON.stringify(value); } if (Array.isArray(value)) { return '[' + value.map(canonicalJson).join(',') + ']'; } const obj = value as Record; return '{' + Object.keys(obj).sort().map(k => JSON.stringify(k) + ':' + canonicalJson(obj[k]) ).join(',') + '}';}function canonicalize(tools: ToolManifest[]): string { const sorted = [...tools].sort((a, b) => a.name.localeCompare(b.name)); return canonicalJson(sorted);}function pin(serverId: string, tools: ToolManifest[]): PinResult { const canonical = canonicalize(tools); const hash = createHash('sha256').update(canonical).digest('hex'); const existing = registry.get(serverId); if (!existing) { registry.set(serverId, { hash, tools, signedAt: Date.now() }); return { outcome: 'registered', hash }; } if (existing.hash === hash) { return { outcome: 'unchanged', hash }; } const diff = diffSchemas(existing.tools, tools); if (diff.cosmeticOnly) { registry.set(serverId, { hash, tools, signedAt: Date.now() }); return { outcome: 'auto-approved-cosmetic', hash }; } operatorQueue.push({ serverId, diff }); return { outcome: 'pending-review', hash };}Copy code
The recursive helper function sorts keys at every level to ensure that the same Manifest always produces the same hash. For production implementations, it is recommended to use established canonical JSON formats such as RFC 8785 (JCS), which also standardizes the encoding of numbers and strings.
/filters:no_upscale()/articles/securing-mcp-production-gateway/en/resources/173figure-2-1784884299695.jpg)
The diff classifier makes this control operationally sustainable. A simple allow/deny gate would generate many false positives, training automated operators to approve requests carelessly. By distinguishing between cosmetic changes and substantive schema modifications, the review queue contains only those items that truly require human decision-making.
It is recommended to combine Manifest pinning with behavioral monitoring. Monitor the endpoints each MCP server actually accesses, data movement volumes, and latency patterns. An alert can be triggered when a server approved as a file search tool begins making outbound HTTP requests to unknown domains, without needing to validate the schema legality of each request individually. Combining static trust with behavioral verification is more powerful than relying on either control alone.
Currently, both the MCP specification and the 2026 MCP roadmap remain silent on Manifest pinning. This is a capability that teams need to implement at the gateway layer.
Positioning the Gateway
The gateway remains necessary. The key is to place the gateway correctly within these layers and clearly define which functions should reside outside the gateway.
The gateway provides a suitable platform for implementing client authentication, policy-based request authorization, maintaining audit trails and request logs, enforcing rate limits (e.g., based on requests per unit of time), managing registration workflows, and centrally evaluating policies applicable to incoming requests. These examples illustrate constraints that must be enforced at the protocol layer, which is precisely what we aim to support with our gateway.
In contrast, checks related to secure execution, control plane isolation, outbound connection restrictions, and semantic drift detection occur at different layers. Therefore, the correct approach is to view the gateway as one component of a multi-layered security solution, applying additional technical controls at each earliest trusted layer.
Uber’s presentation on MCP Gateway and Registry at the April MCP Developer Summit, along with Pinterest’s practice of using dedicated MCP servers across multiple domains via a single registry with a two-tier authorization system, exemplifies this approach. The convergence of this pattern in production environments is occurring at a pace comparable to its adoption in research contexts.
A Four-Week Rollout Plan
The rollout sequence must follow the control model: deploy mature and easily verifiable controls first, followed by MCP-specific controls. Teams do not need to implement all controls simultaneously; a phased approach typically yields better results.
Table 2. Four-week rollout plan.
Week 1: Stabilize Obvious Boundaries
Week 1 focuses on closing boundaries that can be addressed without MCP-specific tooling. Require authentication for every MCP-facing endpoint, migrate inspectors, testing tools, and management consoles to isolated networks, and restrict filesystem access to only what each process actually requires.
Week 2: Enforce Path Security
Week 2 aims to make execution paths more secure. Add CI rules to flag shell interpolation, eval, or subprocess calls with shell=True reachable from tool handlers, and convert hot paths to use array argument passing.
Week 3: Restrict Trust
Week 3 involves limiting outbound trust. Use outbound allowlists at the server layer and replace broad service credentials with tool-specific scoped tokens, making outbound HTTP an explicitly granted capability rather than a default permission held by the server.
Week 4: Begin Implementing Semantic Controls
Week 4 initiates semantic controls. Enable diff-based manifest pinning reviews during registration, collect behavioral baselines for each MCP server over two to three weeks of production traffic (covering endpoints, data volume, and latency), and trigger alerts for deviations once the baseline stabilizes.
Trade-offs Are Inevitable
Each of the above controls sacrifices certain aspects—latency, maintenance overhead, or flexibility—in exchange for security. These costs are not accidental; they must be explicitly identified before implementation.
Containerized or sandboxed execution increases latency. In our tests, ephemeral-runner isolation for each tool call added 50 to 200 milliseconds compared to local execution. This is acceptable for background agent tasks but noticeable for interactive assistants.
Outbound allowlists break MCP servers requiring new external endpoints. In our scenario, most MCP servers connect to three to five internal services and one to two external APIs, making allowlists manageable but requiring manual handling per server.
Manifest pinning introduces some friction for evolving servers. Diff-based automated reviews mitigate this issue but cannot eliminate it entirely.
Behavioral monitoring incurs costs and false positives before the baseline stabilizes. In our deployments, most MCP servers stabilize after two to three weeks of production traffic, while low-traffic servers require longer. I recommend treating this timeframe as operational experience rather than a universal threshold, as the window depends on request volume and variability.
These costs are explicit. However, importantly, they are far smaller than the losses caused by credential leaks or abuse of redefined tools. In my experience, the most common failure mode is not teams over-protecting MCP, but rather underestimating the number of distinct trust boundaries they create.
Direction of Specification Development
Two developments warrant attention. First, the MCP 2026 Roadmap highlights "enterprise readiness" as a repeatedly mentioned priority, although it remains the least defined among the four areas, with core maintainers inviting input from "everyone facing challenges from production to contribution." Second, regarding the NIST AI Agent Standards Initiative launched in February 2026, the NCCoE has published conceptual papers on agent identity and authorization. As a participant in the OASIS CoSAI and IETF AGNTCY working groups, I can state that there is no consensus yet on MCP identity issues, particularly whether servers should carry their own credentials or hold delegated permissions on behalf of human users. CoSAI released research on Agentic Identity and Access Control following RSAC 2026, directly addressing this question.
Both initiatives are promising, but neither is finalized.
Real-world production deployments require Layer 1 security tooling enforcement, Layer 2 isolated control planes, Layer 3 restricted egress trust, and Layer 4 semantic integrity via manifest pinning. Each layer has an earliest point of trusted execution, which is typically not the gateway.
The lesson I leave for teams: protect the first boundary where trust is granted, not just the boundary through which traffic passes.
This is how MCP becomes trustworthy in operation.
View the original English article: Securing MCP in Production: Defense-in-Depth Beyond the Gateway