2026-06-18
MCP: an engineer's guide to the standard interface for connecting enterprise systems to AI
An in-depth look at how the MCP protocol reduces enterprise AI integration complexity from M×N to M+N. Covers the JSON-RPC communication model, authentication, authorization and multi-tenant isolation, security risk mitigation, and deployment architecture selection across five industry patterns. Whether you are evaluating MCP for enterprise integration or weighing it against custom adapters, this engineer's breakdown will help you make an evidence-based decision.
What old pain point MCP actually solves: from M×N to M+N
Let's start with the math. Suppose your team maintains 10 AI agents that need to access 20 internal services—CRM, the ticketing system, the data warehouse, the knowledge base, and so on. Each agent has to understand every service's interface semantics, authentication method, and response structure on its own. That means you are actually holding 10 times 20, or 200, sets of adapter logic. Whenever a service changes a field, you have to track down every agent that references it; whenever an agent switches models, you have to revalidate its compatibility with every service. The maintenance cost of this mesh-like coupling does not grow linearly—it balloons as the product of the two sides.
What really gives engineering teams headaches has never been writing the first version of an integration—it is changing the second. In a mesh, any change can trigger regressions in unexpected places, because nobody can say for sure "who owns this particular connection." During code review there is no unified contract to check against, and when something breaks in production it is hard to pinpoint which adapter is to blame. The bigger the system, the heavier this hidden debt becomes. When you look back at that number, 200, you realize it measures not just workload but how fragile the system is.
The idea behind MCP (Model Context Protocol) is, in essence, to insert a layer of standard contract between the two sides. AI agents no longer integrate with each service directly; they all speak to the protocol. Services no longer customize for each caller; they expose their capabilities once, according to the protocol. The multiplicative mesh is thus split into two additive edges: each of the M AI agents implements a protocol client once, each of the N services implements a protocol server once, and the integration relationships converge from M×N to M+N. The earlier scenario of 200 integrations becomes, under MCP, 10 plus 20—30 independent implementation units, each with clear ownership and well-defined boundaries.
The benefit of this convergence is not just numerical. More importantly, changes become localized: when a server adjusts its interface, clients are unaffected as long as the protocol contract holds; when you add a new AI agent, it can naturally consume every service already on the protocol, with no one-by-one adaptation. Maintenance responsibilities become clear as well—the server is responsible for describing its capabilities accurately, the client is responsible for invoking them correctly, and the two are aligned by the protocol rather than by implicit conventions that live in some engineer's head.
Looking at where the industry is heading, MCP is settling into the de facto interface for connecting large language models (LLMs) to private data and internal enterprise APIs. This matters because a de facto standard means the ecosystem gravitates toward it on its own: tool vendors are willing to ship SDKs for it, platforms are willing to offer hosting for it, and internal enterprise teams find it easier to agree that "this is what we'll use." Once a protocol crosses the tipping point from "optional" to "default," the marginal cost of adopting it keeps falling, while the hidden cost of not adopting it keeps rising.
Notion's evolution is a persuasive case in point. According to what the company has shared publicly, its integration layer went through 5 rewrites before adopting MCP, and the number of tools it exposed externally exceeded 100. Once your capability surface reaches that scale, continuing to feed each client through piecemeal adapters becomes an unsustainable maintenance burden—which is exactly what the M×N mesh looks like in a real business. Notion ultimately chose MCP as its unified interface layer and settled on a remote MCP Server approach to support concurrent access from multiple clients. Read the decision another way: the protocol was not chosen because it was trendy; at a certain scale, a unified contract goes from "nice to have" to "no other choice." What can hold up more than 100 tools with multiple clients hitting them at once is a single set of conventions, not a hundred informal understandings.
So the judgment this section aims to make clear is this: what MCP eliminates is not the workload of any single integration, but the multiplicative structure of integration relationships itself. If you only have a handful of AI agents and services, the pain of mesh coupling may not be obvious, and writing a few adapters yourself will work. But as soon as you foresee growth on both sides, the M×N curve will bite earlier than you expect. Straightening it into M+N ahead of time buys you composure with every change that follows. In the next section we open up the protocol to see how this contract uses JSON-RPC, Schema, and streaming responses to make "speaking the same language" a reality.
The communication model unpacked: JSON-RPC 2.0, Schema, and streaming responses
Strip away the various wrappers and what MCP runs on the wire is JSON-RPC 2.0. This is worth clarifying first, because it directly determines how you debug and troubleshoot. What you see as "calling a tool" is, at the protocol layer, nothing more than a request with a method and params, and a response carrying either a result or an error code. Once you understand this layer, many integration problems that seem mysterious return to familiar territory: capture packets, inspect messages, compare fields.
What really tends to get stuck during integration testing is the structure of tool responses. The client does not unconditionally trust whatever JSON the Server returns; it validates against the Schema. In the response, content is an array, each item in the array must declare its type, and text items must also carry the corresponding text field. If a required field is missing or a type does not match, what the client shows is often not "an error" but "the tool doesn't seem to have worked"—the model receives content it cannot parse and simply acts as if nothing happened. So when troubleshooting this kind of problem, the first step is not to suspect business logic but to pull the raw response and compare it field by field against the Schema. Separating validation failures from business failures saves a great deal of pointless guessing.
Why an enterprise-grade Server can't stop at stdio
Many people get started with MCP in stdio mode: spin up a local process, use standard input and output as the pipe, and you have it running in a few minutes. This mode works for local tools and demos, but it falls short in an enterprise environment. The stdio transport is a one-to-one pipe between processes; it has nowhere to put identity and no boundary to carry tenant isolation. Once you need multiple clients, multiple teams, or even multiple customers to share the same Server, this pipe becomes a ceiling.
To really go to production, remote deployment is unavoidable, usually via one of two routes: HTTP SSE or WebSocket. Switching to network transport is not just making the "pipe" longer—it finally gives you a place to attach authentication headers, carry tenant identifiers on the connection, and perform connection-level audits. In other words, moving from stdio to remote is essentially going from "a tool" to "a service." The former is responsible only for itself; the latter is responsible for the boundaries of every caller. Section three covers authentication, authorization, and multi-tenant isolation in detail; for now, let's establish one judgment: any Server that must support access from multiple parties should have its transport layer planned as a service from the start, rather than migrated after problems appear.
When returning large data, serializing it all at once can be fatal
There is another class of problem that never surfaces at the demo stage and only blows up when real data arrives. Imagine a log-query tool: during development you test it with a few dozen lines of sample data and it runs smoothly. After launch, a query hits a log file several hundred megabytes in size, and the tool tries to serialize the entire result into JSON in one go—memory spikes instantly, and the request either times out or takes the process down with it. This is not a fluke; it is the inevitable outcome of the "return everything" pattern once data grows.
The fix is to hand control of "how much to fetch" and "where to fetch from" back to the caller—in other words, pagination or streaming reads. One approach that has proven fairly robust in practice: each call fetches only a fixed number of rows by default, say a cap of 1,000 rows, and an offset parameter scrolls through the data; if more is needed, send another request. The benefit is twofold. First, the memory footprint of any single call has an upper bound and won't spiral out of control as the underlying data grows. Second, the model can decide whether to keep paging based on what it has already read, instead of being forced to swallow in one gulp content it could never use.
When designing tool interfaces, my recommendation is to treat pagination as a default capability rather than something you bolt on once you hit big data. The criterion is simple: if the amount of data a tool returns is determined by external input and has no controllable upper bound, reserve limit and offset from the very first version. Loading an entire file or table into memory and then returning it will always be correct on sample data and will sooner or later break on production data. Drawing the boundaries in advance is far cheaper than firefighting after the fact.
What enterprise integration requires (1): authentication, authorization, and multi-tenant isolation
When you plug an MCP Server into enterprise systems, the real work is not in defining tools but in whether the identity chain holds up end to end. MCP itself only specifies how messages are transmitted; it does not specify "who initiated this call, and which data they are allowed to see." You have to fill in that part yourself at the transport and session layers. Let's first separate the three concerns of authentication, authorization, and isolation—they solve three different problems and are easily conflated.
Authentication: verify identity the moment the connection is established
Authentication answers "who are you." When MCP runs over Streamable HTTP, whether you are establishing a long-lived SSE connection or making a single POST call, identity information has to ride on the HTTP headers, usually as a JWT inside Authorization: Bearer <token>. There is an engineering pitfall here: don't wait until a specific tool is called to validate; complete the validation during the connection handshake. The reason is that SSE is a long-lived connection—once you let an anonymous connection in, all subsequent streamed pushes assume the channel is trusted. It is like letting someone slip through the front door and then questioning people room by room: costly, and easy to miss someone.
The advantage of a JWT is that it is self-contained: the Server can read the issuer, expiration time, and user identifier without going back to a session store every time. But self-contained also means hard to revoke—once a token is issued, it remains valid until it expires. So in enterprise scenarios, keep the validity period short and pair it with a refresh mechanism, rather than taking the easy route of issuing a long-lived token and parking it in the Agent's configuration. If such a long-lived token leaks, the scope of investigation extends to every tool it can reach.
Multi-tenant isolation: Session and User jointly define the data boundary
Once authentication passes, the next and thornier problem is isolation. An MCP Server often has to serve multiple tenants and multiple users at the same time, and SSE is a stateful long-lived connection, which calls for an entirely different approach from traditional stateless REST. A workable approach is to bind a Session ID to each SSE connection, with the Server maintaining the mapping between the Session and the connection; when actually drawing the data boundary, layer on the User ID decoded from the token. The Session ID tracks "which active connection this is," and the User ID tracks "who is behind this connection"—only by combining the two can you carve out an isolated data sandbox for each tool call.
Why not rely on just one of them? With only a Session ID, identity is lost when the connection drops and reconnects, and the Session itself carries no permission semantics. With only a User ID, you can't distinguish between multiple concurrent sessions of the same user, nor promptly reclaim the context held by a connection that has already dropped. Only by putting the two together can you both trace the person and manage the connection lifecycle. In code, every tool's execution entry point should be forced to carry the "current User + current tenant" context, rather than trusting each tool implementation to decide for itself—pushing isolation logic down into the framework layer is far more reliable than scattering it across individual tools.
Authorization: from single sign-on to trust propagation across Servers
When an enterprise has more than one MCP Server, authorization shifts from a problem "inside a single Server" to one of "trust propagation across Servers." To complete a task, an Agent may need to call several Servers in sequence—documents, tickets, databases, and so on. If every Server makes the user go through a separate OAuth authorization, users fall into authorization fatigue, and IT loses its global view of "who authorized what"—which is exactly where large-scale Agent deployments most easily spin out of control.
The industry currently has two approaches worth referencing. The first uses the identity provider as the trust hub: WorkOS's Cross-App Access treats the IdP as a trust bridge between MCP Servers. The user signs in once via SSO at the IdP, and the issued Identity JWT is recognized across multiple Servers, eliminating repeated authorization Server by Server. Its value is not just fewer consent clicks; more importantly, authorization relationships are consolidated back into the IdP as a single auditable source. The second hands OAuth over to a gateway: Cloudflare's Managed OAuth lets internal applications already protected by Cloudflare Access become accessible to AI agents with one click, without building and maintaining a separate OAuth authorization flow for each MCP Server. For operations teams, this means consolidating scattered authorization configurations into one place.
The shared logic of these two approaches is worth remembering: authorization fatigue and audit blind spots arise, at bottom, because trust is scattered into the hands of each Server. Whichever vendor's solution you use, the engineering direction should be to consolidate the issuance and revocation of trust into a centralized, auditable step, rather than letting it grow unchecked throughout the system.
A few judgments for deployment
- Put authentication at the connection handshake, not at the tool call; with long-lived SSE connections especially, beware of the "once it's in, it's trusted" assumption.
- Keep JWT validity short and pair it with refresh; never hard-code long-lived tokens in Agent configurations.
- Push isolation logic down into the framework layer: the Session ID governs the connection, the User ID governs identity, and together they define the data sandbox for each call.
- In multi-Server scenarios, prioritize consolidating authorization in the IdP or a gateway; get the audit trail working first, then worry about fine-grained permissions.
These three concerns have an order: without working authentication, isolation is out of the question; without clear isolation, cross-Server authorization only amplifies risk. Verify them layer by layer in that order—don't skip ahead.
MCP or custom adapters: an engineer's trade-off checklist
The essence of this choice is not "which technology is more advanced" but "how much of your existing assets you want to reuse, and how much uncertainty you are willing to take on." Let's first lay out the cost of each path.
The benefit of custom adapters is certainty: how interfaces are defined, how authentication is attached, how errors are handled—all of it is in your hands, and when something goes wrong you know where to fix it. The cost is that every new model client and every new internal system requires another round of glue code, and once the numbers grow it degenerates into M×N repetitive work. MCP's value lies precisely in collapsing that repetition into a single unified surface, but the standardization it buys also means accepting the protocol's current design trade-offs.
Understanding MCP's engineering role makes things much clearer. In the Java ecosystem, the Spring AI community uses it as an isolation layer between models and enterprise systems—you don't need to refactor existing services to make them callable by an LLM; instead, within a Spring Boot project, you wrap existing capabilities in a protocol layer and expose them. The key is that this wrapper doesn't force you to tear down what already works: existing authorization chains, monitoring instrumentation, logging and alerting, and release and rollback processes can all stay as they are. For a team with a mature operations system, this approach of "adapt at the boundary, keep the interior unchanged" is far less risky than starting over.
The reuse mindset applies equally to client-side compatibility. In practice, a number of model clients only support stdio, the local standard input/output channel—desktop tools such as Claude Desktop and Cursor, for example—and cannot connect directly to remote services deployed on the intranet. Rather than changing the server to accommodate the client, a local relay script is enough: the script reads the Token from an environment variable, forwards local stdio requests to the remote enterprise Server, and passes the results back. A thin layer of forwarding logic connects "local-only" clients to a "centrally deployed" backend, without leaking credentials into code and without the cost of rewriting the service. Small tools like this are often the most cost-effective first step in early deployment.
But don't treat MCP as a silver bullet. One fact must be acknowledged: the protocol was not designed with enterprise scenarios as its primary goal, and that has left several obvious gaps. Authentication and authorization capabilities are weak—multi-tenant isolation, fine-grained authorization, and integration with existing identity systems all have to be added on top by the enterprise itself. Deployment forms are fragmented—local, behind a gateway, in the cloud, each with its own way of connecting, and no one-size-fits-all standard path. On top of that, the protocol is still evolving, and the integration logic you write today may be tomorrow's technical debt. These uncertainties add up, and they are exactly why many enterprises are still watching from the sidelines and holding off on production.
So the decision line can be drawn like this:
- You want to consolidate integration surfaces, there are many client types and many internal systems, and the team already has a reliable security and operations foundation—prefer MCP for boundary adaptation, wrapping existing assets rather than rewriting them.
- Authentication, authorization, tenant isolation, and compliance audits are hard constraints that the protocol's current capabilities cannot fully cover—either build your own control plane outside MCP to close the gaps, or keep custom adapters for those scenarios.
- Very few integration targets, a short lifecycle, and almost no expansion expected—build it yourself; the abstraction cost of introducing a protocol is not worth it.
In the end, MCP solves the scale problem of "repeated integration," not the depth problem of "enterprise governance." Position it as a standardized access layer rather than a complete enterprise solution, and your trade-offs won't go astray: reuse existing assets wherever you can, and diligently build the governance capabilities the protocol doesn't cover.
Three categories of security risk to think through before going to production
When you move MCP from a development machine into production, the real concern is not whether the protocol itself works, but that it takes a set of capabilities that used to be scattered everywhere and exposes them uniformly as tools the model can call directly. This unification brings convenience, but it also concentrates risk onto a single surface. Working backward from engineering consequences, the recurring problems in production deployments fall into three categories, which are clearer when examined one by one.
The first is authorization abuse. The symptom is often this: an Agent uses a set of credentials to call a tool it should never have touched, and when you want to cut it off immediately, you find there is no central place to revoke access with one click. MCP lets a model chain together multiple Servers; once credentials are issued, they flow through the call chain, and if revocation is scattered across each Server's own implementation, you lose a unified emergency brake. This is not an abstract worry—when an erroneous call involves write operations or external interfaces, the lack of centralized revocation means the time to stop the bleeding stretches from seconds to minutes or even hours.
The second is prompt injection. Content returned by an MCP Server enters the model's context, and much of that content comes from untrusted sources: scraped web pages, user-uploaded documents, responses from third-party APIs. An attacker only needs to hide instructions in that data, and the model may misread "data" as "commands," triggering unauthorized calls. What makes this class of attack stealthy is that it doesn't target your protocol layer; it targets the model's assumption of trust in its context. Any pipeline that feeds external content directly to the model should assume that content may carry malicious instructions.
The third is supply chain risk. Connecting a third-party Server brings it inside your trust boundary. If that Server is controlled by an attacker, or an update smuggles in malicious logic, every result it returns and every tool it declares can become an attack vector. Unlike traditional dependency poisoning, an MCP Server is a live runtime endpoint: what you are trusting is not just a piece of code, but its behavior for as long as it stays online.
These three risks are not isolated. Loose authorization enlarges the blast radius of injection, and a compromised Server is the source of both injection and unauthorized access. So pre-launch security design must treat them as a whole, rather than as three separate tickets each handling its own piece.
| Risk type | Trigger scenario | Engineering response |
|---|---|---|
| Authorization abuse | Credentials flow through the call chain and cannot be revoked centrally | Build centralized credential issuance and revocation; apply least privilege to tool calls |
| Prompt injection | External content enters the model's context via a Server | Tag and isolate content from untrusted sources; add human or rule-based checks before critical calls |
| Supply chain risk | A third-party Server is compromised or an update is poisoned | Maintain an allowlist of trusted Servers; monitor for anomalous changes in their behavior and responses |
This approach is not theoretical. Cloudflare's published enterprise MCP reference architecture is built around exactly these three lines—the company has fully shifted internally to MCP-driven Agent workflows, and it makes authorization control, prompt injection protection, and supply chain mitigation core components of the reference architecture. What's worth borrowing is not any specific configuration but its orientation of moving security forward into the architecture layer rather than patching after the fact: when you decide which Servers to trust, how credentials are consolidated, and how context is isolated, those decisions already define your attack surface.
Here is an actionable test: if you haven't worked out "how to stop all calls within one minute when something goes wrong," don't rush to connect MCP to systems that can write data or move money. Centralized revocation, boundaries for untrusted content, and an allowlist of trusted Servers—if any one of these three is missing, production risk will spread rapidly along MCP's connections. Security here is not an item on a compliance checklist; it is the precondition for whether this architecture can go to production at all.
What MCP doesn't solve: registries, supply chain, and access control
Getting the MCP protocol itself working and running it stably in production are two different things. The protocol specifies how clients and servers talk, but not where those servers come from, who vouches for them, or who may call them. In other words, MCP solves the problem of "how interfaces align" but leaves a gap in "how the system is governed." An organization that focuses only on the protocol layer to get a demo running will very likely stumble in three places after launch: not finding the right tools, bringing in untrusted tools, and being unable to constrain who uses which tools. These three gaps map exactly to three pieces of infrastructure: a registry, the software supply chain, and access control. I call them MCP's "homework beyond the protocol."
Start with the registry. When your environment has only three to five MCP servers, hand-writing addresses in a config file is perfectly adequate. But once internal tools scale up to dozens or hundreds of servers scattered across teams and network segments, and the LLM has to pick the right tool for the current task from a long list of candidates on every inference, tool selection itself becomes a retrieval problem. Without a unified registration and discovery mechanism, you either rely on people to maintain an ever-growing list or let the model repeatedly trial-and-error among too many irrelevant tools, and both latency and erroneous calls go up. Registering all MCP servers in a centralized directory and retrieving them by capability description is what lets the model choose from a reasonable candidate set, and that is the prerequisite for scale.
The second piece is the software supply chain. An MCP server is essentially executable code written by someone else, running in your environment, that a model can trigger automatically. The risk level of this combination is far higher than casually running npm install on a front-end dependency—it carries the usual risk of dependency-chain poisoning, plus the amplifying effect of "the model autonomously deciding to call it." So the source of each server must be controlled: who published it, which version it is, whether it has passed internal review—this information must be locked down at the point of introduction, not traced after something goes wrong. Put another way, if you can't answer "which version is the tool the model is calling on this machine, and where did it come from," you haven't passed the supply chain test.
The third piece is access control. Protocol-level authorization answers "is this connection legitimate," but the governance layer has to answer "can this role use this tool, and should this tenant's request be routed to that server." Call policies for local and remote servers often differ, and so do the authorization boundaries for sensitive and ordinary tools. All of this requires a separate layer of control on top of the protocol—MCP won't do it for you.
Ready-made engineering solutions are already filling these three gaps. Nacos MCP Router is designed along exactly these three lines: it provides MCP server search, letting the model efficiently locate the right tool from the registry directory, directly easing the tool-selection efficiency problem; it provides server onboarding, bringing the introduction step into a controlled process, which addresses supply chain security; and it adds proxied tool invocation, giving switching between local and remote servers a single unified entry point instead of hard-coded addresses scattered everywhere. Together, the three capabilities link "finding tools, trusting tools, and calling tools" into a single governance chain.
On the foundation side, the official release of Nacos 3.0 extended its core strengths (service discovery, registration, and dynamic configuration) to MCP scenarios. This means MCP servers can reuse mature registry capabilities rather than building a separate discovery mechanism for AI tools. For teams already on Nacos 2.0, Alibaba Cloud's MSE Nacos commercial edition (Platinum) offers a smooth migration path from 2.0 to 3.0 with MCP-oriented enhancements, keeping migration costs relatively manageable. This is especially important for existing systems—most enterprises won't tear down their current service governance system just to connect AI, and being able to grow MCP capabilities on the existing foundation greatly reduces resistance to deployment.
In terms of hands-on cost, the barrier to adopting this kind of tool is not high. Nacos MCP Router can be launched with zero installation via npx (nacos-mcp-router@latest), supports both stdio and SSE, and its configuration is pared down to three parameters: the Nacos address, username, and password. For engineers, this means you can first quickly verify locally whether registration and routing behave as expected, then gradually push toward production, instead of committing to heavy integration work up front.
The judgment to emphasize is this: the registry, supply chain control, and access control are not solved once and for all by adopting some tool; they are ongoing operational responsibilities. Tools only make governance actions executable. What truly determines your security posture is whether you have turned "tool registration, source review, and permission assignment" into institutionalized processes. Getting the MCP protocol working is the entry point; only by filling in these three pieces of surrounding infrastructure are you truly qualified to use MCP at scale in the enterprise. Think this layer through first, then decide how to roll out—it is much easier than doing it the other way around.
Choosing a deployment architecture: five patterns matched to industries
When discussing deployment architecture, first isolate two variables: where the proxy sits and where the MCP servers sit. The proxy consolidates internal enterprise requests into a single entry point—handling authentication, rate limiting, auditing, and protocol conversion; the servers are the end that actually holds the tools and data. Each variable has two possible placements, "local" and "remote," and the combinations, plus the question of "whether to have a proxy layer at all," produce several typical forms. Choosing an architecture is not about picking the most advanced one; it is about whether your data can leave the domain, how much concurrency you face, and where your compliance lines are drawn. Get the constraints clear first, and the pattern will naturally narrow down.
First ask whether data can leave the premises, then how much concurrency you must handle
The decision logic really comes down to two main threads. The first is the data boundary: whether sensitive data is allowed to leave the data center or the local network. The harder this line, the more the servers must be pulled on-premises and the more remote calls must be cut. The second is the load profile: whether requests are steady internal traffic or public-facing open traffic with sudden spikes. The less predictable the load, the more you need servers in a remote environment that can scale elastically. Where the two threads intersect, industry differences emerge.
Finance: local proxy plus local servers
Financial scenarios are characterized by highly sensitive data and close regulatory scrutiny, but relatively manageable internal concurrency. The sensible form here keeps both the proxy and the servers local—requests come in from the intranet, pass through the local proxy for unified authorization and auditing, and then hit MCP servers also deployed on the intranet, never leaving the domain. The extra local proxy layer is not redundant; it handles access control and the operational audit trail, which directly meets finance's hard requirement for auditability. The cost is somewhat weaker elasticity, but internal financial traffic doesn't rely on burst elasticity anyway, so the trade-off is worth it.
Government: connect directly to local servers and keep the path as short as possible
Government systems often have data security requirements a notch stricter than finance; red lines around classified information and China's Multi-Level Protection Scheme (MLPS) allow no unnecessary relays. In this case even the proxy layer can be dropped, with clients connecting directly to local servers—the shorter the path, the smaller the exposure surface and the greater the control. This is the most secure of the patterns, essentially trading "one fewer hop" for "one fewer attack surface." The cost is weaker unified governance—without a proxy layer to consolidate things, authentication, rate limiting, and auditing must be implemented separately by each server. It therefore suits scenarios with few services, closed boundaries, and a willingness to sacrifice flexibility to push security to the maximum.
Internet: proxy plus remote servers, built for elasticity
The tension in internet businesses is the reverse of the previous two: concerns about data leaving the domain are relatively low, but concurrency pressure is high and the gap between peaks and troughs is large—being unable to handle the load is the real problem. Here the servers go to the cloud, fronted by a proxy that acts as the traffic entry point and distributes requests to a horizontally scalable server cluster. Add instances when traffic rises, scale back when it falls, and both cost and capacity track the load. The core requirement of this pattern is elastic scaling under high concurrency; the fixed-capacity mindset of on-premises deployment is actually a burden here.
Hybrid cloud: local proxy plus remote servers, switching smoothly between the two
Large enterprises are rarely purely on-premises or purely in the cloud; more often they follow a hybrid cloud strategy—some sensitive workloads stay local, some elastic workloads move to the cloud, and there may also be multi-region, cross-geography global deployments. The corresponding form is a local proxy plus remote servers: the proxy stays local and controls the entry point and policies centrally, while the servers can be placed locally or in the cloud as needed. Its value lies in "smooth switching"—under the same access layer, which side resources sit on can be scheduled, and the business doesn't need to be re-engineered when underlying resources migrate. For global enterprises with multi-region deployments, this approach of keeping the control plane consolidated locally while spreading the execution plane across regions preserves unified governance and gains the latency advantage of nearby deployment. The cost is the most complex architecture: cross-domain networking, consistency, and failure domains all need dedicated design, and small and midsize teams have no need to force it.
Break governance into three concerns, then find a home for each
The five patterns above address "where things go," but an enterprise-grade solution also has to answer "how they are managed." Broken down, governance needs amount to roughly three concerns: how services are discovered and registered, how dependency sources are controlled, and how access entry points are consolidated and secured. These can be assigned to a registry, a supply chain control gateway, and a secure access gateway, respectively. A common combination in the open-source ecosystem maps onto this—Nacos as the MCP registry for service discovery, Nacos Router for fine-grained control of the software supply chain, and Higress as the entry gateway for secure access; together, the three cover the backbone of an enterprise-grade MCP solution. The emphasis here is not on any specific tool but on the way of splitting the problem: registration, supply chain, and access are three surfaces, each managed on its own. Matching candidates to these three responsibilities during selection is more reliable than hunting for an "all-in-one suite" from the start.
One final reminder: these five patterns are not a ladder, and none ranks above another. The cost of choosing wrong usually comes from "using a finance architecture to carry internet-scale concurrency" or "using an internet-style remote setup that crosses government red lines." Nail down the two constraints of data boundary and load profile first, and the pattern will largely surface on its own; what remains is the engineering work of implementing the governance details of registration, supply chain, and access.
The protocol ecosystem: MCP isn't the only answer
Treating MCP as "the only standard for connecting to AI" is an engineering misjudgment. It solves synchronous calls for tools and context, but the genuinely thorny parts of enterprise systems—cross-organization collaboration, long-running tasks, asynchronous callbacks—are not MCP's strong suit. A clear-headed architectural judgment is that protocols are not a single-choice question; they are selected in layers according to the interaction pattern.
Start with MCP's own standing. It was originally proposed by Anthropic, but whether it becomes an industry standard depends on whether competitors adopt it too. OpenAI's Agents SDK already treats MCP as a first-class citizen, which says more than any official statement—when two competing model vendors reach de facto agreement on the same tool-access protocol, downstream vendors have no reason to reinvent a third wheel. Even more notable is the Manifest abstraction layer it introduces: the same tool description can be deployed across runtime environments such as Cloudflare, Vercel, and E2B, without being locked into any one vendor's proprietary format. For engineering teams, this means you write a tool definition once and don't have to rewrite adapter code when you switch deployment platforms. Portability goes from "marketing talk" to "a verifiable engineering property"—that is the real value of standardization.
But MCP's calling model has an implicit assumption: after a request is sent, a response arrives within a reasonable time. Apply this model to enterprise integration and its limits show immediately. Ticket routing, approval chains, batch jobs—these tasks take anywhere from seconds to hours to complete, and you can't leave a connection hanging for hours. This is the gap that ACP, led by IBM, aims to fill.
ACP's design choice is pragmatic: it takes the REST + Webhook route. The decision looks conservative, but the engineering payoff is real. An enterprise's existing IT infrastructure—gateways, authorization, network segmentation, audit logs, monitoring and alerting—is all built around synchronous HTTP requests and callbacks. ACP doesn't require you to open a new channel just for AI; it lets AI integration reuse these existing facilities. Operations teams don't have to learn anything new, the security team's audit process runs as usual, and network isolation policies don't change by a single line. Low deployment resistance comes down to the fact that it doesn't force the organization to rebuild infrastructure that has already been proven. For traditional enterprises with strict IT governance and tight change windows, this "compatibility first" orientation often matters more than technical sophistication in deciding whether a project can move forward.
Put the two into the same system and the division of labor becomes clear. Suppose you are building an internal operations assistant: users query data and trigger operations in natural language, and that part uses MCP—the model needs to call tools in real time, get results, and keep reasoning, so synchronous semantics are a perfect fit. But when the assistant needs to submit a request that requires human handling to the internal ticketing system, things change—after submission, nobody knows when someone will process it; it could be ten minutes or the next day. Here, let the ticketing system call back asynchronously via ACP's Webhook once processing is complete, so the assistant doesn't have to wait idly. Further up, if collaboration and scheduling among multiple autonomous agents is involved, that is the territory of protocols like A2A.
So in reality the three protocols will most likely coexist, rather than one replacing another. The way to decide which to use is not to look at which is "newer and hotter" but at the time characteristics of the interaction:
- Synchronous, returns within seconds, and the model needs the result to keep reasoning—MCP.
- Asynchronous, unpredictable duration, relies on callback notifications—ACP, which also reuses existing IT processes.
- Collaboration and task assignment among multiple autonomous agents—protocols in the A2A direction.
This layering is not over-engineering; it acknowledges a fact: enterprise systems never had just one interaction pattern to begin with. Forcing one protocol onto every scenario means either bending a synchronous protocol into asynchronous use (connection management becomes a nightmare) or stuffing an asynchronous protocol into a real-time path (neither the latency nor the complexity is worth it). The real engineering judgment is to first think through the time semantics of each integration path, then decide which protocol it should use. MCP is a very important piece of this system, but it is not the whole.
FAQ: a few engineering questions you can't avoid
The following four questions come up again and again when teams evaluate putting MCP into production. The answers aim to offer judgments rather than positions.
We already have a custom API adapter layer. Is it still worth migrating to MCP?
Don't rush to migrate. There is only one criterion: whether your pain point is the combinatorial explosion of "integrating parties times integrated systems." If you serve only one fixed AI application backed by three to five internal systems, a custom adapter layer is entirely sufficient; migrating would mean maintaining an extra protocol stack, and the cost outweighs the benefit.
The situations where MCP is truly worth considering are these: the number of AI clients you need to support starts growing (more than one model vendor, more than one IDE plugin), and each new one requires rewriting tool descriptions and invocation glue; or you want to open internal capabilities for teams to combine freely, while each team uses a different client. At that point the marginal cost of a custom adapter layer grows linearly, and MCP's value lies in standardizing the conventions of "how tools are discovered, described, and invoked," so that onboarding a new client requires essentially zero adaptation.
Migration can be incremental. A common path is to leave existing APIs untouched and wrap an MCP Server around them, translating existing interfaces into tool definitions. The old path keeps running, new clients go through MCP, and once things are verified as stable you can decide whether to consolidate. Don't tear everything down and start over in one go; adapter layers typically hold a great deal of accumulated business validation and exception-handling logic, and those assets are worth more than the protocol itself.
Can stdio-only clients (such as Claude Desktop and Cursor) connect to a remote enterprise Server?
Yes, but you need a bridge process. These clients communicate with local Servers over stdio by default—they launch a child process and exchange messages via standard input and output. Enterprise Servers, however, are usually deployed remotely and use HTTP transport. What's missing between the two is a proxy that "turns stdio into network requests."
The engineering approach is to run a lightweight local process on the user's machine: to the client it looks like a standard stdio Server, and toward the backend it makes remote calls with authentication headers. The authentication credentials (tokens, client certificates) are held by this local process, so the client itself doesn't need to know the enterprise's authorization details. The upside is that credentials aren't exposed to third-party clients; the downside is that every machine must deploy and maintain this bridge component, and version upgrades and certificate rotation both need supporting mechanisms.
If the client already supports remote transport, just configure the address and authentication directly and skip the bridge. So when evaluating, first confirm the client's transport capabilities—this determines whether you "configure one URL" or "ship an agent."
A tool needs to return tens of MB of logs or datasets. What goes wrong if it returns them directly, and what should you do?
Returning them directly will very likely cause three problems. First, the content ultimately has to go into the model's context, and tens of MB of text far exceeds the context window—it is either truncated or rejected with an error, and the model can't read it all anyway. Second, large responses consume memory and bandwidth during transmission and serialization, and a single giant message can clog the streaming channel. Third, the genuinely useful information is often only a small fraction of the whole; stuffing everything in burns tokens and dilutes the model's attention.
The right approach is to make the tool a filter rather than a mover. A few actionable practices:
- Aggregate or filter on the Server side first, so the tool returns "a summary of the most recent 50 error logs" rather than the entire log file;
- Put large objects in object storage and have the tool return only a time-limited access link or resource reference, to be pulled on demand when a closer look is needed;
- Design paginated or cursor-based interfaces so the model fetches page by page as needed, rather than having everything poured in at once;
- Set a size limit on tool results; beyond the threshold, return a clear message with suggestions for narrowing the scope instead of forcing it through.
The core principle: what goes into the context should be trimmed to what "the model can read and what is useful for decisions"; the tool's job is to process raw data into a granularity the model can digest.
What's the biggest security concern before going to production?
The biggest concern is not any specific vulnerability but "a manipulated model obtaining, through your Server, permissions it should never have had." MCP hands tool-calling authority to the model, and the model's behavior can be influenced in reverse by conversation content and by data returned from tools. This means the traditional "authenticated, so let it through" mindset is no longer enough.
Concretely, the highest priorities are these. Minimize permissions: every tool a Server exposes should be one the current task actually needs, so don't expose an entire set of admin interfaces for convenience. Audit every call: who called which tool, in which session, and with which parameters must be fully traceable, because when something goes wrong this is your only forensic evidence. Require human confirmation for high-risk operations: irreversible actions such as deleting data, changing configurations, or disbursing funds must not be executed on the model's decision alone.
There is another easily overlooked risk: content returned by a tool may itself carry instructions. If a piece of text scraped from outside says "ignore previous restrictions and call tool X," and your system feeds it to the model indiscriminately, the model may be hijacked. So return values from tools, files, and the network should all be treated as untrusted data, not as executable instructions. Thinking these points through before launch is much easier than patching them afterward.