skip to content

How do you decide what belongs in the MCP host versus in an MCP server?

level: principalimportance: should knowfreq 32%

answer

  1. place it by what it can see
  2. servers see only their own requests
  3. would another host want this server?
  4. the model and the user stay host-side
  5. 2026-07-28 narrowed the server's reach

basics

~20 s

Put anything needing the model, the user, or several integrations at once in the host; put one system's capabilities and its credentials in a server. A server sees only its own requests, so anything requiring wider context cannot live there.

solid answer

~50 s

Start from what each role can actually see. A server sees the requests its own client sends and nothing else — not the conversation, not the user's screen, not another server — so any behaviour that needs those must sit in the host. The host owns the model, the user, consent, and all cross-server orchestration. A server owns one integration: its domain operations as tools, its data as resources, its reusable prompt templates, and the credentials to reach the backing system. Revision 2026-07-28 pushes this split harder by deprecating the mechanisms that let servers borrow host capability — sampling (migrate to calling LLM provider APIs directly), roots (pass paths as tool parameters, resource URIs or configuration) and logging (use `stderr` on stdio, OpenTelemetry for observability). The test that decides most cases: would this server be useful, unchanged, to a different host?

go deeper

for a junior

Know the simple version: the host has the model and the user, a server wraps one integration. If something needs the conversation or an approval prompt, it is not the server's job.

for a middle

Be able to justify a placement from what each side can see, and name what a server cannot see at all — the conversation, the user's screen, and other servers.

for a senior

Bring the operational tests: which failures should be contained to one integration, where the credentials live, and why a design that needs per-conversation server state conflicts with the 2026-07-28 statelessness rule.

for a principal

Own the boundary as an architectural decision — portability across hosts, the server as the host's unit of trust and containment, and the read that 2026-07-28's deprecations of sampling, roots and logging deliberately narrow what a server may reach for.

## Decide by visibility, not by convenience The cleanest way to place a responsibility is to ask what it needs to see. A server sees the requests its own client sends it. That is all. It cannot read the conversation, cannot see the user's interface, and cannot see or call the other servers the same host is connected to. A host sees everything: the model, the user, the full conversation, and every server it wired up. So the rule falls out directly. Anything that needs the conversation, the user's attention, model inference, or the combined output of two integrations is host work. Anything that is one system's domain capability plus the credentials to reach it is server work. ## What belongs in the host - **The model.** The host runs inference. Revision 2026-07-28 deprecated sampling — the mechanism by which a server could ask the client to run a completion — with the stated migration being that servers integrate directly with LLM provider APIs. Deprecated features remain for at least twelve months, but a new server should not be architected around borrowing the host's model. - **The user.** Consent, approval and any interface are the host's. The protocol cannot enforce its principles; the host is the enforcement boundary. - **Aggregation and routing.** Merging tool, resource and prompt lists from several servers for the model, keeping provenance, and routing a chosen call back to the right client. - **Cross-server orchestration.** Feeding one server's output into another's input is host logic, because neither server can see the other. - **Trust policy.** `ToolAnnotations` (`readOnlyHint`, `destructiveHint`, `idempotentHint`, `openWorldHint`) and `io.modelcontextprotocol/serverInfo` are hints and self-reported respectively, and must be treated as untrusted unless the server is trusted. Deciding which servers are trusted is the host's call. ## What belongs in a server - **One integration's operations**, exposed as tools with real input schemas, its data as resources, and its reusable instructions as prompts. - **The credentials for its own backing system.** A remote server is its own OAuth 2.1 resource server with audience-bound tokens, and passthrough of a token to anything else is prohibited. - **Domain knowledge the host cannot be expected to have** — how to page an API, which fields matter, what a safe default filter is. Encoding that in the server is what makes it worth writing, because every host then gets it for free. ## Sharpening tests **Portability.** Would this server be useful, unchanged, to a completely different host? If it only makes sense inside one product — it assumes a particular UI, a particular prompt, a particular agent loop — the logic has leaked out of the integration and belongs in the host. **Statelessness.** Revision 2026-07-28 removed protocol sessions and states that every request is self-contained and that an open connection, including a stdio process, is not a conversation. A design that needs the server to be a stateful backend for one conversation is fighting the protocol; continuity must be explicit in what each request carries. **Granularity of the trust unit.** The host's isolation unit is the server: one client, one credential, one trust decision, one failure domain. If you fold six systems of different sensitivity behind one server, the host can no longer distinguish them, and your server has silently taken over the containment the topology was providing. **Who suffers the failure.** If the thing failing should degrade one integration, it is server-side. If it should degrade the whole application, it is host-side. ## The deprecations as a design signal The 2026-07-28 revision reads, in places, as a deliberate narrowing of the server's role. Sampling, roots and logging were all deprecated together, and all three were ways a server reached back into host capability: to the model, to the client's filesystem view, and to the host's log sink. The migrations point outward instead — call an LLM provider directly, take paths as tool parameters or resource URIs or configuration, write to `stderr` and use OpenTelemetry. Server-initiated JSON-RPC requests were removed entirely in the same revision; a server that needs something now returns an `InputRequiredResult` and waits to see whether the client retries. The direction of travel is unmistakable: servers offer capability and answer requests; hosts hold everything that is about the user, the model, or the whole picture. ## Common bad splits A "gateway" server that fronts many unrelated systems collapses the host's isolation to one unit. A server that embeds a full agent loop duplicates the host's job and cannot see the conversation it is trying to reason about. A server designed as per-conversation state storage assumes sessions that no longer exist. And a server that needs privileged knowledge of one host's configuration is not really an MCP server — it is a plugin for that product wearing the protocol's clothes.

  • A team wants their server to summarise results with an LLM before returning them. Where should that run?
    Either in the host, which already holds a model and the conversation, or in the server against its own LLM provider account. What it should not do is borrow the host's model through sampling: that mechanism was deprecated in revision 2026-07-28, with the stated migration being direct integration with provider APIs. If the summary needs conversation context, it is host work by definition, since the server cannot see the conversation.
  • Is one big server exposing forty tools across six systems a good split?
    Usually not. The host's isolation unit is the server — one client, one credential, one trust decision, one failure domain — so folding six systems together removes the host's ability to enable, approve or contain them separately, and the containment burden moves inside your server. It can be right when the six systems share a trust level and an owner, and the aggregation genuinely simplifies the model's view.
  • How do you tell that logic has leaked from the host into a server?
    Ask whether a different host would want the server as-is. Signs of leakage are a server that hardcodes prompt wording for one product, that implements its own agent loop, that assumes a particular approval UI, or that expects to remember a conversation across calls — which the 2026-07-28 statelessness rule rules out anyway. Each of those is host territory wearing a server's clothes.

saying these in an interview costs you the question

  • Putting cross-server orchestration inside one of the servers
  • Designing a server that needs the conversation to work
  • Building new servers around sampling to borrow the host's model
  • Treating one server as per-conversation state storage
  • Assuming a server can enforce consent without the host

context