Why must an MCP server re-read capabilities from every request rather than caching them?
answer
- Self-contained requests, no memory between them
- A pipe is not a conversation
- Same client, different call, different declaration
- Any node may serve any request
- MUST NOT infer capabilities from prior requests
basics
~20 sMCP revision 2026-07-28 makes each request self-contained, so a server MUST NOT infer capabilities from prior requests. A cached capability set can be stale, belong to a different caller, or be unavailable to the node now handling the call.
solid answer
~50 sThe specification states that MCP is a stateless protocol and that servers MUST NOT rely on prior requests over the same connection to establish context — an open connection, including a long-lived stdio process, is explicitly not a conversation. Concretely, a server must read `io.modelcontextprotocol/clientCapabilities` out of the `_meta` of the request it is currently serving and make its decision from that alone. Caching breaks in three ways: the declaration can legitimately differ per request, because a client may offer interactive capabilities on a user-facing call and none on a background one; a horizontally scaled deployment routes requests to nodes that never saw the earlier one; and cached capability state is a correctness trap after any reconnect. If the current declaration is insufficient the server answers `-32021` `MissingRequiredClientCapabilityError`; it never quietly falls back to what a previous request claimed.
go deeper
Remember the rule as stated: capabilities arrive on every request, and a server reads them from the request it is handling right now rather than from anything it saw before.
Explain why this is a correctness rule and not an optimisation — the same client may declare differently per call, and absence of a capability key is an error rather than a repeat of the last one.
Demonstrate it in system terms: no affinity, no session store, request-scoped capability state in the handler chain, and -32021 rather than a silent fallback when the present declaration falls short.
Own the tradeoff you are buying: metadata on every call in exchange for restartable, affinity-free servers. Be able to argue where that stops paying off and what you would measure before adding any caching layer above the protocol.
## The rule Revision 2026-07-28 removed the `initialize` handshake and, with it, the idea that capabilities are established once and then remembered. The replacement rule is stated as a MUST NOT: servers must not infer capabilities from prior requests. The capability object arrives in `params._meta` under `io.modelcontextprotocol/clientCapabilities` on every single request, and the server evaluates the request in front of it using only that object. The specification's framing is worth quoting because candidates often soften it: "MCP is a stateless protocol: every request is self-contained and carries its own protocol version and capabilities," and "an open connection, such as a STDIO process, is not a conversation or session." The stdio clause exists precisely because a subprocess *feels* like a session — one long-lived pipe, one peer at the other end — and implementers are tempted to hang state off it. ## Why caching is wrong, not merely unnecessary ### The declaration can legitimately change Capabilities are a property of the request, not of the client as an organism. The same client may declare interactive capabilities on a call it is making on behalf of a user sitting at a screen, and declare none at all on a scheduled background call where nobody can answer a question. A server that remembered the first declaration would happily start an interaction the second call cannot complete. Nothing in the protocol forbids the variation; the client is behaving correctly and the caching server is wrong. ### The next request may land elsewhere A remote MCP deployment is ordinary HTTP infrastructure. Requests are POSTs, and any node can serve any of them. A capability set cached in one process's memory is invisible to its peers, so a cache-dependent server works in single-process testing and fails intermittently in production — the worst possible failure signature. ### There is nothing to hang the cache on Protocol-level sessions and their identifying header were removed in 2026-07-28, so a server has no protocol-blessed key under which to file remembered capabilities. Keying off a transport connection re-introduces exactly the connection-as-session assumption the revision deleted, and keying off self-reported client identity is worse, because that identity is untrusted. ## What a correct server does instead Validate `_meta` in one place, before dispatch. Parse the capability object into a value that travels with the request through your handler chain — a request-scoped context object, not a field on a connection or a process-wide map. Every decision that depends on what the client can do reads from that value. When the current request's declaration is insufficient for the work, the server returns `-32021` `MissingRequiredClientCapabilityError` with `data.requiredCapabilities` naming what was needed. It does not consult history, does not assume the client "probably still" supports what it declared a minute ago, and does not degrade silently. If the client does support the capability, it re-issues the call with the declaration corrected. Equally, a server must not treat an absent `clientCapabilities` key as an empty object or as a repeat of the last one. Absence is invalid params: `-32602`, delivered with HTTP 400 on Streamable HTTP. ## The client side of the same rule Clients carry a matching obligation: build `_meta` for every outgoing request rather than once at start-up. In practice the reliable pattern is to stamp it in the transport layer so that no individual call site can forget, and to derive the capability object from the *situation* — is a human present, is this interactive, is this a batch job — rather than from a single constant. ## What this costs and what it buys The cost is real: a few hundred bytes of metadata on every call, and a small amount of validation work per request. The payoff is that a server needs no affinity, no session store, no expiry sweeper, and no reconnect logic to rebuild negotiated state. A process can be restarted mid-flight and the next request is served normally. For a protocol whose servers are frequently small, short-lived, and deployed behind ordinary load balancers, that trade is deliberate. ## How this shows up in interviews The question separates candidates who read the current specification from those recalling the handshake era. Weak answers say caching is "an optimisation" or "fine within one connection". A strong answer names the MUST NOT, gives at least one concrete way a cache produces a wrong result rather than merely a stale one, and points at `-32021` as the sanctioned response when the present declaration falls short.
- A client declared elicitation support on one call and omits it on the next. What must the server assume?That the current request has no elicitation support, and it must be served without it — or refused with `-32021` if it cannot be. The variation is legal: capabilities describe the request, not the client, and a client may reasonably suppress interactive capabilities on background work where no human can answer.
- Doesn't a long-lived stdio process make caching safe, since there is exactly one peer?No — the specification singles this case out: an open connection, such as a stdio process, is not a conversation or session. The peer being constant does not make its per-request declarations constant, and the MUST NOT is unconditional across transports. Read capabilities from the request being served.
- Where should a server keep the parsed capability object in its own code?In request-scoped state that travels with the handler chain, never on the connection object or in a process-wide map keyed by connection. That placement makes the correct behaviour structural: there is simply no field in which a previous request's declaration could survive to influence the next one.
saying these in an interview costs you the question
- Says caching capabilities per connection is a safe optimisation
- Treats a long-lived stdio process as a session
- Assumes a client's capabilities never change between calls
- Falls back to the last known declaration instead of erroring
- Keys remembered capabilities off untrusted clientInfo