What does MCP 2026-07-28 statelessness change for a remote server behind a load balancer?
answer
- nothing to route on any more
- default routing is now fine
- one request, one node, still true
- watch the idle timeout on streams
- rolling deploys can split the catalogue
basics
~20 sSession affinity stops being required: with sessions removed, any node can serve any request, so ordinary round-robin routing works. What remains is that long-lived response streams still occupy one node, and any state a tool mints must live in shared storage.
solid answer
~50 sRemoving protocol sessions in revision 2026-07-28 turned an MCP endpoint into an ordinary stateless POST endpoint. There is no `Mcp-Session-Id` to route on and no per-connection context to preserve, so you need no sticky sessions, no session store to keep nodes in agreement, and no drain step that waits for sessions to end. Three things still need attention. **Long-lived streams**: a request whose response is streamed, and a `subscriptions/listen` stream in particular, occupies one node for its lifetime, so idle timeouts, response buffering and drain windows in the balancer matter. **Shared state**: anything a tool mints as a handle must be resolvable from every node, or you have re-created affinity invisibly. **Fleet homogeneity**: the spec says a server's tool, prompt and resource set MUST NOT vary per connection (it MAY vary by the authorization presented), so a mixed-version rolling deploy can show one client two different tool lists.
go deeper
Know the headline: because 2026-07-28 has no sessions, an MCP server behind a load balancer does not need sticky routing — any node can answer any request.
Be able to say what the affinity was for in the first place — the Mcp-Session-Id and the connection-scoped context behind it — and why self-contained requests remove the need for it.
Show the operational detail: idle timeouts and response buffering for streamed responses and subscriptions/listen, drain windows sized to in-flight work, shared storage for handle-backed state, and clients re-issuing lost calls with new ids.
Own the rollout policy — catalogue and era changes as fleet-wide flips rather than gradual rollouts, per-request token validation as the identity story, and the capacity question of how many concurrent long-lived streams a node can hold.
## What removing sessions actually bought Under revisions up to 2025-11-25, a remote MCP server could mint an `Mcp-Session-Id` after `initialize` and require it on every later request. Operationally that is a sticky session: the node holding the session's context had to receive the follow-ups, so you configured affinity on the balancer, or you externalised session context into a shared store and paid a lookup on every request, or you accepted that a client would periodically get a 404 meaning its session had expired and have to re-handshake. Revision 2026-07-28 removed the whole layer. Every request is self-contained, carrying its own protocol version and capabilities in `params._meta`, and servers MUST NOT rely on prior requests over the same connection. The consequences on the infrastructure side are direct: - **No affinity.** Round-robin, least-connections, whatever your balancer does by default is fine; any node can serve any POST. - **No session store to keep nodes agreeing** about who has initialized and with what capabilities. - **Scale-in is cheap.** Terminating a node costs the requests currently in flight on it, not a population of sessions. - **Rolling deploys are ordinary.** There is no session lifetime to drain past; you drain in-flight requests. - **Cold start is one round trip.** A one-shot `tools/call` needs no handshake first, which is what makes serverless and edge deployments practical. ## What still pins work to a node Statelessness is about the protocol, not about TCP. A single request still lives on one connection to one node from start to finish, and MCP has requests that can live a long time: - A `tools/call` whose response the server upgrades to a stream stays on that node until it completes. Losing the node loses the request, and the client must re-issue it as a **new request with a new JSON-RPC id** — 2026-07-28 removed resumability, so there is no replay. - `subscriptions/listen` is deliberately one long-lived POST whose response stream carries change notifications. That stream can outlive many ordinary requests, so it is the thing most likely to be killed by an idle timeout or a deploy. Practical implications: raise the balancer's idle and read timeouts on the MCP path, ensure the proxy does not buffer streamed responses (buffering turns a live stream into a hang), give the balancer a drain window long enough for in-flight work, and accept that clients will reconnect and re-establish listens. Note the client-side rule that a client MUST re-send `subscriptions/listen` after a stdio reconnect; the same practical need exists whenever a stream is lost over HTTP, because nothing about the previous listen survives. ## The state that remains, and where it must live Cross-call state now takes the form of handles a tool mints and the client passes back as ordinary tool arguments. The handle carries no routing information, so its backing state must sit in storage every node can reach — a database, a cache cluster, object storage. Keeping it in process memory is the classic failure: it works in a single-node staging environment and produces intermittent unknown-handle errors in production, because it is affinity that the load balancer does not know about and therefore cannot honour. ## Fleet homogeneity and rollouts Two spec rules bite during a rolling deploy. The tool, prompt and resource set MUST NOT vary per connection, although it MAY vary by the authorization presented — so if version N and N+1 of your server expose different tools, a client hitting different nodes sees an unstable catalogue and any cached listing goes stale unpredictably. And era compatibility is a property of the **server**: if half your fleet is dual-era and half is modern-only, a legacy client's behaviour depends on which node it lands on. Treat catalogue changes and era changes as fleet-wide flag flips rather than as things that ride along with a gradual rollout. ## Auth, not sessions With no session to carry identity, every request stands on its own credential. For remote servers that is OAuth 2.1 with audience-bound tokens and no token passthrough; the MCP server is an OAuth resource server. That is good news for a fleet — validating a token is stateless work any node can do — but it also means authorization is per request, including for calls that present a handle. Possession of a handle MUST NOT be treated as authentication. ## How to answer Lead with what you no longer need (affinity, session store, session drain), then show you know what is left (long-lived streams, shared handle storage, catalogue consistency across versions, per-request auth). Naming the load-balancer specifics — idle timeouts and response buffering — is what separates someone who has run this from someone who has read the spec.
- You are draining a node for a deploy. What actually needs to finish before you kill it?Only in-flight requests, since there are no sessions to expire. In practice that means ordinary calls, any `tools/call` currently streaming a response, and open `subscriptions/listen` streams. Stop accepting new POSTs, allow a grace window sized to your slowest tool, then terminate: clients recover by re-issuing lost calls with new ids and re-establishing their listen streams.
- Your streamed tool responses hang for exactly 60 seconds and then deliver everything at once. What would you look at first?Response buffering in the proxy or balancer in front of the MCP endpoint. A proxy that buffers the response body defeats an SSE-upgraded stream: nothing reaches the client until the buffer flushes or the request ends. Disable buffering on that route, then check idle and read timeouts, since a long-lived stream with sparse traffic is otherwise a prime candidate for being cut.
- Half your fleet runs a build exposing a new tool and half does not. What breaks?Catalogue consistency. A client calling `tools/list` twice can get two different answers depending on routing, and any cached listing (results can carry `ttlMs` and `cacheScope`) may describe tools the node it next reaches does not implement, producing method or name errors. The rule is that the tool set MUST NOT vary per connection, so gate catalogue changes behind a fleet-wide flag flipped after the rollout completes.
- Does statelessness let you run an MCP server on a serverless platform?Yes, and that is much of the point: an ordinary `tools/call` is a single self-contained POST with no handshake to warm up, so a function invocation can serve it. The constraint is the long-lived pieces — a `subscriptions/listen` stream and any slow streamed tool run up against per-invocation duration limits — and any handle-backed state must go to an external store rather than to instance memory.
saying these in an interview costs you the question
- Configures sticky sessions on the MCP endpoint out of habit
- Keeps handle state in node memory behind a load balancer
- Assumes a dropped stream will be resumed after failover
- Rolls out a changed tool catalogue node by node
- Believes statelessness means no request can be long-lived