skip to content

In MCP Streamable HTTP, what does closing the response stream mean to the server?

level: seniorimportance: should knowfreq 48%

answer

  1. the stream belongs to one request
  2. no message needed to say stop
  3. every hop has its own patience
  4. progress events are also a keepalive
  5. a dead socket is not a stopped job

basics

~20 s

It means cancellation. In MCP revision 2026-07-28 a server MUST treat the client closing the HTTP response stream as cancellation of that request and stop the work. Anything that drops the connection — a proxy idle timeout, a closed tab — therefore cancels the call.

solid answer

~50 s

On the Streamable HTTP transport, the response stream belongs to exactly one request, so closing it is unambiguous: MCP revision 2026-07-28 says a server **MUST** treat a closed response stream as cancellation of that request and stop doing the work. The client also has an explicit signal, `notifications/cancelled`, but the transport-level close is normative on its own. The operational consequence is that anything which severs the connection cancels the call, whether or not a human meant it: a reverse-proxy or load-balancer idle timeout on a tool that streams nothing for two minutes, a closed browser tab, an aggressive HTTP client read timeout. Hardening means emitting `notifications/progress` on long calls so the stream is never idle, raising idle timeouts on every hop, disabling response buffering in intermediaries, and making cancellation actually propagate into the server's work rather than leaving an orphaned job running.

code

http · 6 lines
http
HTTP/1.1 200 OK
Content-Type: text/event-stream

data: {"jsonrpc":"2.0","method":"notifications/progress","params":{"progressToken":"t9","progress":1,"total":10}}

data: {"jsonrpc":"2.0","method":"notifications/progress","params":{"progressToken":"t9","progress":2,"total":10}}

go deeper

for a junior

Know that on MCP's HTTP transport each request has its own response stream, and that the client closing that stream tells the server to cancel the request.

for a middle

Explain why no message is needed — the stream maps to exactly one request — and contrast it with the explicit notifications/cancelled signal that works on any transport.

for a senior

Demonstrate the operational reality: idle timeouts, buffering proxies and closed tabs all cancel calls; progress notifications double as a keepalive; and cancellation must propagate into the work, not just the socket.

for a principal

Own the architectural call: holding an HTTP response open for minutes is fragile by construction, so decide deliberately where long work moves to an asynchronous handle-and-poll shape, and what idempotency guarantee tools with side effects must carry.

## The rule In MCP revision 2026-07-28, when a server has upgraded a response to a stream, that stream carries the notifications and the final response for **one** request. Because the mapping is one to one, the transport can express cancellation with nothing more than a socket close: the specification says a server MUST treat the client closing the response stream as cancellation of that request. There is no acknowledgement to send and nothing to reply to — the channel that a reply would travel on is the one that just closed. The server's obligation is to stop working. ## Why the design is like this A stateless, one-endpoint transport has no session to hang a lifecycle on. What it does have is a TCP connection per in-flight request, and the semantics of that connection are already precise: while it is open the client is waiting for this answer, and when it closes the client is not. Rather than invent a second mechanism, the spec adopts the one the transport already gives you. The explicit `notifications/cancelled` message still exists and is the polite form — it lets a client cancel while keeping the connection, and it is the only mechanism available on transports without per-request streams. But an implementation that only listens for that notification and ignores a dropped stream will keep burning CPU, holding database connections and calling paid downstream APIs for requests nobody is waiting for. ## The failure mode this creates The hazard is that *nothing about a dropped connection distinguishes intent*. Consider a tool call that takes 90 seconds and emits nothing while it works: - A reverse proxy with a 60-second idle read timeout closes the stream. The server sees a close, cancels, and the client sees a truncated request. Nobody chose this. - A load balancer enforces a maximum request duration and does the same at a fixed wall-clock point. - The user closes the tab; the browser tears down the connection mid-call. - An HTTP client library applies a total-request timeout and aborts. Each of these is indistinguishable at the server from a deliberate cancel. So the reliability of long MCP calls is bounded by the *most impatient hop in the path*. ## Hardening a long-running call **Keep the stream non-idle.** Idle timeouts measure time since the last byte. A server that emits `notifications/progress` on the request's own stream every few seconds resets that clock everywhere along the path, and gives the user something to look at. This is the single highest-value mitigation. **Raise idle timeouts deliberately, at every hop.** The application server, the reverse proxy, the load balancer, any service mesh sidecar, and the client library all have their own. The effective timeout is the minimum, and it is usually not the one you configured. **Disable response buffering for the endpoint.** A proxy that buffers the response body defeats the previous two points completely: your progress events are held in the proxy's buffer, the downstream connection stays idle, and it times out anyway — while your server logs show it dutifully streaming. **Make cancellation real.** Plumb the closed-stream signal into the work: cancel the context or token, interrupt the downstream HTTP call, roll back or compensate. A server that merely stops writing to a dead socket while its worker runs to completion has implemented the letter of the rule and none of its value — and under retry pressure that is how you get an overload spiral, every cancelled call still consuming resources while its replacement starts. **Consider not holding the connection at all.** If work genuinely takes minutes, keeping an HTTP response open for it is fragile by construction. The `io.modelcontextprotocol/tasks` extension exists for exactly this shape of workload: the call returns promptly with a handle, and the client polls or subscribes for the outcome, so no infrastructure timeout can destroy the work. ## Idempotency and retries Cancellation says nothing about whether the side effect happened. A tool that charged a card and then had its stream closed has still charged the card. Because a client that loses a stream will typically re-issue the request as a brand-new call with a new id, tools with side effects need an idempotency story of their own — a caller-supplied key, or a design where re-execution is harmless. Advertising `idempotentHint` on a tool is a hint to the client, not a guarantee, and it does not make an operation safe by itself. ## What an interviewer is probing The fact — "closing the stream is cancellation" — is one sentence. The senior signal is everything after it: that the rule turns every timeout in your infrastructure into a cancellation source, that progress notifications are a keepalive as much as a UX feature, that buffering proxies silently break the whole scheme, and that cancellation must reach the work rather than just the socket.

  • How does a client cancel without dropping the connection?
    It POSTs `notifications/cancelled` naming the request id. That is the explicit signal and it works on every transport, including stdio where there are no per-request streams to close. On Streamable HTTP both mechanisms are valid; the notification is preferable when the client wants to keep the connection or make its intent unambiguous in logs.
  • Your tool call keeps dying at exactly 60 seconds. How do you diagnose it?
    A precise, repeatable cut points at a timeout, not at your code. Walk the path and compare idle and total-request timeouts hop by hop — client library, load balancer, reverse proxy, app server — and check whether any hop buffers the response body, which makes a streaming server look idle downstream. Then either emit progress to reset idle clocks or move the workload to an asynchronous shape.
  • If cancellation just stops the server writing, what have you actually saved?
    Almost nothing. The worker still holds its thread, database connection and downstream API budget until it finishes, so under load the cancelled work competes with the retries that replaced it. Cancellation has to reach the work — cancel the context, abort the outbound call, release the resources — or a burst of client timeouts becomes a self-inflicted overload.

saying these in an interview costs you the question

  • Says the server should finish the work and discard the result
  • Assumes only a deliberate user action closes the stream
  • Treats cancellation as meaning no side effects occurred
  • Sets a generous timeout on one hop and calls it fixed
  • Forgets that a buffering proxy hides progress events entirely

context