skip to content

How do you detect a stalled Anthropic SSE stream, and what timeout should you set?

level: seniorimportance: should knowfreq 38%

answer

  1. Idle gap, not total duration
  2. Pings exist to prove liveness
  3. Reset the watchdog on every frame
  4. Read timeouts are already per-chunk
  5. Buffering proxies fake a stall

basics

~20 s

Time the gap between events, not the whole request. Reset a watchdog on every frame including ping keepalives, and fail the stream when the idle gap exceeds your budget. A total wall-clock timeout kills legitimate long generations instead.

solid answer

~50 s

A streaming Messages call can legitimately run for many minutes, so a total request timeout is the wrong instrument — set it low and you cancel good generations, set it high and a wedged connection hangs for that long. The right control is an inter-event idle timeout: keep a deadline that every received frame resets, including `ping` events, which exist precisely as keepalives so that a healthy but slow stream still produces traffic. HTTP clients help here, because a read timeout is applied per chunk rather than to the whole response, so it already behaves like an idle timeout when you consume a stream — but if you relay events onward, you need your own watchdog in that layer too. Independently, verify completion: a stream that ends without `message_stop` was truncated regardless of timing. And check your infrastructure — a proxy that buffers responses will hold frames back and make a healthy stream look stalled.

go deeper

for a junior

Know that ping events are empty keepalives you skip when rendering, and that a stream is only finished when message_stop arrives.

for a middle

Explain why an inter-event idle timeout beats a total-request timeout for streaming, and that HTTP read timeouts already apply per chunk when you consume a response incrementally.

for a senior

Show the operational picture: a watchdog reset by every frame including pings, separate liveness and completion checks, buffering proxies as the first suspect, and stalls tracked apart from clean errors.

for a principal

Own the end-to-end latency contract across every hop — origin, gateway, relay, browser — deciding where liveness is enforced, what the client sees on a stall, and how timeout budgets are derived from measured tails rather than defaults.

## Why the obvious timeout is the wrong one Every HTTP client you have used defaults to a total-request timeout, and for ordinary JSON APIs that is exactly right. Streaming breaks the assumption behind it. A long generation can legitimately keep the response open for minutes; Anthropic's own guidance is to stream long requests rather than wait on a buffered response. So a 30-second total timeout does not protect you from anything — it just cancels valid work — while a 10-minute one lets a genuinely wedged socket hold a request slot for ten minutes. The useful question is not "how long has this run?" but "how long since anything happened?". A healthy stream produces frames continuously: content deltas while the model generates, and `ping` events, which carry no payload and exist so a stream that is briefly quiet still shows signs of life. Silence for longer than a few tens of seconds is anomalous; five minutes of steady deltas is not. ## Building the watchdog Keep a deadline that is reset on **every** received frame, not only on ones you find interesting. This is where implementations go wrong: a relay that resets its timer only when it forwards visible text will consider a stream dead during a stretch of pings or of non-text blocks. Skip unknown and uninteresting events for *processing* purposes, but count them as liveness. HTTP clients give you part of this for free. A read timeout is measured per read rather than across the whole response, so when you consume a stream chunk by chunk it already behaves as an idle timeout — the SDK's timeout setting flows down to that behaviour. What it cannot cover is the layers you add: if you buffer events, hand them to a queue, or re-emit them into your own SSE or WebSocket stream to a browser, each hop needs its own liveness notion, and the browser side needs one too, since a healthy backend that stops producing looks identical to a network partition from the client's seat. ## Completion is a separate check from liveness Timeouts answer "is it still alive?". They do not answer "did it finish?". Those are different failures and both must be handled. A stream that terminates without `message_stop` was truncated, no matter how promptly its bytes arrived; treat the absence of the terminator as an error in the accumulator itself, rather than trusting callers to notice. Conversely, a stream that is slow but eventually delivers `message_delta` and `message_stop` succeeded, and any timeout that killed it was misconfigured. ## Infrastructure that fakes a stall Before blaming the provider, check your own path. Reverse proxies, CDNs, API gateways and some serverless platforms buffer responses by default, holding output until a size threshold or the response ends. Behind such a hop, a perfectly healthy stream arrives as one lump at the end — which destroys the entire point of streaming and reads, from the client, as a long stall followed by a burst. The fix is per-platform: disable response buffering on that route, or use whichever mechanism your proxy honours to mark the response as unbuffered. Compression middleware that buffers to compress causes the same symptom. Any "streaming is broken in production but fine locally" report should send you here first. ## Reasonable settings There is no universal number, but the reasoning is portable. Pick an idle budget comfortably above the largest normal inter-event gap you observe — the interesting gap is usually the one before the *first* content event, since the model may think before it speaks, and that gap grows when extended thinking or a large cached prefix is involved. Instrument time-to-first-event and the maximum inter-event gap as histograms, then set the budget from the tail rather than from intuition. Keep a generous ceiling on total duration as a backstop against pathological cases, but do not make it your primary control. ## What to do on a stall A stalled stream is a mid-flight failure like any other: you may already have shown text to the user. Cancel the request to release the connection, emit an explicit failure into your own client protocol rather than just hanging up, and decide deliberately whether to retry — remembering that a retry regenerates from the start and re-bills the input. Track stalls separately from clean errors; a rising stall rate usually points at your own network path or a newly introduced buffering hop, not at the model.

  • What are ping events for, and should your application logic react to them?
    They are keepalives with no content, sent so that a stream which is momentarily producing nothing still shows traffic. Your rendering and accumulation logic should ignore them entirely — but your liveness watchdog must count them, because treating only content events as signs of life makes a healthy quiet stream look stalled and gets it cancelled.
  • Streaming works locally but arrives as one lump in production. What do you check first?
    Response buffering somewhere on the path — a reverse proxy, CDN, API gateway or compression middleware that holds output until a threshold or until the response ends. That converts a stream into a delayed single payload and looks exactly like a long stall followed by a burst. Disable buffering on that route, then re-test end to end rather than only against the origin.
  • How do you pick the idle timeout value rather than guessing it?
    Instrument two histograms: time to first event and the maximum inter-event gap per request. Set the budget above the observed tail, not the median, and re-check after prompt or configuration changes, since a large cached prefix or extended thinking lengthens the pre-first-token gap. Keep a generous total-duration ceiling only as a backstop against pathological cases.
  • Is a slow stream that eventually completes a failure?
    No. Liveness and completion are different checks. If message_delta and message_stop arrive, the message is complete and correct however long it took; the only question is whether your latency budget is acceptable to the product. The genuine failures are silence beyond the idle budget, and a stream that ends without the terminator regardless of how fast its bytes arrived.

saying these in an interview costs you the question

  • Using a short total-request timeout for streaming calls
  • Resetting the watchdog only on visible text events
  • Treating ping events as errors or as content
  • Assuming a completed-looking stream needs no message_stop check
  • Blaming the provider before checking proxy response buffering

context