How do you handle an Anthropic SSE error event that arrives mid-stream?
answer
- 200 already sent, headers already flushed
- Failure travels inside the stream
- No resume, no cursor, no offset
- Retry re-bills the input tokens
- Missing message_stop is also a failure
basics
~20 sTreat it as a failed request even though bytes already reached the user. The HTTP 200 and headers are long gone, so surface the failure in your own stream protocol, keep or discard the partial text deliberately, and retry the whole call — there is no resume point.
solid answer
~50 sA streamed Messages call returns `200` the moment the stream opens, so a failure that happens later cannot be expressed as an HTTP status. Instead an `error` event appears in the stream — for example `{"type": "error", "error": {"type": "overloaded_error", "message": "Overloaded"}}` — and the SDKs raise it out of the event iterator. Two things follow. First, your own transport needs a way to say "this failed after it started": if you relay to a browser, define an explicit error frame in your protocol rather than closing the socket, because a silent close is indistinguishable from a short answer. Second, there is no resume — retrying re-sends the whole request, regenerates from scratch and bills the input again (cache reads soften that). For transient types such as `overloaded_error` retry with exponential backoff and jitter; for a request-shaped error, retrying is pointless. A stream that simply dies without `message_stop` is the same class of failure and needs the same handling.
go deeper
Know that a streaming call returns 200 immediately, so an error later arrives as an event inside the stream and the SDK raises it out of the loop you are iterating.
Explain why there is no HTTP status left to change, which error types are worth retrying, and that a retry regenerates from scratch and re-bills the input rather than resuming.
Show the operational plan: an explicit error frame in your own protocol, a deliberate decision about partial output, backoff with jitter and attempt caps, separate metrics for mid-stream versus pre-stream failure.
Own the user-visible contract and the cost of failure — when auto-retry is acceptable versus when replacing visible text is worse than an error, and how idempotency is guaranteed for effects triggered from a stream that may be replayed.
## Why this question is asked at senior level Non-streamed failures are easy: you get a 4xx or 5xx, you branch on it. Streaming inverts that. The status line and headers are flushed when the first byte leaves, which means the request has already "succeeded" from HTTP's point of view before the model has generated anything. Every failure after that point has to be carried *inside* the stream. Candidates who have only built demos have never hit this; candidates who have shipped a chat UI have, and it shows. ## The two failure shapes **An `error` event.** The stream carries a frame whose payload is an error object with a `type` and a `message`. `overloaded_error` (capacity pressure) and `api_error` are the ones you will see mid-stream; request-shaped problems such as `invalid_request_error` and `authentication_error` normally fail before the stream opens. In the SDKs this is surfaced as an exception thrown from the iterator or streaming context manager, so your `for event in stream` loop needs a `try` around it, not just around the call that opened it. **A dead connection.** The socket closes, or goes silent forever, with no `error` event and no `message_stop`. Because you may already hold plausible-looking text, the only reliable detector is the terminator: if you never saw `message_stop`, the message is incomplete regardless of how good the prefix looks. Bake that assertion into the accumulator rather than leaving it to the caller. ## What you do about it **Decide the fate of the partial output explicitly.** There are only three honest choices, and the right one is product-dependent: discard it and show a failure; keep it on screen but mark it visibly as interrupted; or keep it and offer a continue action. What you must not do is leave partial text on screen styled exactly like a completed answer — that is how users end up quoting a sentence the model never finished. **Retry the whole request, with eyes open.** There is no cursor, offset or resume token; a retry is a fresh request that regenerates from the beginning and re-bills the input tokens (prompt caching reduces the cost of that re-send, not the fact of it). Use exponential backoff with jitter, cap the attempts, and only retry error types that are actually transient. If regeneration is expensive and you have a long partial answer, an alternative is to start a new request whose last message is an assistant turn containing the text you already have, so the model continues rather than restarts — a product decision with its own risks (the seam can be awkward, and it changes the prompt), not a protocol feature. **Make your own protocol able to say "failed".** If your backend relays to a browser, do not forward Anthropic's frames verbatim and do not signal failure by hanging up. Emit your own typed error frame, so the client can distinguish an interrupted generation from a completed one and can decide whether to offer retry. This is also where you strip provider details you do not want to leak and attach your own request id for support. **Observe it.** Count mid-stream failures separately from pre-stream failures. They have different causes and very different user impact: a pre-stream 529 is invisible behind a retry; a mid-stream failure at token 800 is visible, wasteful and, if you auto-retry, may show the user a second, differently-worded answer replacing the first. ## The idempotency angle Because a retry regenerates, any side effect you performed from partial output — a tool you already invoked, a row you already wrote, a message you already sent — can happen twice. The discipline is the same as anywhere else in distributed systems: do not act on a content block before its `content_block_stop`, and make downstream effects idempotent, keyed on something stable rather than on the stream itself.
- Can you resume an Anthropic stream from where it broke?No — there is no resume token or byte offset, and a retry is an entirely new request that regenerates from the start and bills the input again. The nearest workaround is a product-level one: issue a fresh request whose final message is an assistant turn containing the text you already received, so the model continues from it. That changes the prompt and can leave a visible seam, so treat it as a deliberate design choice.
- How does a mid-stream failure differ from a 429 or 529 on the initial call?A pre-stream failure never reached the user: you back off, retry, and nobody notices. A mid-stream failure has already painted text on screen, so retrying replaces visible output with a differently-worded answer. That is why you count and alert on them separately, and why the UI needs an explicit interrupted state rather than a silent swap.
- You relay the stream through your own backend to a browser. How do you signal the failure?With an explicit error frame in your own event protocol, sent before you close the connection. Simply hanging up is indistinguishable from a normal end-of-stream, so the browser would render a truncated answer as if it were complete. Forwarding the provider's raw frames is also a poor default: you leak vendor error detail and lose the chance to attach your own request id.
- What stops a mid-stream failure from causing duplicate side effects?Never act on a content block before its content_block_stop, and make anything downstream idempotent under a stable key. If a tool-use block was only half-streamed, its arguments are truncated and must not be executed; if you already executed a completed tool call and then the stream died, the retry will likely call it again, so the tool itself has to tolerate that.
saying these in an interview costs you the question
- Expecting an HTTP 500 for a failure after streaming starts
- Believing the stream can resume from the last event
- Retrying blindly without backoff on overloaded errors
- Closing the socket as the only failure signal to clients
- Showing partial output styled as a completed answer