skip to content

Client Tolerance Rules

The rules a client follows when the stream is not clean: what it ignores, what it reads as an empty value, and when it stops retrying. Interviewers use it to separate spec readers from library users.

part ofServer-Sent Events (SSE)overview, primer and where to startread it →
on this pageshow

questions

4

During a rolling restart your arrivals feed answered 503 Service Unavailable for two minutes, and afterwards no client ever came back - why?

level: seniorimportance: must knowfreq 58%

answer

  1. reestablish versus fail, two different outcomes
  2. no answer recoverable, bad answer fatal
  3. any non-200 status ends it permanently
  4. Retry-After is never consulted here
  5. 204 No Content is the deliberate stop signal

basics

~20 s

Every client received an HTTP response whose status was not 200 OK, so each one failed the connection: no further attempt, ever. Only a network-level failure or an ended stream makes a client reestablish the connection.

solid answer

~50 s

A Server-Sent Events client separates two outcomes. If the request fails at the network level, or a stream that opened correctly later ends, the client *reestablishes the connection*: it waits its reconnection time and reissues the request. But if a response actually arrived and was unacceptable - any HTTP status other than `200 OK`, or the wrong media type - the client *fails the connection* and makes no further attempt at all. `503 Service Unavailable` is such a response, and `Retry-After` on it is not consulted. So during the restart every client got exactly one terminal answer and left. Had the instance refused or dropped the connection instead of answering, the same clients would have come back on their own. The rule applies whatever produced the status, including `401 Unauthorized` or `403 Forbidden`; the two statuses the specification singles out as followed rather than fatal are `301 Moved Permanently` and `307 Temporary Redirect`.

go deeper

for a junior

Remember that reconnection is not unconditional. A response the client rejects ends the relationship with that endpoint, while a broken connection does not.

for a middle

State both branches precisely: a network failure or an ended stream is reestablished, and any status other than 200 OK, apart from the redirects that are followed, fails the connection.

for a senior

Show the operational consequence: drain by dropping connections rather than answering errors, use 204 No Content deliberately, and alarm on stream requests disappearing rather than on errors.

for a principal

Decide fleet-wide how stream endpoints behave under deploys and partial failure, given that one bad response silently and permanently removes clients with no error signal left behind.

This is the failure that teaches the leaf, because everything about it looks like good practice. The service was drained politely, it answered honestly while it was unavailable, it even said when to come back - and the entire client fleet disappeared. ## What a client decides, and when A conforming client makes its decision on the response, once, as the stream opens, and the vocabulary is the specification's own. - **Reestablish the connection** - wait the current reconnection time, then reissue the same request, carrying `Last-Event-ID` if one has been set. The client does this on its own with no application code in the path. - **Fail the connection** - make no further attempt. There is no delay, no backoff, no ceiling; there is simply no next request. | what the client met | outcome | |---|---| | connection refused, reset, or dropped before any response | reestablish | | a good stream whose response body later ends | reestablish | | `200 OK` with `text/event-stream` | open the stream | | `301 Moved Permanently`, `307 Temporary Redirect` | follow the redirect like any HTTP request | | `204 No Content` | fail - the defined way to say "stop reconnecting" | | `401`, `403`, `500`, `503`, or any other non-`200 OK` status | fail, whatever caused it | The boundary is the surprising part: **not getting an answer is recoverable; getting a bad answer is not.** Someone who has only used a client library remembers "it reconnects by itself" and has no idea that half the responses a server can send end the relationship permanently. ## Why the restart emptied the fleet The mechanism, step by step. Each client was holding an open response. The instance shut down, so those bodies ended - a reestablish case, and every client dutifully came back after its reconnection time. That reissued request arrived while the service was still draining, and this time something answered it: `503 Service Unavailable`. That is a response, it is not `200 OK`, so each client failed the connection. `Retry-After: 120` changed nothing, because a failed connection is never retried and the header is not consulted. Two minutes later the service was healthy and the logs were quiet - not because clients were waiting, but because there were none left. Nothing short of the receiving application opening a fresh stream will bring them back. ## Draining a stream endpoint without losing the fleet 1. **Do not answer with an error status on a stream endpoint you want clients to return to.** During a restart, refusing the connection or dropping it is strictly better than answering: a refused connection is a reestablish case and the client comes back by itself. 2. **If a request must be answered, answer it correctly.** Open the stream with `200 OK` and `text/event-stream` and carry the condition as an event inside it. The stream stays alive, the application learns what is happening, and no terminal decision is made. 3. **Make the drain end the bodies, not the requests.** Ending open responses is the safe half of the sequence; it is the next request that is dangerous. 4. **Use `204 No Content` when you actually mean it.** Retiring an endpoint, decommissioning a feed, moving a client population off a host - this is the one response defined to say "stop reconnecting", and it is precise. Never send it for a transient condition. 5. **Move traffic with a redirect rather than an error.** `301 Moved Permanently` and `307 Temporary Redirect` are followed as on any HTTP request, so relocating a feed does not cost you the fleet. 6. **Alarm on the absence.** The signature of this failure is a drop in the arrival rate of new stream requests, not an error rate. If your alerting only watches responses, a fleet that has quietly gone terminal looks exactly like a quiet night. ## The design behind the rule It seems harsh until you invert it. The client reconnects with no application code involved, so if every bad response were retried, a broken endpoint would be hammered by its entire fleet forever, with nobody able to call it off. Making an answered-and-rejected response terminal puts a ceiling on that loop and gives the server one deliberate lever - `204 No Content` - to shed clients on purpose. The cost of that safety is this failure mode, and it is why the health of a stream endpoint is about which status it can be made to emit, not only whether it is up.

  • You are retiring this feed endpoint for good. How do you tell conforming clients to stop reconnecting?
    Answer the next request with `204 No Content`. That is the response defined for exactly this purpose: the client fails the connection and makes no further attempt. It beats simply refusing connections, which the client reads as a network failure and keeps retrying, and it beats an error status, which happens to work but says something untrue about the endpoint.
  • The feed moves to a different path. Does that cost you the connected clients?
    No. A redirect is followed as it would be on any HTTP request - the specification singles out `301 Moved Permanently` and `307 Temporary Redirect` - so the client issues the request to the new location and the stream opens there. Relocating is one of the few changes that is safe to make in front of live clients.
  • Does it matter to the client why the status was 401 Unauthorized rather than 503 Service Unavailable?
    Not for this decision. The rule is about the response, not its cause: any status other than `200 OK`, apart from the redirects it follows, fails the connection. What the credential problem was, and how the stream should be authorised, is a separate question - the stop-versus-reconnect rule treats every unacceptable response the same way.
  • How would you detect this happening in production before someone reports a blank board?
    Watch the rate of incoming stream requests rather than the error rate. A healthy fleet produces a steady trickle of reconnects; a fleet that has failed its connections produces none, and the endpoint looks perfectly healthy because nothing is failing any more. Alert on that arrival rate falling, per endpoint.

A shutter you find closed is one you come back to tomorrow; a notice on the door saying the shop has moved out is one you read once and never return. A conforming client reads an error response as the notice.

saying these in an interview costs you the question

  • Says the client retries any failure forever
  • Expects Retry-After on a 503 to schedule the next attempt
  • Thinks an error status pauses the stream rather than ending it
  • Cannot tell a refused connection from an error response
  • Believes only the receiving application can stop reconnection
open as a page

A Server-Sent Events endpoint answers 200 OK but with Content-Type: application/json - what does a conforming client do?

level: juniorimportance: should knowfreq 45%

basics

~10 s

A conforming client fails the connection permanently. It accepts only 200 OK carrying Content-Type: text/event-stream; anything else means no events are dispatched and no reconnection is ever attempted, so the feed simply stays empty.

open as a page

Your text/event-stream feed emits a line with no colon and a field name the grammar never defines - how does a conforming client treat each?

level: middleimportance: should knowfreq 42%

basics

~20 s

Both are tolerated, not rejected. A non-empty line with no colon is processed as a field name with an empty value, and a field name the grammar does not define is ignored outright. Neither ends the stream.

open as a page

In a Server-Sent Events stream a server sends retry: 5s and an id value containing U+0000 NULL - what happens to each?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Both fields are ignored. A retry value is only accepted when it is entirely ASCII digits, and an id value containing U+0000 NULL is discarded, so each buffer silently keeps the value it already held.

open as a page