Your Server-Sent Events response is cut every sixty seconds whenever the feed goes quiet, though the origin never ended it — what cuts it, and what does the receiving client do?
answer
- silence is not liveness
- the clock resets on forwarded bytes
- the status already left the building
- reads as a stop, not an error
- reconnect load is the visible symptom
basics
~20 sAn intermediary's idle timeout ends it. That clock measures the gap since the last byte it forwarded, so a live but silent response looks dead. Because the 200 OK already went out, the body simply stops, and the receiving client reads that as a dropped stream and reconnects.
solid answer
~40 sA reverse proxy or edge tier applies an idle timeout to a response in flight, and it measures **silence**, not liveness: a stream with nothing to report for sixty seconds is indistinguishable, from the intermediary's side, from a stalled upstream. When it fires, the intermediary has no way to signal an error -- the `200 OK` status line left when the response opened and cannot be taken back -- so the body just ends or the connection is reset. The receiver sees a stream that stopped rather than a failure, and a conforming client reconnects on its own after the delay it was given. The tell is that the drops correlate with quiet periods rather than with load, and land on a suspiciously round interval.
go deeper
Recall that a stream can be ended by something in the middle rather than by either end, and that the receiving client reopens it without any application code doing so.
Explain that an idle timeout measures the gap between forwarded bytes, and that no error status can be sent once the 200 OK has already gone out.
Show the diagnosis: drops anchored to the last event rather than to the request, a round interval, survival against the origin directly, and the reconnect load the cut produces.
The judgment is how much of a fleet's stream reliability should depend on every intermediary's default idle bound, versus a standing requirement that long-lived responses keep traffic on the wire regardless of path.
## What an idle timeout actually measures An intermediary holding a response open needs some protection against an upstream that has died without saying so. The usual mechanism is a clock that resets on every byte it forwards downstream; if the clock reaches its bound, the intermediary gives up on the response. This is correct and necessary behaviour for the traffic such a tier mostly carries, where a minute of silence between the status line and the last byte of a body really does mean something is broken. A Server-Sent Events response inverts that assumption. A quiet bottling line is a healthy bottling line: no jams, no speed changes, nothing to report. The response is alive and correct and produces no bytes, which is precisely the shape the timeout was built to kill. ## Why it reads as a drop and not as an error The `200 OK` status line and the response header fields went out the moment the stream opened, possibly an hour ago. HTTP gives an intermediary no way to revise that afterwards. So when the clock fires, its options are to end the body or to reset the connection -- and from the receiver's side those look the same: a response that stopped. That is why this failure does not show up as an error status anywhere. There is no `504`, no error page, nothing to alert on. A conforming client treats a stream that stops as a stream to reopen, waits the reconnect delay it was given, and reissues the same request. Ten minutes later the operator's dashboard is still working and nobody has filed a ticket. The failure surfaces instead as load: a fresh request per reader per timeout interval, arriving at the origin forever. ## Telling the three clocks apart | What you observe | Which clock fired | |---|---| | Drops only during quiet periods, on a round interval | An intermediary's idle timeout on the response | | Drops on the same interval even while events flow steadily | A bound on total response duration, which activity does not reset | | The same drop against the origin directly, no intermediary in the path | A lifetime bound on the emitting side, which is a different subject entirely | The first two are easy to confuse and the distinction is the whole diagnosis: an idle timeout is reset by **any** forwarded byte, so a busy stream never trips it. A maximum-duration bound is not reset by anything. ## Diagnosing it 1. Plot drop times against event times. Idle-timeout drops sit a fixed interval after the **last event**, not a fixed interval after the **request**. 2. Check whether the interval is a round number in seconds. Defaults cluster on thirty, sixty and three hundred. 3. Request the same stream from the origin directly and leave it quiet. If it survives, the clock belongs to the intermediary. 4. Look at the origin's request rate for that path. A step change that matches the interval is the reconnect load this failure produces. ## What it costs, and how it is addressed The visible cost is a gap: between the cut and the reconnect, events produced by the line are not being delivered to that reader, and whether they are recoverable afterwards depends on machinery this leaf does not own. The invisible cost is steady reconnect load on the origin, plus whatever each reconnect costs to authorise and set up. Two levers exist, and they belong to different owners: - **Keep bytes flowing.** Any traffic on the response resets the intermediary's clock, so a periodic write from the emitting side prevents the timeout from ever firing on a healthy quiet stream. What that write looks like and how often it is sent belong to the emitting side's own discipline rather than here. - **Raise or remove the bound on that route.** Configuration in the intermediary, usually owned by another team, and the durable fix when the path is dedicated to streaming. Most deployments end up doing both: the periodic write because it works through intermediaries you do not know about, and the route configuration because it stops you relying on one hop's default being generous. ## Why this is a senior question It requires having operated a stream rather than having built one. The candidate must realise that silence is not health from the middle of the path, that a status cannot be retracted once sent, that the absence of an error is the reason nobody noticed, and that the visible symptom -- reconnect load at the origin -- is several steps removed from the cause. None of that is in a specification. All of it shows up on the third week of running a live feed.
- The drops happen exactly every ten minutes even while events flow steadily. Is an idle timeout still the cause?No. An idle timeout is reset by every byte forwarded, so a busy stream never reaches its bound. A cut that survives a steady flow of events points to a bound on the total duration of a response, which nothing resets, or to a periodic recycle of the connection on that hop. The distinction matters because keeping traffic on the wire fixes the first and does nothing for the second.
- What is the operational signature of this failure at the origin?A steady, periodic request rate for that path that tracks the number of concurrent readers divided by the timeout interval, with each request authorising and setting up a stream that lives exactly as long as the interval. It looks like healthy traffic on a dashboard, which is why it usually goes unnoticed until someone asks why a one-reader feed produces sixty requests an hour.
- Why is there no HTTP status that reports this to the receiver?Because the status line is the first thing sent. By the time the intermediary gives up, the `200 OK` is long gone and the receiver has been reading a body for minutes; HTTP offers no way to revise a status already on the wire. All the intermediary can do is end the body or reset the connection, which is why the failure presents as a stream that stopped rather than one that failed.
saying these in an interview costs you the question
- Expects an error status when an intermediary gives up on a response
- Thinks an idle timeout fires on a schedule regardless of traffic
- Assumes a drop during a quiet period means the origin closed the response
- Believes the receiving application must detect and reopen the stream itself
- Treats steady reconnect load as normal rather than a symptom