skip to content

An irrigation controller's long-poll feed drops some events and repeats others under load; what happens between each response and the next request?

level: seniorimportance: must knowfreq 55%

answer

  1. nobody is listening between responses
  2. the gap is where events vanish
  3. position, not wall-clock time
  4. echo the position the response returned
  5. retain a bounded window per channel

basics

~20 s

Nothing of that client is open in the gap. Unless the server keeps each client's position and the events published since it, anything emitted between a response and the next request is lost — and re-asking by timestamp repeats events instead.

solid answer

~50 s

Long polling answers one request and then has nowhere to put the next event until the client asks again. That interval — response written, client processing, next request in flight — is the gap, and it is structural, not a bug you can tune away. Two failures come out of it. **Drops**: an event published in the gap is gone if the server retains nothing. **Duplicates**: a client that re-asks by wall-clock time gets everything at that instant again, or misses events sharing it, depending on which side of the boundary you take. The fix is a position, not a time: give each channel a monotonic sequence, return the next position with every response including the empty ones, and have the client echo it in the next request. The server then retains a bounded window of events per channel and answers with everything after the position it was given.

code

http · 9 lines
http
HTTP/1.1 200 OK
Content-Type: application/json
Content-Length: 37

{"events":[{"seq":1739}],"next":1740}

GET /pivot/42/events?since=1740&wait=45 HTTP/1.1
Host: fleet.example.net
Accept: application/json

go deeper

for a junior

Remember that a long poll has moments when the client has no request open, and that events published in those moments need somewhere to wait.

for a middle

Explain the position parameter: the server names the next position in every response, the client echoes it, and the server answers with everything after it.

for a senior

Diagnose from the symptom — drops clustered at bursts mean nothing was retained, duplicates at burst boundaries mean the client resumed by time — and size the retention window from the measured gap rather than guessing.

for a principal

Own the delivery promise this design makes. At-least-once with a bounded window is a commitment to consumers, and the alternative to stating it is every team inventing its own answer to a hole they cannot see.

This is the failure that separates people who have read about long polling from people who have run it. It only appears under load, because the gap only matters when something is published inside it. ## The gap is structural One cycle looks like this: the server writes a response, the client reads and processes it, the client issues the next request, the server parks it. Between the first step and the last, **no request of that client is open on the server**. Nothing is listening. The gap is small — a round trip plus whatever the client does with the response — but it exists on every cycle, and at a hundred events a minute something will land in it. This is the concrete difference between long polling and a transport that is genuinely held open: not that long polling is slower, but that it has moments of not being there at all. ## The two failure shapes | symptom | mechanism | what it looks like | |---|---|---| | events missing | the server retained nothing, so a publish during the gap had no reader and no store | gaps that correlate with bursts, worse on slow clients, invisible in the server's own logs because the publish succeeded | | events repeated | the next request asked for everything since a timestamp, and events at that instant qualify again | duplicates clustered at burst boundaries, the same event arriving on two consecutive cycles | Both are silent. The publish path reports success in each case, and only the consumer knows. ## What the server must remember 1. **Assign a monotonic position per channel.** A counter the server controls, increasing on every publish. Not a timestamp. 2. **Return the next position with every response**, including the empty one at hold expiry, so the client always has an exact place to resume from. 3. **Retain a bounded window of published events per channel** — by count, by age, or both — so events published in a gap are still there when the next request arrives. 4. **Answer with everything after the position supplied**, batched into one response, rather than one event per cycle. 5. **Park the request only when the supplied position is already at the head.** If the client is behind, there is nothing to wait for; answer immediately. 6. **Say so explicitly when the supplied position has fallen out of the retained window**, instead of returning the oldest event you still have and implying continuity that is not there. The client can then refetch current state knowingly. ## Why wall-clock time is the wrong key Timestamps look like positions and fail as positions: - **Ties.** Several events can share a millisecond, and then an inclusive boundary repeats them all while an exclusive one skips them all. - **Skew.** If more than one process stamps events, their clocks disagree, and an event can be stamped earlier than one already delivered. - **Non-monotonicity.** Clocks are adjusted, and an adjustment backwards means a later event sorts before an earlier one. A per-channel counter has none of these properties, and it also gives the client something it can log: the exact position it had before a burst went wrong. ## What the client must do - Send back exactly the position the last response returned — not a position it computed itself from the events it saw. - Treat handlers as **idempotent** anyway. Retention windows, redeliveries after a failed write and restarts all make at-least-once the honest expectation, and the cheapest defence is a handler that can process the same position twice without harm. - Keep the position across a restart if the device is expected to resume rather than re-sync, and be explicit about which of those two it is doing. ## The cost of the window Retention is not free, and it is the tuning knob this design gives you. - Too short a window and slow clients fall out of it during a burst, which converts a silent drop into an explicit re-sync — better, but still a cost. - Too long a window and you are storing a growing history for clients that may never come back. - The window only needs to cover the worst realistic gap plus whatever outage you intend to survive, and that is a number you can measure rather than guess: it is the distribution of time between a client's response and its next request. The summary worth saying out loud in an interview: long polling gets you low latency without a held connection, and in exchange it makes delivery the server's bookkeeping problem, because there are moments in every cycle when nobody is listening.

  • Why is a timestamp a poor key for resuming a long-poll feed?
    Three reasons: events can share the same instant, so an inclusive boundary repeats them and an exclusive one skips them; several publishing processes disagree about the time; and clocks move backwards when adjusted, which can sort a later event before one already delivered. A per-channel monotonic counter has none of those properties.
  • What should the server do when the position it is given is older than what it retains?
    Say so explicitly — answer with a marker that the gap is unrecoverable and the current head position, so the client refetches state and continues from there. The failure mode to avoid is returning the oldest retained event as though nothing were missing, which hides a hole the consumer will never detect.
  • Why must the empty hold-expiry answer also carry a position?
    Because it is the value the client will send next. Without it the client reuses a position from an older response, and any client-side bookkeeping between the two can drift. Carrying the head position on every response, including the empty one, means the next request always resumes from a point the server named.

saying these in an interview costs you the question

  • Assumes the server can push into the gap because something is still open
  • Keys the next request on a timestamp and gets boundary duplicates
  • Believes an empty answer needs no position, so the client reuses a stale one
  • Keeps an unbounded per-client buffer and calls it delivery
  • Treats duplicates as harmless without making handlers idempotent
  • Thinks the gap disappears if the client just re-issues fast enough