skip to content

After an outage, what must a reader group's reading rate exceed before its backlog of unread records starts to shrink?

level: juniorimportance: must knowfreq 62%

answer

  1. two rates, one difference
  2. surplus, not throughput
  3. steady-state capacity only holds the line
  4. divide by the difference
  5. peak arrival rate, not average

basics

~20 s

The arrival rate. Only the surplus between reading rate and arrival rate drains a backlog of unread records, so capacity sized to keep pace in steady state merely holds the backlog at whatever depth the outage left it.

solid answer

~40 s

Writers do not pause while the reading side recovers, so a catch-up plan works on two rates, not one. If records arrive at `R` per second and the reader group completes `D` per second, the backlog shrinks at `D - R` per second and the time to close it is the backlog divided by that surplus — never by `D`. A group provisioned for normal traffic has `D` roughly equal to `R`, a surplus near zero, and will sit at the same depth indefinitely while looking perfectly healthy. So the first question of any drain is where extra capacity above the arrival rate is going to come from, and the second is whether the arrival rate you used is the average or the peak, because a daily peak can turn the surplus negative again.

go deeper

for a junior

Remember there are two rates, and only their difference empties anything. Backlog divided by (reading rate minus arrival rate) is the catch-up time, and matching the arrival rate exactly means never recovering.

for a middle

Be able to run both units and show they agree: the unread count falls by the surplus each second, and the age of the oldest unread record falls by the ratio of the two rates minus one. Disagreement means an input is wrong.

for a senior

Quote drain rate as completed work, not records fetched, and compute against the arrival rate present during the drain window rather than the daily mean. State the assumption you used out loud when you give an estimate.

for a principal

Decide whether surplus capacity for recovery is a standing cost or something bought at the time, and make sure the answer exists before an incident. A group with no reachable surplus has no recovery plan, only a hope.

A **backlog of unread records** is the set of records that exist in a stream and have not yet been handled by the reader group responsible for them. It appears whenever the reading side stops, slows or crashes while the writing side carries on, and it is the normal aftermath of a deployment gone wrong, a downstream outage or a night of undersized capacity. The moment the readers are healthy again, the operational question is not *are they running* but *are they running fast enough*, and that turns out to be arithmetic rather than judgment. ## Three numbers describe the whole situation - **Arrival rate (`R`)** — records written into the stream per second. This is the number people forget, because writers are unaffected by how far behind the reading side is. - **Drain rate (`D`)** — records the reader group actually completes per second, measured end to end, including the downstream work each record triggers. Not what the readers could fetch; what they finish. - **Backlog (`B`)** — the unread count at the moment the drain starts. Where a stream is measured in time rather than records, the equivalent is the **unread age** of the oldest unread record, and `B` is roughly that age multiplied by `R`. The backlog falls at `D - R` records per second. That difference is the drain; `D` on its own is not. ``` surplus = D - R time to close = B / (D - R) example: R = 4,000/s D = 5,000/s B = 18,000,000 surplus = 1,000/s time = 18,000,000 / 1,000 = 18,000 s = 5 hours ``` Read with `D` instead of the surplus, the same example gives one hour — a five-fold under-estimate, and the reason so many recovery estimates are announced and then quietly revised. ## Why steady-state capacity never recovers Most reader groups are sized so that they keep up with normal traffic and no more, sometimes with a modest safety factor. That sizing gives `D` approximately equal to `R` and a surplus of approximately zero. Such a group, restarted after a two-hour outage, is not broken and is not falling further behind — it is simply frozen at the depth the outage created. Every record it handles is matched by a new arrival. Nothing on a dashboard looks like an error; the gap is flat rather than rising, and it stays flat for as long as you leave it. Recovery therefore requires deliberately creating capacity that does not exist in normal operation, which is what makes a drain a plan rather than a wait. ## The same arithmetic in time units Where the gap is tracked as the age of the oldest unread record rather than a count, the drain shows up differently. While the group is behind, it reads historical records, advancing through `D / R` seconds of history for every second of wall clock. The unread age therefore changes at `1 - D / R` per second: it falls only when `D > R`, holds when `D = R`, and climbs at up to one second per second when the readers are stopped altogether. | Quantity | What it is | Behaviour while `D > R` | |---|---|---| | unread count | records not yet handled | falls by `D - R` each second | | unread age | age of the oldest unread record | falls by `D / R - 1` each second | | arrival rate | records written each second | unchanged by the drain | Both units give the same answer for when the group is level, which is a useful cross-check on a plan: if they disagree, one of the two inputs is wrong. ## Peak, not average A surplus computed against a daily average is optimistic for the hours that matter. If traffic doubles for three hours each evening, a group with a small surplus at the average rate has a **negative** surplus during the peak and gives back part of what it gained. Compute the plan against the rate that will actually be arriving during the drain window, and if that window spans a peak, say so in the estimate rather than discovering it. ## What varies between platforms The arithmetic is universal, but what limits `D` is not. Where a stream is split into a fixed set of parts and each reader takes some of them, the number of readers that can usefully work at once is capped by the stream's shape, and the surplus is capped with it. Where readers merely compete for records from a shared queue, there is no such cap, and the ceiling appears somewhere else — a broker-side rate ceiling on the client, or, more often than anyone expects, the downstream system each record is written into.

  • Arrivals are not flat across the day. How does that change the estimate?
    The surplus is a function of time, not a constant. Compute it against the arrival rate that will be present during the drain window: a group with a thin surplus at the daily average can have a negative surplus through a peak and lose ground for those hours. Quote catch-up times against peak, and if the window spans one, say which part of the estimate assumes recovery only happens off-peak.
  • What if the reading side can only ever match the arrival rate exactly?
    Then the backlog freezes at its current depth and the age of the oldest unread record holds flat. That is stable, not recovering, and it can persist for days without triggering anything. The options are to create surplus capacity, to reduce what arrives, or to take a decision about the group's recorded read position — which is not an operator's call alone and belongs with the stream's owner.
  • Why measure the drain rate end to end rather than at the broker?
    Because the surplus is set by completed work, not by records handed to the reader. If each record triggers a write to a downstream store, the store's throughput is the real `D`, and a group that fetches far faster than it finishes only builds an in-memory queue of its own. Measure where the work actually ends, and the plan stops being optimistic.

Bailing out a boat that is still taking water: what empties the hull is buckets per minute minus litres per minute coming in, and a crew that bails exactly as fast as the leak keeps the boat afloat forever without ever getting it dry. There is also a limit on how many people physically fit at the gunwale, which is why the answer is rarely just 'more hands'.

saying these in an interview costs you the question

  • Divides the backlog by total reading rate and promises an hour
  • Assumes writers pause while the reading side catches up
  • Says any running reader group must eventually catch up
  • Calls a flat, non-zero gap recovery in progress
  • Uses the daily average arrival rate for an evening drain
  • Measures the drain rate at the fetch instead of at completion