A reader group was down six hours on a stream taking 10,000 records a second and can now drain 12,000 — when is its backlog of unread records gone, and what could still be lost?
answer
- two answers, not one
- backlog over surplus
- cross-check in time units
- the race is against expiry
- unread age falls while surplus holds
basics
~20 sThirty hours: 216 million unread records divided by a 2,000-a-second surplus. Nothing is lost while that surplus holds and the six-hour unread age sits inside the retained window — loss arrives only if the surplus goes to zero or negative and the oldest unread records age out unhandled.
solid answer
~40 sSix hours at 10,000 a second is 216 million unread records. The surplus is 12,000 minus 10,000, so the backlog closes in 216,000,000 / 2,000 = 108,000 seconds, thirty hours. Cross-check in time units: the group advances through 1.2 seconds of history per second, so the age of the oldest unread record falls by 0.2 hours each hour — six hours of age, thirty hours to clear, which agrees. The deadline that matters is not patience but expiry: while the surplus is positive the unread age only falls, so nothing ages out. If a peak pushes arrivals above 12,000, the surplus inverts, the unread age climbs again, and once it reaches the retained window the oldest records are removed before anyone reads them.
go deeper
Get the first number right: 216 million unread, a surplus of 2,000 a second, thirty hours. Dividing by the full drain rate is the usual mistake and gives a fivefold under-estimate.
Show the cross-check in time units — the unread age falls by the ratio of drain to arrival rate minus one — and explain why the two methods must agree. Name the retained window as the second deadline.
Demonstrate that you plan against the smaller of the two deadlines, that you check whether the drain window crosses a peak, and that you know expiry during a backlog is silent because nothing errors and the gap falls anyway.
Frame the case where the arithmetic does not close: that is a capacity or retention decision taken in advance, and a deliberate choice about what may be lost, not something an operator improvises at three in the morning.
This is the arithmetic a drain plan is made of, and it has two answers rather than one: when the group is level, and whether anything is destroyed before it gets there. ## Working the numbers ``` arrival rate R = 10,000 records/s drain rate D = 12,000 records/s outage 6 hours = 21,600 s backlog B = R x outage = 10,000 x 21,600 = 216,000,000 records surplus = D - R = 2,000 records/s time = B / surplus = 216,000,000 / 2,000 = 108,000 s = 30 hours ``` The same plan in time units, as a check. While the group is behind it is reading historical records, so it advances through `D / R` = 1.2 seconds of history for every second of wall clock. The age of the oldest unread record therefore falls at `D / R - 1` = 0.2 per second of wall clock, which is 0.2 hours per hour. Starting at six hours of unread age, that is thirty hours — the two methods agree, and disagreement between them would mean one input was wrong. ## The deadline is expiry, not patience Thirty hours is a long time to be behind, but the question that decides whether this is an incident or an inconvenience is different: **will the oldest unread record still exist when the group reaches it?** Records do not wait indefinitely. A cluster holds a **retained window** — a bound on how far back records survive — and designs differ in how it is expressed: as an age or size bound on a whole stream, as a per-subscriber bound, or as a maximum lifetime for an individual waiting record. In every case there is a boundary the drain is racing, and it is not one that patience wins. The useful property is that while the surplus is positive, the unread age falls monotonically. It starts at six hours and only shrinks, so as long as six hours was already inside the retained window when the drain began, nothing ages out during the recovery. Here the drain is safe — thirty slow hours, but not a lossy thirty hours. ## Where the loss actually comes from | Situation | Unread age behaviour | Outcome | |---|---|---| | Drain rate above arrivals | falls steadily | Group levels; nothing expires during the drain | | Drain rate equal to arrivals | flat | Never levels, and never loses records either | | Drain rate below arrivals | climbs | Reaches the retained window; records expire unhandled | | Readers stopped | climbs at one second per second | Fastest path to the boundary | So loss during a drain means one thing: the surplus went to zero or below for long enough for the unread age to reach the retained window. In this scenario that needs arrivals to exceed 12,000 a second — a daily peak, a retried flood from an upstream system that also just recovered, or a second incident. It is worth checking explicitly whether the thirty-hour window crosses a peak, because an estimate computed on an average arrival rate can hide a period of negative surplus inside it. ## Why this loss is quiet Expiry during a backlog is one of the least visible failures in broker operations, for three reasons: 1. **Nothing errors.** Removing records that have passed the retained window is routine housekeeping; the cluster is doing exactly what it was configured to do. 2. **On designs where expiry removes unread records, the unread count can fall with nothing processed at all.** A falling gap reads as recovery when it is actually data leaving the other end. 3. **The reading side sees no gap.** A reader continues from where it was; it is not told that the records between its position and the new oldest record ever existed. The defence is to plan the drain against the smaller of the two deadlines — the time to level and the time until the unread age reaches the retained window — and to state both when reporting. If the second is the binding one, the honest conclusion is that the current capacity plan does not recover this backlog intact, and the decision that follows is not an operator's alone. ## What this question is really testing That you treat catching up as a plan with numbers in it rather than as waiting. An engineer who answers "thirty hours, and nothing is lost because the gap only shrinks while we are ahead of arrivals, but if the evening peak takes arrivals over twelve thousand this stops being true" has demonstrated the whole of it: the surplus, the cross-check, the second deadline, and the condition under which the safe answer stops being safe.
- The evening peak lifts arrivals to 14,000 a second for three hours. What changes?The surplus inverts to minus 2,000 a second for those hours, so the backlog grows by about 21.6 million records and the unread age climbs by roughly 0.17 hours per hour instead of falling. The drain is not lost, just longer, and it is still safe while the age stays inside the retained window. If peaks are frequent enough that the average surplus approaches zero, the plan does not recover at all.
- How would you know that records had already expired unread?Compare the group's recorded read position against the oldest record the cluster still holds, rather than watching the gap. On designs where expiry removes unread records the unread count falls when they go, which looks like progress. A reader that resumes beyond a boundary it never crossed, or a falling gap with a flat completion rate, are the two signals worth checking early.
- If the arithmetic says the backlog cannot be drained intact, what is the decision?It stops being an operations question. The options are to buy more drain capacity than the plan assumed, to extend how long records survive if that is possible and affordable, or to accept the loss deliberately and act on the group's recorded read position — which is a decision for the people who own what the records mean, not for whoever is on call.
- Why measure the unread age at all if the record count already gives a time?Because the two deadlines are expressed in different units. The count tells you when the group levels; the age tells you how close the oldest unread record is to the retained window. They agree on the first answer and only the age answers the second, which is why a plan quotes both.
saying these in an interview costs you the question
- Divides 216 million by 12,000 and answers five hours
- Says records cannot be lost because the readers are healthy
- Treats a falling unread count as proof of progress
- Assumes the arrival rate holds flat across a thirty-hour drain
- Never checks the oldest unread record against the retained window
- Claims every platform expresses the retained window the same way