A job sits twelve million records behind, processes nine thousand records a second and receives six thousand a second — when is it caught up?
answer
- the gap, not the rate
- arrivals continue during the drain
- subtract input rate first, then divide
- zero or negative gap means never
- error factor is processed over net
basics
~20 sDivide the backlog by the gap between the rates, not by throughput. Nine thousand processed against six thousand arriving drains three thousand a second, so twelve million records take about four thousand seconds — a little over an hour.
solid answer
~50 sThe drain rate is the difference between the two: `9,000 - 6,000 = 3,000` records a second. Twelve million divided by three thousand is `4,000` seconds, near seventy minutes. The tempting division, twelve million over nine thousand, gives about twenty-two minutes and is wrong by exactly the ratio of processed rate to net rate — here three times — because it assumes nothing arrives while the job is draining. Two corollaries fall out of the same formula. If the processed rate is equal to or below the input rate the denominator is zero or negative and there is no finite answer, which is the arithmetic statement of a permanent deficit. And the answer is only good while both rates hold; it is a decision aid, not a prediction, so re-derive it from fresh readings rather than quoting the first estimate an hour later.
code
python · 12 linesdef drain_seconds(backlog_records, processed_per_s, input_per_s):
net = processed_per_s - input_per_s
if net <= 0:
return None # at these rates the backlog never clears
return backlog_records / net
drain_seconds(12_000_000, 9_000, 6_000) # 4000.0 -> about 67 minutes
drain_seconds(12_000_000, 9_000, 8_500) # 24000.0 -> about 6 h 40 m
drain_seconds(12_000_000, 6_000, 6_000) # None -> frozen, not draining
12_000_000 / 9_000 # 1333.3 -> the wrong formgo deeper
Recall the shape of the formula: backlog divided by the difference between what the job processes and what arrives. The subtraction is the part people forget.
Derive it and show the error the naive division makes, including why that error grows without bound as the two rates converge. State what the zero and negative denominators mean.
Attach the assumptions: sustained averages over the same window, the backlog's per-record cost, the rate the job achieves while draining rather than while idle, and a destination that may cap the net rate from outside the job.
Treat the number as a commitment, not a calculation. An hour is a notice to consumers, several hours is an intervention, and no finite answer is a capacity decision — decide in advance which band triggers which response so nobody negotiates it during the event.
## The arithmetic An unplanned backlog — input the job fell behind on — drains at the **net rate**, which is processed throughput in records per second minus input rate in the same units. Everything else follows from that one subtraction. ``` net rate = processed rate - input rate = 9,000 - 6,000 = 3,000 records/s drain time = backlog / net rate = 12,000,000 / 3,000 = 4,000 s ≈ 67 min ``` The job is not idle for those sixty-seven minutes; it is doing roughly thirty-six million records of work, because six thousand a second keep arriving throughout. Nine thousand a second for four thousand seconds is thirty-six million processed, of which twelve million was the backlog and twenty-four million arrived during the drain. ## Why dividing by processed throughput is wrong The wrong form treats the backlog as a fixed pile and the input as paused: | Form | Result | What it assumes | |---|---|---| | `backlog / processed rate` | 1,333 s ≈ 22 min | nothing arrives while the job drains | | `backlog / (processed − input)` | 4,000 s ≈ 67 min | arrivals continue at their current rate | The error factor is the ratio of processed rate to net rate, which here is exactly three. That ratio explodes as the two rates converge: at nine thousand processed against eight thousand five hundred arriving, the same backlog takes twenty-four thousand seconds, nearly seven hours, while the naive division still says twenty-two minutes. **The closer a job is to merely keeping up, the more dangerously optimistic the wrong formula becomes**, which is why it is worth being pedantic about the denominator. ## When the formula returns no answer If the processed rate equals the input rate the denominator is zero and drain time is infinite: the backlog is frozen. If the processed rate is below it, the denominator is negative and the formula returns a meaningless number — the honest reading is that the deficit is growing, and the same subtraction tells you how fast. A negative net rate of three hundred records a second is twenty-six million more records behind by tomorrow. ## What has to be true before you trust the number - **Both rates must be sustained averages over the same window.** Instantaneous figures taken one second apart are noise, and on a job built from repeated small finite runs — the engine slices an endless input into small finite jobs run back to back — an instantaneous rate mostly tells you where in the slice you sampled. - **The backlog's work may not resemble the live work.** Older records can be cheaper or dearer per record, and a job whose per-record cost depends on state held under a key can be slower on the backlog's key distribution than on today's. - **The processed rate during catch-up is often not the steady-state rate.** A job with a deep queue may run at a higher rate than it does when it is caught up and waiting on arrivals, so the drain estimate made from the caught-up rate understates its own recovery. - **The destination can be the binding constraint.** If the system the job writes to throttles, the processed rate cannot rise to its own ceiling and the net rate is set outside the job entirely. - **The input rate is not a constant.** Compute against the arrival rate you expect during the drain window, not the one you measured during the peak that caused the backlog. ## The same arithmetic in a sliced job For a job built from repeated small finite runs, the same statement is usually easier to make in time rather than records: if each slice covers a period of sixty seconds but takes ninety seconds to process, the job falls thirty seconds further behind per slice, and a backlog of one hour drains only if the slice duration is brought below its period. The net rate is `period − duration` per slice, and it is negative when duration exceeds period. A record-at-a-time job states the same thing directly in records per second, and an interval-launched two-phase model states it per run: a run that must process ninety minutes of accumulated input inside a sixty-minute window is the same deficit in a different unit. ## Using the number as a decision The point of the arithmetic is rarely the estimate itself. It is the decision it supports: an hour is something you tell consumers and wait out; seven hours is something you intervene in; no finite answer is an incident. Re-derive it from fresh readings each time, and quote it with the assumption attached — `about an hour, if arrivals stay near six thousand a second`.
- The same job's processed rate drops to six thousand a second while arrivals stay at six thousand. What is the new drain time?There is none. The net rate is zero, so the twelve million records stay twelve million: the backlog is frozen rather than draining. The job is still doing full work and its own throughput figure looks healthy, which is why the processed rate must always be read against the input rate rather than on its own.
- Why can an estimate made while the job is caught up understate how fast it will recover later?Because a caught-up job's processed rate is often limited by arrivals rather than by capacity — it is partly idle. Given a deep queue it may run measurably faster, since work is always available and per-record overheads amortise better. Measure the processed rate during the drain, not before it, and re-derive the estimate from those readings.
- How would you state the same arithmetic for a job that slices its input into small finite runs?In time rather than records: compare each slice's processing duration with the period it covers. A ninety-second duration on a sixty-second period loses thirty seconds per slice, so the deficit grows; a forty-second duration on the same period recovers twenty seconds per slice, and an hour of backlog needs one hundred and eighty such slices to clear.
saying these in an interview costs you the question
- Divides the backlog by processed throughput and ignores arrivals during the drain
- Quotes a drain estimate hours later without re-reading the two rates
- Treats a healthy-looking processed rate as sufficient without the input rate beside it
- Assumes the backlog's records cost the same per record as live traffic
- Ignores that a throttling destination, not the job, may set the net rate
- Reports a negative drain time as a number instead of as a permanent deficit