A continuous job reports it is four million records behind its input — why does that number alone not say whether it will recover?
answer
- level now, direction over time
- one reading answers nothing
- two readings, far enough apart
- processed rate minus input rate
- flat means no headroom, not health
basics
~20 sFour million records behind is a level, and recovery depends on the trend. Compare two readings taken minutes apart: a distance that is growing means the job is losing ground, a shrinking one means it is already draining.
solid answer
~50 sThat figure is lag in the sense of unread input — how far the newest record the job has finished with sits behind the newest record its input already holds. It is a **level**: a quantity at one instant. A level cannot separate a job quietly draining an unplanned backlog it built up during an earlier peak from a job that will never catch up, because both can read four million. Take two readings far enough apart and look at the direction. Falling means processed throughput, in records per second, currently exceeds the input rate in the same units, so the backlog will clear. Flat means the two rates are equal and the backlog is frozen where it is. Rising means a real deficit, and nothing clears until something about the job, the input or the machines changes.
go deeper
Recall that the distance behind is a level and that recovery is a direction. Get a second reading before you say anything about whether a job will catch up.
Explain the subtraction underneath: processed throughput in records per second minus input rate in the same units, and what each of the three signs implies. Say why a flat line is a warning rather than a pass.
Show that you sample against the job's own rhythm — slices, snapshot intervals, daily peaks — and that you alarm on a sustained direction rather than on a threshold crossing that a draining job will trip anyway.
Argue what the team should be promised and what it should watch: a distance-behind threshold is a poor contract because it fires late on slow deficits and early on recoveries. The defensible commitment is stated as freshness delivered to consumers, with the trend as the leading indicator.
## The number in front of you A running job that reads an unbounded input publishes some form of distance behind. The neutral statement of it is **lag in the sense of unread input**: how far the newest record the job has finished with sits behind the newest record its input already holds. It may be published as a count of records behind, or as the age of the oldest record the job has not reached, in seconds behind. Either way it is a **level** — a quantity measured at one instant, like the water sitting in a tank. A level answers `how much work is waiting`. It does not answer the question anybody actually cares about, which is `will this clear`. Two jobs can both read four million records behind at nine in the morning and have opposite futures. One ran all night well above its input rate and is eating through a queue that built up during an overnight peak; it will be clear by ten. The other has been losing a few hundred records a second for a week, and four million is simply where a slow, steady deficit has got to; it will read six million tomorrow. ## Level, trend, and the subtraction underneath The trend is the direction the level moves, and its sign is the sign of one subtraction: **processed throughput, in records per second, minus input rate, in records per second**. Always state the unit and the side, because both quantities get called throughput and the whole subject is the gap between them. | Two readings show | The rates underneath | What it means | |---|---|---| | distance behind falling | processed rate above input rate | an unplanned backlog that is draining; it will clear on its own | | distance behind flat | the two rates are equal | frozen: demand exactly consumes capacity, with no headroom left | | distance behind rising | processed rate below input rate | a real deficit; nothing clears until something changes | The flat case is the one that gets misread. A distance that has not moved all morning is not evidence of health — it says the job is exactly keeping up and has nothing spare, so the next small rise in arrivals, or any slowdown per record, turns it straight into a growing deficit. ## Sampling: how far apart, and against what rhythm The trend is only as good as the interval you sampled it over, and different execution models give the same underlying situation different shapes: - **A job built from repeated small finite runs** — the engine slices an endless input into small finite jobs run back to back — lets unread input climb while a slice is being processed and fall when that slice completes. The line saws. Two readings taken inside one slice show a slope that is not the trend; compare trough with trough, or the envelope across many slices. - **A record-at-a-time job**, where each record moves through the graph as it arrives, produces a smoother line, though on engines where writing a saved snapshot of a running job briefly slows processing, the line ripples at the snapshot interval. - **A two-phase model that writes each phase's whole output to shared storage before the next phase reads it**, launched once per interval, does not move continuously at all: the distance behind is a step function, comparable only at run boundaries. - **Daily rhythm beats all of it.** A job that is behind every afternoon and clear by midnight has no deficit. Sample across a full cycle before calling anything permanent. ## What the trend still does not tell you Three things, and saying so is part of the answer: 1. **Why.** A rising distance says the deficit is real; it does not say whether the input grew, the work per record grew, or capacity was lost. 2. **When it clears.** A falling distance says it will clear, not when. That comes from the gap between the rates, not from the level — dividing the level by processed throughput ignores everything that arrives during the drain. 3. **Freshness.** Zero unread input only means the job has reached the newest record its input holds. If the source itself has gone quiet, the newest record the job has processed can still be hours old. Those are two different senses of lag and they move independently. ## Answering it in an interview Say the level-versus-trend distinction first, then the subtraction, then the sampling caveat. An interviewer asking this is checking one habit: that your first move on a lag alarm is to get a second reading rather than to act on the first.
- Two readings ten minutes apart show exactly the same distance behind. What do you conclude?That processed throughput and input rate are equal, so the backlog is frozen rather than draining — it will not clear on its own. It also says the job has no headroom: capacity currently consumes demand exactly, so any rise in arrivals or in work per record turns a flat line into a rising one. Treat flat as a warning, not as health.
- Can a job have zero unread input and still be serving stale results?Yes. Zero unread input only means the job has reached the newest record its input holds. If the source stopped producing an hour ago, the newest record the job has processed is an hour old and anything computed from it is that stale. Name the sense you mean — unread input, or freshness of the newest processed record — because they move independently.
- Why is a lag threshold a poor alarm on its own?Because it fires on the level. A job draining a backlog crosses the threshold on its way down and pages somebody about a situation that is already resolving, while a job losing a hundred records a second stays under the threshold for days before it trips. An alarm on the direction, sustained over a window longer than the job's own rhythm, catches the second case earlier and the first not at all.
saying these in an interview costs you the question
- Treating a single distance-behind reading as a verdict on the job's health
- Assuming any nonzero backlog means the job is broken
- Reads a high processed rate as proof of recovery without comparing it to the input rate
- Calling a flat distance behind healthy when it actually means zero headroom
- Confusing unread input with how old the newest processed record is
- Comparing two readings taken inside one slice of a job that runs in slices