Which signals on a running job move first when it degrades, and which only confirm the damage afterwards?
answer
- when does it move, not what it says
- accumulations arrive late
- spill and retries move early
- alarm on the promise, investigate with the rest
basics
~20 sLeading signals move while output still looks fine: bytes spilled to local disk, shrinking memory headroom, retried units of work, and the fraction of a step's time spent unable to pass its output on. Distance behind the input and run duration are lagging — they accumulate the damage.
solid answer
~50 sSplit the numbers by when they move. **Leading**: bytes a worker writes to its own local disk because a working set no longer fits, memory headroom falling, the count of units of work being retried, and *reported resistance* — the fraction of its time a step spends unable to hand its output on, which is an engine's way of saying something after it is the bottleneck. These move while the job is still producing correct output on time. **Lagging**: distance behind the input, wall-clock duration, and freshness at the consumer. They are accumulations, so they only cross a threshold after the underlying shortfall has run for a while. One caveat: distance behind the input also jumps when the input rate jumps, with nothing wrong in the job at all — so it is lagging as evidence of the job's own health, not lagging in general.
go deeper
Learn which numbers accumulate. A distance behind the input and a total duration are sums over time, so they always report a problem later than the figures that move the moment it starts.
Explain why spilled bytes, memory headroom and retried units move early: each one degrades while the output is still correct and on time. Then say why none of them should alarm without a duration qualifier.
Demonstrate the two-sided use: alarm on the lagging signals because they encode the promise to consumers, investigate with the leading ones. Name what you had to instrument in advance to have either.
Take a position on how many signals a team can genuinely watch. Every extra alarm costs attention, and an alarm nobody trusts is worse than one nobody has.
## Why the split matters Every number on a running job answers a question about time as well as about health: *how long after the underlying problem starts does this figure move?* Sorting the small set of available figures on that axis is what lets you say 'I would look at this one first' and defend it — which is exactly what an interviewer is after. A *leading* signal changes while the job is still meeting its obligations. A *lagging* signal is an accumulation: it integrates a shortfall over time, so by the time it is alarming, the shortfall has been running for a while and the damage is already done. ## The signals, sorted | Signal | Leading or lagging | Why it sits there | |---|---|---| | bytes spilled to a worker's local disk | leading | appears the moment a working set stops fitting, long before output slows | | memory headroom on a worker process | leading | falls continuously toward the point where a unit of work fails | | retried units of work | leading | a failing machine or a bad piece of input shows here before the run's duration moves | | reported resistance on a step | leading | a step already spends time waiting to pass output on before any consumer notices | | processed throughput | mixed | responds immediately to a change, but a healthy figure proves nothing about whether it is enough | | distance behind the input | lagging | an accumulation of the gap between work done and work arriving | | wall-clock duration of the run | lagging | known in full only at the end | | output freshness as the consumer sees it | lagging | the last thing to move and the first thing anyone complains about | Definitions used above, since the names differ between engines: a **worker process** is one process on one machine that runs pieces of the job and owns the memory they use; a **unit of work** is the smallest thing the engine hands a worker and the smallest thing it retries by itself; **spill** is bytes a worker writes to its local disk when a working set will not fit in memory; and *distance behind the input* is lag in the sense of unread input — how far the newest record the job has finished with sits behind the newest record its input already holds. ## The trap inside the lagging column Distance behind the input is lagging evidence *of the job's own health*, and people over-read it in both directions. - It rises when the job slows, slowly, because it accumulates. - It also rises immediately when the **input rate** jumps, with the job performing exactly as it always has. That is a change in the world, not a degradation. - It can be flat and large: a job holding a steady deficit is keeping up with arrivals while remaining permanently stale. - It can be zero while the newest record the job has processed is hours old, because the source went quiet. That is a freshness question about event time and belongs to a different subject; do not read it as the job being healthy or unhealthy. So distance behind the input is a good alarm for the consumer's promise and a poor first diagnostic for the job. ## What varies across execution models Which leading signals you actually get depends on the runtime, and an answer that assumes one model is wrong for the next. 1. On a **continuous job built from repeated small finite runs** — an endless input sliced into small finite jobs run back to back — the earliest signal is usually a slice taking longer than its own period. It is leading because the slices have not yet piled up. 2. On a **record-at-a-time runtime with key-bound state**, the earliest signal is the per-step busy fraction and the fill level of the in-memory queue between two steps, because records are pushed onward as they are produced. 3. On the **two-phase disk-to-disk model**, where each phase writes its whole output to shared storage before the next reads it, there is very little in flight to watch: the useful leading signals are the retried-unit count and spilled bytes, and the phase boundary is where you learn the rest. ## How to use the split - **Alarm on the lagging ones**, because they express the promise you made to a consumer, and a promise is what an alarm should protect. - **Investigate with the leading ones**, because they move early enough to act on, and because they point at a mechanism rather than at a symptom. - **Never alarm on a leading signal alone** without a duration qualifier. Spill happens by design in a healthy job; a retried unit is normal; a step waiting briefly is normal. It is spill that keeps growing, retries that keep repeating and waiting that becomes constant that mean something. What this leaf owns is the sorting. Why a particular step is the one under resistance, and whether a growing deficit will ever drain, are the two neighbouring exercises this one hands off to.
- Why is 'any spill at all' a bad alarm condition?Writing a working set to local disk is a designed fallback, not a fault: many correct, well-tuned jobs spill on every run and finish on time. The signal is a change — spilled bytes rising run over run, or growing within a run — not the presence of a non-zero figure. An alarm on presence fires constantly and is then ignored.
- Distance behind the input doubled overnight and nothing about the job changed. What happened?Most likely the input rate rose. The figure is a gap between two sides, so it moves when either one does. Confirm by reading the arrival rate beside the processed rate: if arrivals doubled and processed throughput held at its usual ceiling, the job is behaving as designed and the question is a capacity decision, not a fault.
- Can a leading signal be healthy while the job is already failing its consumers?Yes, and it is common. A job with ample memory, no spill and no retries can still sit permanently behind because its steady throughput is simply below the arrival rate. Nothing is degrading; the capacity was never sufficient. That is why the lagging signals carry the alarms.
A supermarket checkout. The queue length is the lagging number: by the time it is visibly long, the till has been struggling for several minutes and the damage to everyone's evening is already done. The leading number is the till itself hesitating — a scan that has to be repeated, an assistant called over. Both are real, but only one of them gives you time to open a second lane.
saying these in an interview costs you the question
- Treats distance behind the input as the first diagnostic to reach for.
- Alarms on any spill at all, though spilling happens by design.
- Assumes a rising distance behind means the job degraded, not that arrivals rose.
- Says a run's duration is a leading indicator of that same run.
- Reads a healthy throughput figure as proof the job is keeping up.