Every health check is green and the process is up, yet consumers say the numbers stopped updating hours ago — what should the alarm have measured?
answer
- up is not the same as working
- alarm on the result, not the process
- silence must trip it
- evaluated from outside the run
- freshness, then volume, then distance behind
basics
~20 sLiveness is not progress. The useful alarm is a threshold on what the job emits — the freshness and volume of its own output, or its distance behind the input — evaluated by something outside the job, so that silence itself trips the alarm.
solid answer
~50 sA health check answers 'is the process running', which is not the question anyone cares about. A job can be alive and producing nothing: a unit of work retrying in a loop, a step waiting on a destination that stopped accepting writes, a source that returned no records, a coordinating process alive with every worker gone. The alarm belongs on the **result**: the age of the newest output record against the wall clock, the count of results emitted per interval against the expected count, or the job's distance behind the input crossing a threshold. Two properties make it work. It must be evaluated **outside** the job, or it dies with the thing it watches. And **absence must be the trip condition** — a check that only fires when a bad value arrives never fires when nothing arrives at all.
go deeper
Remember that a process can be running and doing nothing useful. Checking that something is alive is not the same as checking that results are still coming out.
Explain the failure modes that keep a process up while progress stops: a unit of work retrying forever, a blocked write to the destination, a source returning nothing, workers lost under a live coordinating process.
Show the design: a threshold on output freshness or volume, evaluated outside the run, with absence as the trip condition. Say what you instrumented in advance to make that possible.
Frame it as the promise rather than the dashboard. Decide what consumers are actually owed, how stale is too stale for each of them, and who is woken when it is breached.
## Liveness against progress A process being up is the weakest possible claim about a running job. It is also the easiest thing to check, which is why so many jobs are monitored that way and why this scenario is a standing senior interview question. The gap has a precise shape: **liveness** says a process exists and responds; **progress** says useful work is still coming out. Everything that goes wrong in the middle keeps the first true while making the second false: - A unit of work — the smallest thing the engine hands a worker process and the smallest thing it retries by itself — fails and is retried repeatedly on the same poisonous input. The run is alive and permanently stuck. - A step is blocked writing to a destination that has stopped accepting writes. Everything is up; nothing lands. - The source returns no records, either because it genuinely went quiet or because a credential expired and the read is failing silently. The job dutifully processes nothing. - The coordinating process — the single process that turns the program into pieces, hands them out to workers and tracks which finished — is alive, but the workers it needs were reclaimed and never replaced. - The job is producing output, but the output is empty because a filter upstream now matches nothing. Only the last of these is visible in the job's own logs as anything unusual, and none of them is visible in a liveness check. ## Alarm on what the job emits The fix is to measure the thing consumers actually depend on. Three shapes, in rough order of usefulness: 1. **Freshness of the output.** The age of the newest record at the destination, against the wall clock. This is the closest thing to what the complaint was, so it is the closest thing to the right alarm. 2. **Volume against expectation.** The number of results emitted per interval compared with what that interval normally produces. This catches a job that is running and emitting near-empty results, which freshness alone will not. 3. **Distance behind the input crossing a threshold** — lag in the sense of unread input, how far the newest record the job has finished with sits behind the newest record its input holds. This is the earliest of the three on a job reading a growing source, and it exists only on such a job. | Alarm on | Catches | Misses | |---|---|---| | the process being up | the machine or container dying | every form of stuck-but-alive | | the run exiting non-zero | an outright crash | a run that never exits at all | | output freshness | stuck, blocked, starved, crashed | a job emitting fresh but wrong or empty results | | output volume per interval | empty or collapsed results | a slow drift in correctness | | distance behind the input | a growing shortfall, early | a job with no growing input to be behind | ## The two properties that make it work **It must be evaluated outside the job.** An alarm the job raises about itself is not an alarm, it is a log line: a job that has stopped doing anything has also stopped raising alarms. The evaluation belongs to something whose life is independent of the run. **Absence must be the trip condition.** Most checks are written to fire when a bad value arrives. The failure in this scenario produces *no* value, so the condition has to be inverted: fire when no acceptable value has been seen within a window. Practically, that means a check with a deadline attached rather than a comparison against an incoming figure. ## What varies by execution model The words 'the job emits a result' do not mean one thing across this family, and an alarm designed for one runtime misreads another. - On a **continuous job built from repeated small finite runs**, each slice completes and can report success while having processed nothing. A slice-success alarm is nearly worthless here; slice output volume is what you want. - On a **record-at-a-time runtime with key-bound state**, output is continuous rather than periodic, so 'a result per interval' has to be defined by you — usually as a rate over a window — before it can be thresholded. - On the **two-phase disk-to-disk model**, where every phase writes its whole output to shared storage before the next reads it, output arrives in one lump at the end, so freshness at the destination is a coarse signal and the alarm has to be a deadline on the run itself. ## Where this stops being ours A threshold on this job's own signals, or on the result a grouping emits, is an operating concern of the running job. Promises expressed across a *schedule* — this pipeline must have completed by 06:00, this task may be retried three times before the schedule gives up — belong to the system that decides when runs start, which is a different graph with different vocabulary. Keep the two apart in an interview answer; conflating them is the commonest way this question is half-answered.
- Why must the alarm be evaluated outside the job?Because the failures worth catching are exactly the ones that stop the job doing anything, including raising alarms. Self-reporting covers only the cases where the job is healthy enough to notice, which are the cases you least need help with. An independent evaluator with a deadline turns silence into a signal.
- Output is fresh and the volume is normal, but the numbers are wrong. Does this alarm help?No, and it is important to say so. Freshness and volume prove the machinery is turning, not that the answers are right. Catching wrong-but-plausible output needs checks on the values themselves — ranges, totals against an independent count, comparisons against the previous period — which is a correctness exercise rather than an operating one.
- How do you choose the deadline before absence trips the alarm?From the consumer's tolerance, not the job's schedule. Work out how stale the output can be before someone makes a bad decision on it, subtract the time needed to react, and set the deadline there. Then check the deadline is comfortably longer than the job's normal worst-case gap, or the alarm trains people to ignore it.
saying these in an interview costs you the question
- Treats a green liveness check as evidence the job is making progress.
- Puts the alarm inside the job, so it dies with what it watches.
- Writes a condition that fires only when a bad value arrives.
- Alarms on the run exiting non-zero, though a stuck run never exits.
- Assumes fresh output implies correct output.
- Confuses a deadline across a schedule with a threshold on this run's own signals.