skip to content

Which evidence should a team carry on every job run, and which should it capture only on a deliberate re-run?

level: principalimportance: should knowfreq 40%

answer

  1. always-on costs every run forever
  2. capture later needs a later to exist
  3. small and irreproducible goes always
  4. exercise the switch before the incident
  5. the window is a promise, not a dashboard

basics

~20 s

Carry always what is small and impossible to reconstruct afterwards: a few totals per step, the run's identity and shape, per-unit status. Raise the expensive evidence on demand, and only where a re-run can genuinely recreate the failure.

solid answer

~50 s

The split is between evidence that is cheap and irreproducible and evidence that is expensive and reproducible. Always-on costs on every run for the life of the pipeline; capture-on-demand costs nothing until the day you need it, but it assumes a re-run will show the same thing. That assumption fails more often than teams expect: the source may retain only a window, the cluster shape and concurrent load are gone, and a failure caused by a reclaimed machine or a particular interleaving may never recur. So carry a handful of totals per step, the run's identity and width, and per-unit status — none of which any later run can recover for the run that already failed. Raise per-record traces and wide sampling deliberately. And exercise the raising path routinely, because a switch first used during an incident is a switch that does not work.

go deeper

for a junior

Understand the two options exist at all: carry evidence on every run, or turn it up later. The second only works if the same failure can be made to happen again.

for a middle

Explain which evidence is constant in size regardless of data volume, and why that property is what makes something affordable to carry on every single run.

for a senior

Name the conditions that destroy a re-run — a source that retains a window, a cluster shape that is gone, a failure that is not deterministic — and instrument ahead of them.

for a principal

Turn it into a promise with a time on it: how long may a wrong number stand unexplained. That figure sets the always-on budget and belongs beside freshness and availability.

## Two kinds of evidence, two kinds of cost Every pipeline pays for its own diagnosability twice, and the two payments look nothing alike. - **Always-on evidence** is paid on every run, forever, whether or not anything goes wrong. A few integers per step are free at any scale. A trace per record is a second pipeline. - **Capture-on-demand** is paid only when invoked, but it buys a *future* observation, not the one you missed. It is worth exactly as much as your ability to make the failure happen again. The decision is therefore not "how much telemetry do we want" but "which observations, if we do not make them now, can never be made". ## What makes a later capture impossible Before relying on re-running with more capture, check whether the re-run can exist: 1. **The input has moved on.** A source that retains a window will have discarded the records by the time anyone looks. Whether the input can be re-read at all is its own subject, but its answer decides whether this strategy exists. 2. **The conditions are gone.** The cluster width, the machines you were given, the other work sharing the pool and the reclaimed host that triggered the recomputation are not reproducible by asking nicely. 3. **The failure is not deterministic.** Anything caused by interleaving, by placement or by a machine's own trouble may simply not happen twice. 4. **The code has moved.** By the time a wrong number is noticed downstream, the logic that produced it may be two deployments old. Where any of these holds, the evidence must have been carried, or it does not exist. ## A defensible default split | Evidence | Cost per run | Reconstructible later? | Verdict | |---|---|---|---| | A few totals per step, named by condition | Negligible | No, not for the run that failed | Always | | Run identity, width, start and end, per-unit status | Negligible | No | Always | | A bounded sample of rejected records | Small and fixed | Sometimes | Always, capped | | Per-record traces through every step | Large, on every run | Usually yes | On demand | | Full payloads of every rejected record | Large and sensitive | Sometimes | On demand | | A counter per grouping key | Unbounded in the summing process | Yes | Never | The first three rows share a property: they are the same size whether a run handles a thousand records or a trillion, and they describe the run that actually happened. The middle rows scale with the data and can usually be recovered by asking again. The last row is the standing temptation that quietly makes the coordinating process the fragile part of the system. ## What a mature team deliberately stops measuring There is a real cost to keeping everything, and it is not only money: - **Attention.** Evidence nobody has read in a year still has to be understood by whoever finds it during the next incident. - **Sensitivity.** Captured records accumulate whatever personal data the input carried, in a place with weaker handling than the pipeline itself. - **Change friction.** Instrumentation nobody uses is instrumentation nobody maintains, and it breaks silently at the worst moment. So pruning is part of the discipline: retire counters nobody has queried, keep the vocabulary small enough that an unfamiliar pipeline is legible, and prefer four totals everyone understands to forty that only their author does. ## The switch has to be exercised A capture-on-demand strategy rests on a path that, by construction, is almost never used. The failure is predictable: - it needs a code change and a deployment, so it takes a day at the moment it is needed in an hour; - nobody remembers where the raised output lands, or it lands somewhere with no capacity; - it was written against an older shape of the job and no longer compiles; - it raises capture on a run that can no longer read the input that failed. Run it periodically on an ordinary day, on a real pipeline, and confirm the evidence arrives where the runbook claims. ## The argument to make at the top The honest framing is not a dashboard question, it is a promise. Carrying no evidence is cheapest until the first unexplained number, and the cost of that number is absorbed by the people who trusted it, not by the team that owns the job. So ask how long an unexplained wrong figure may stand before someone must be able to say why — a day, a week, never — and buy exactly the always-on evidence that answers within that window. That number is a service promise, it belongs in the same conversation as freshness and availability, and it is the only sound basis for deciding what the pipeline pays for on every single run.

  • A failure only appears in production at full width. How does that change the split?
    It moves evidence into the always-on column, because the cheap alternative does not exist: you cannot recreate full width and the concurrent load around it in a re-run. Carry enough per-step totals and per-unit status to reconstruct the shape of the run afterwards, and accept the standing cost as the price of a failure you can only observe once.
  • What most often makes a raised-capture path unusable at the moment it is needed?
    That it has never been run. It typically needs a code change and a deployment when there are minutes to spare, it was written against an older shape of the job, or the output lands somewhere nobody remembers or nobody sized. Exercising it on an ordinary day, against a real pipeline, is what turns it from a plan into a capability.
  • How do you argue for always-on evidence to someone looking only at the bill?
    Convert it into a promise with a time on it: how long may a wrong number go unexplained. A few integers per step cost almost nothing at any scale and are the only thing that can answer for a run that already ended. The comparison is not telemetry against zero, it is telemetry against an incident whose cost lands on consumers.

saying these in an interview costs you the question

  • Assuming any failure can be reproduced later by re-running the job
  • Carrying full per-record traces on every run because storage seems cheap
  • Leaving the raised-capture path untested until the first real incident
  • Treating an unexplained wrong number as free because the job still runs
  • Keeping every counter ever added rather than retiring unread ones