Neither RED nor USE fits a batch job or an event consumer. What do you instrument on those components instead?
answer
- The checklists assume a caller or a capacity
- Measure distance from good, not busy-ness
- Absence of a sample is not zero
- Last-success time beats a success counter
basics
~20 sInstrument what the component owes someone: for a batch job, the time of its last successful completion; for an event consumer, the age of the oldest unprocessed item; for both, items processed, failed and skipped.
solid answer
~50 sRED assumes a caller waiting for an answer and USE assumes a resource with a capacity; a scheduled job or an event consumer has neither. Four measurements replace them. **Freshness** — time since the last successful completion, or the age of the oldest unhandled item — is the closest analogue to duration and usually the most valuable. **Throughput with per-item outcome** is RED's shape with the item as the unit. **Backlog**, read in time rather than items, is the saturation analogue. **Completeness** asks whether the run covered everything. The critical emission detail: a failed-run counter cannot express a job that never started, since a counter that never increments looks exactly like a quiet healthy period. Record the last success as a timestamp you subtract from now. A cache is the exception: it serves lookups, so RED applies at its interface, plus hit ratio and miss cost.
code
pseudocode · 7 lineson_run_success:
last_success_epoch_seconds = now()
items_processed_total += processed
items_failed_total += failed
staleness_seconds = now() - last_success_epoch_seconds
oldest_unprocessed_age = now() - oldest_item_enqueued_atgo deeper
Know that RED assumes a caller waiting and USE assumes a resource with capacity, and that a nightly job has neither. The first question to ask about such a component is when it last completed its work successfully.
Explain the substitutes — staleness, per-item throughput and outcome, backlog, completeness — and why a counter of failed runs cannot distinguish a quiet healthy job from one whose schedule stopped firing.
Show that you have designed this. Talk about emitting a last-success timestamp so staleness is computable, about backlog measured in time rather than items, and about whether a run actually covered every input it was given.
Own the gap. Components with no obvious checklist are the ones that go uninstrumented across an estate, and a pipeline silently stale for eleven days is an organisational failure rather than a bug. Decide what every non-request component must emit, and who verifies that it does.
## Why the two checklists do not reach these components Both checklists carry an assumption. RED assumes something is waiting for an answer, so a rate, a failure and an elapsed time are all properties of one observable interaction. USE assumes a resource with a finite capacity, so "how much is in use" and "how much work is waiting" are both meaningful. A batch job has no caller waiting. An event consumer has a caller who left long ago. A cache has aspects of both and is fully described by neither. Applying RED to a nightly job by calling each run "a request" yields a rate of about one per day, an error count of zero or one, and a duration that is really a completion time — numbers that exist and answer nothing. ## The four measurements that replace them 1. **Freshness.** Time since the component last did its job successfully, or the age of the oldest thing it has not yet handled. It is the closest analogue to duration and usually a pipeline's single most valuable number, because it is already in the units people downstream care about. 2. **Throughput and per-item outcome.** Items processed, items failed, items skipped or set aside. This is RED's shape with the unit changed from the request to the item, and it is emitted per item rather than per run so that a run which processes half its input is distinguishable from one that processes all of it. 3. **Backlog, read as time.** How much unprocessed work is waiting — the saturation analogue — but converted into time wherever possible, because depth in items cannot be interpreted without knowing the drain rate. 4. **Completeness.** Whether the run covered everything it was supposed to. A job that processed 1,900 of 2,311 inputs and exited cleanly is a success by every other measurement here. | Component | Closest thing to rate | Closest thing to duration | Saturation analogue | The signal with no RED or USE equivalent | |---|---|---|---|---| | scheduled batch job | items processed per run | wall time of the run | inputs left unprocessed when the run ended | time since the last successful completion | | event consumer | items consumed per second | per-item processing time | backlog depth | age of the oldest unprocessed item | | cache | RED applies at its lookup interface | lookup duration | entries against the capacity limit, eviction pressure | hit ratio, and what a miss costs downstream | ## The run that never happened This is the failure mode that makes the whole leaf worth knowing. A counter of failed runs and a counter of successful runs cannot express a job whose schedule stopped firing, because a counter that never increments looks exactly like a component with nothing to report. Absence of a new sample is not a zero, and a dashboard reading "0 failures" over a stopped job is not merely unhelpful — it is actively reassuring. The fix is to record a value that moves on its own: the wall-clock time of the last successful completion, written each time the run succeeds. Read as "now minus that value", it yields a staleness number climbing every second the job does not run, whatever the reason. It cannot separate a crash from a scheduler outage, and it does not need to: it separates *working* from *not working*, which counters cannot do at all. The same reasoning covers an event consumer. Its throughput falling to zero is ambiguous — an idle stream and a stalled consumer look identical — while the age of the oldest unprocessed item rises only in the second case. ## The awkward case of a cache A cache is worth naming because it defeats a naive reading of both checklists at once. It genuinely serves lookups, so RED applies at its interface. It genuinely has a bounded capacity, so USE applies to that capacity: entries against the limit as utilization, eviction pressure as saturation. And neither of those captures what the cache exists for. That takes two more numbers: the hit ratio, and the cost of a miss to whoever asked. A cache with a 99% hit ratio whose misses take 900 milliseconds against a backing store may be hurting the median caller far less than the tail, and no utilization number will say so. ## A worked example The wind-farm maintenance planner runs a nightly job that builds dispatch plans for 2,311 turbines, and an event consumer that ingests vibration readings from them. Mid-quarter, migrating off a hosted monitoring vendor, the team lost that vendor's built-in job-monitoring integration and did not notice, because the application's own metrics were intact and reported no failures. They reported no failures for eleven days, during which the job did not run once: the migration had removed its scheduler trigger, the failure counter had nothing to increment, and of 2.3 million series in the new store exactly four described the job — all counters. A recorded last-success time would have shown the staleness on the first morning. ## What this is not Everything above is the instrumentation checklist: what a component with no caller and no obvious resource must emit before it ships. Which of these numbers becomes a promise, what target it carries, and which may wake a human are separate disciplines that consume them rather than replace them.
- Why is backlog depth alone a poor signal for an event consumer?Depth carries no fixed meaning. Forty thousand waiting items is minutes of work for a fast consumer and hours for a slow one, and the same number means something different after any throughput change, so nobody can say from the number alone whether it is bad. The age of the oldest unprocessed item is already in time units, rises the moment the consumer stalls, and stays comparable when the consumer is rewritten or rescaled.
- A nightly job has reported zero failures for two weeks. What would you check first?Whether it ran at all. A failure counter that never increments and a job that never executes produce identical telemetry, so "no failures" is not evidence of health. Look for a last-success time and read the elapsed staleness; if there is none, check whether the job's series still receive fresh samples, and whether the schedule that triggers it still exists. Then add the staleness measurement so the next occurrence is visible immediately.
- Does a cache need RED, USE, or something else?All three shapes at once, which is what makes it awkward. It serves lookups, so RED applies at its interface. It has a bounded capacity, so USE applies to that — entries against the limit as utilization, eviction pressure as saturation. Neither captures what a cache exists for, which takes two more numbers: the hit ratio, and what a miss costs the caller downstream. A high hit ratio with expensive misses can still hurt the tail badly.
A success counter on a nightly job is a smoke alarm nobody tests: silence means nothing has burned, and it equally means the battery is dead.
saying these in an interview costs you the question
- Applies RED to a batch job by calling each run a request
- Relies on a failure counter that cannot express a job that never ran
- Reports backlog only in items and never in time
- Treats hit ratio as the only thing worth measuring on a cache
- Says these components need nothing until somebody complains
- Reads a flat zero-failure line as proof the component is healthy