skip to content

How do you get metrics out of a batch job that exits before any collector could poll it?

level: seniorimportance: nice to knowfreq 42%

answer

  1. the job is gone before anyone asks
  2. hold it, send it, or have someone else say it
  3. a held value never looks stale
  4. attribution moves to whoever reports
  5. alert on age of last success, not presence

basics

~20 s

Three options: push to an intermediary that holds the last value for later collection, push straight to a receive-capable backend before exiting, or have a long-lived process report on the job's behalf. Each trades away either freshness or correct attribution.

solid answer

~50 s

A job that runs for fifteen seconds against a twenty-second collection cycle can never be polled reliably, so something else must carry its numbers. **A last-value gateway** accepts a push and keeps serving it, which fits a polling estate neatly but means the value outlives the job and must be deleted explicitly when the job is retired. **Pushing directly** to a backend that accepts writes avoids the stale value but needs the flush to finish before exit, so a killed job reports nothing. **Reporting through a long-lived process** - the scheduler, a node shipper, or an endpoint reading a job-status table - is the durable option but attributes the measurement to the reporter, so job identity must be carried deliberately. Whichever you pick, record a last-success timestamp, a duration and an outcome, then alert on the *age* of the last success.

code

text · 4 lines
text
batch_job_last_success_timestamp_seconds{job="hearing-reconciliation"} 1757124023
batch_job_duration_seconds{job="hearing-reconciliation"} 15.4
batch_job_records_processed_total{job="hearing-reconciliation"} 8412
batch_job_exit_ok{job="hearing-reconciliation"} 1

go deeper

for a junior

Recall the basic problem: a process that starts and finishes between two collections is never seen, so a job's numbers must be sent somewhere rather than waited for. Know that a job's success time is the useful thing to record.

for a middle

Explain the three routes - an intermediary holding the last value, a direct send before exit, or a long-lived process reporting on the job's behalf - and describe the cost of each in terms of stale values, lost flushes and misattribution.

for a senior

Show that you have operated these. Talk about deleting retired entries from a holding gateway, jobs killed before they flush, keeping run identifiers out of metric dimensions, and alerting on the age of the last success rather than on missing data.

for a principal

Own the standard. Decide which route the whole estate uses for scheduled work, who deletes stale entries, and how job outcomes stay auditable months later - the case where a report has to prove a control ran, not merely that nobody was paged.

## Why the pull model has no answer here A nightly reconciliation job in a courtroom-scheduling estate starts at 02:00, runs for 15.4 seconds and exits. A collector polling every 20 seconds might catch it, might catch it half-initialised, and will usually miss it entirely. Shortening the interval does not fix it - it narrows the race. The job cannot meet the pull model's basic demand, which is to be alive and answering at a moment chosen by someone else, so the numbers have to leave the process by some other route. There are exactly three routes, and each gives something up. ## Option 1: an intermediary that holds the last value The job pushes its results to a small service that stores them and exposes them for normal collection; Prometheus's Pushgateway is the canonical implementation of this shape. It fits a polling estate with no change to the collector, and it is genuinely the right answer for job-completion facts. Its costs are specific: - **The value outlives the job.** The gateway keeps serving the last thing it was given, so a job that stopped running weeks ago still presents fresh-looking numbers. Nothing expires on its own. - **Deletion is manual.** Retiring a job means explicitly removing its entries, or they linger indefinitely and quietly pollute every aggregate over them. - **The gateway's health masquerades as the job's health.** The collector reports that it reached the gateway, which says nothing about whether the job ran. Per-run liveness is gone. - **It is a shared chokepoint.** One process now holds the results of many jobs and becomes both a single point of failure and a place where pushes from different runs can overwrite one another if they share the same grouping identity. The anti-pattern that follows is using it for long-running services: a service that pushes into the gateway every minute gets all four costs and none of the benefits of being polled directly. ## Option 2: push directly from the job If a backend accepts writes, the job can send its own measurements and exit. Nothing is held on your behalf, so there is no stale value to explain. What you take on instead: - **The flush must complete before the process ends.** A job killed by a timeout, an out-of-memory condition or a node eviction reports nothing - the very failures you most wanted to see are the ones that go unreported. - **There is no retry budget.** A long-lived sender can retry for minutes; a job that is trying to exit has seconds, and holding it open to retry distorts its own duration. - **Per-run identity is a trap.** Attaching a unique run identifier as a dimension is the obvious move and the expensive one: a job running 40 times a day for six months is thousands of series that will never be queried individually. ## Option 3: report through a long-lived process The scheduler, an orchestrator controller, a node-level shipper or an endpoint that reads a job-status table can publish job outcomes for as long as anyone cares. This is the most durable option, because the reporter is still alive when the collector calls, and it can report facts the job itself could not - most importantly *that a job never started at all*, which neither of the first two options can ever tell you. Its cost is **attribution**: the measurement belongs to the reporter, so the job's own identity must be carried deliberately as a dimension, and the reporter becomes another thing whose failure looks like the jobs' failure. | Option | Freshness problem | Attribution problem | Catches a job that never started | | --- | --- | --- | --- | | Last-value gateway | Value persists forever | Grouping identity can collide | No | | Direct push from the job | None | Per-run dimensions explode | No | | Long-lived reporter | Depends on the reporter | Belongs to the reporter, not the run | Yes | ## The shape that actually works: timestamps, not presence Whichever route the data takes, record the job's outcome rather than its existence: **the timestamp of the last successful completion**, the **duration** of that run, the **exit outcome**, and one or two counts describing the work done. Then alert on the *age* of the last success - for a job scheduled daily, alert when nothing has succeeded in about twenty-six hours. This works where presence-based alerting fails in both directions. A gateway never lets the series disappear, so "missing data" will never fire. A direct-push job that failed before it flushed leaves nothing to go missing, so "missing data" will never fire there either. An ageing timestamp, by contrast, is true in every one of those cases - the job did not succeed, and the alert says exactly that. It is also the form that survives an audit: when a regulator asks for six months of evidence that a nightly control actually ran, a series of success timestamps is the answer, while an absence of alerts is not.

  • Why not simply alert when the batch job's series disappears?
    Because in both push shapes it never does. A last-value gateway keeps serving the entry until someone deletes it, so nothing goes missing however long the job has been broken. A job pushing directly that dies before it flushes never created a series to lose, so there is nothing to notice either. The age of the last successful completion is true in both cases and in the healthy case too.
  • What goes wrong when a long-running service pushes through a last-value holding gateway?
    You get an intermediary with none of a metrics store's properties: values are overwritten by grouping identity rather than accumulated per instance, per-instance liveness disappears because the collector only ever confirms it reached the gateway, and the last value persists after the service is gone. A long-lived service can be polled directly, which restores all three.
  • How do you keep a run identifier without exploding the series count?
    Keep it out of the metric's identifying dimensions entirely. Metrics carry the job name, the environment and the outcome; the run identifier belongs in the job's logs or its trace, where per-run detail is cheap and queryable. If you truly need to join a run to its numbers, record the identifier as a log field alongside the same timestamp the metric carries.

A last-value gateway is a note left on a whiteboard: it keeps saying the same thing long after whoever wrote it went home.

saying these in an interview costs you the question

  • Suggests keeping the job process alive purely so it can be polled
  • Alerts on the presence of the series rather than its age
  • Forgets that a holding gateway serves its last value indefinitely
  • Routes long-running service metrics through a batch job gateway
  • Puts a per-run identifier into the metric's identifying dimensions
  • Assumes a job killed mid-run still flushed what it had