skip to content

Operating Jobs in Production

The job once it is running: the few numbers that say it is healthy, how it degrades under load, how you debug work done on machines you no longer have, and what the run cost.

on this pageshow

explore

questions

27

A step in a running job spends most of its time unable to hand its output on — what does that indicate?

level: juniorimportance: must knowfreq 62%

answer

  1. a symptom, not a fault
  2. blocked waiting, not busy working
  3. the cause lies further along
  4. a resisting step has spare capacity
  5. some engines publish no such figure

basics

~20 s

Backpressure: a step further along the job graph cannot accept work as fast as this one produces it, so this one blocks. The figure locates a bottleneck ahead of the resisting step, not a fault inside it.

solid answer

~40 s

That is **reported resistance** — the fraction of a step's time spent unable to hand its output on, which some engines publish per step and others leave you to infer. Read it as **backpressure**: a step that cannot accept more work stops the step feeding it, which stops the one feeding that, back towards the source. So a resisting step is transmitting someone else's shortage, not causing one; it is *idle waiting*, not busy, and therefore has spare capacity. Its value is directional — it tells you which way to walk. What it does not tell you is why the constrained step is constrained: uneven work across pieces, a worker process short of memory, and a destination that rate-limits writes all look identical from here.

go deeper

for a junior

Recall the one sentence: a step reporting that it cannot hand its output on is telling you something after it is slower, not that it is itself the problem. Say the word backpressure and say it is a pointer.

for a middle

Explain the mechanism — a bounded in-memory queue between two steps, full, so the producer stops, and that stopping propagating back towards the source. Say explicitly that a resisting step is idle and therefore has spare capacity.

for a senior

Show that you treat the figure as a locator and nothing more, and that you know it is a normal steady state rather than an incident. Then say what you would look at next, and in what order, to turn the location into a cause.

for a principal

The judgment is what you alarm on. Resistance is a routine steady state, so paging on it produces noise; the defensible position is to alarm on the promise made to consumers and keep resistance as the first diagnostic you open, not the trigger.

## What the figure measures An engine of this class runs a job as a graph of **steps** — one transformation applied to every piece of the input — with each step handing records to the next. Between two adjacent steps sits a **bounded in-job buffer**: a small in-memory queue holding records the earlier step has produced and the later one has not yet taken. *Bounded* is the whole point. When that queue is full, the earlier step has nowhere to put its next record, so it stops. That stopping is **backpressure**: the mechanism by which a step that cannot accept more work halts the step feeding it, which halts the step feeding that, all the way back to the read from the source. **Reported resistance** is the number some runtimes publish for exactly this — the fraction of a step's wall-clock time spent unable to hand its output on. Others publish no such figure and you infer the same thing from buffer occupancy or from per-step throughput. ## Backpressure is a diagnosis, not a fault The single most useful thing to understand about the number is that it is **not an error**. A job can run at exactly the rate its slowest step allows, for months, with several steps reporting high resistance the whole time, and be perfectly healthy: the resistance is simply how the graph settles into one common rate. Nothing is dropped, nothing is retried, no unit of work has failed. What the number does is **locate**. A resisting step is one whose consumer is slower than it is, so the constraint lies somewhere *away from the source* — further along the graph. The resisting step itself is the last place to look for a cause, because by definition it is spending its time idle. ## What it tells you - **Direction.** The bottleneck is ahead of the resisting step, in the direction the records flow. - **Spare capacity.** A step blocked 80% of the time could do roughly five times the work if something accepted it; widening *that* step buys nothing. - **A steady state.** Resistance that is high and flat means the job has already settled to the constrained step's rate. Resistance that appears and clears means an intermittent cause. ## What it does not tell you - **Why** the constrained step is constrained. Uneven work across pieces, a worker process running short of memory, an expensive per-record computation, and a destination that throttles writes all present here as a downstream step that accepts slowly. - **Whether the job is behind.** Resistance is about the relationship between two steps; how far the job sits behind its input — lag in the sense of unread input — is a separate number, and a job can resist heavily while still keeping up with everything that arrives. - **That anything is broken.** See above: it is the normal steady state of any graph whose steps are not equally fast. | Signal | What it says | What it does not say | |---|---|---| | One step blocked most of the time | Its consumer, further along the graph, is the constraint | Which consumer, or why | | Every step blocked | The constraint is at the very end — the write, or the destination | Whether the destination is slow or merely rate-limited | | No step blocked, and the job still behind | The constraint is the read from the source or the source itself | Whether more read parallelism is available | | Resistance flapping on and off | The cause is intermittent | Whether it is a heavy piece, a pause, or a periodic consumer | ## Where the number comes from, and where there is none This is the part candidates most often get wrong by generalising from the one engine they have used. Engines in this family disagree: 1. Some runtimes instrument the handover directly and publish a blocked-time fraction or a busy/idle split per step. 2. Others publish only the occupancy of the in-job buffers, or the per-step processed throughput, and you read resistance off those. 3. In **the two-phase disk-to-disk model** — where every phase writes its whole output to shared storage before the next phase reads it — there is nothing in flight between phases at all, so there is no producer to block and no resistance to report. Slowness there shows up as one phase's elapsed time, not as a blocked step. So the right sentence in an interview is "the figure is called different things and some runtimes do not publish it; what I am looking for is a step that is idle because its consumer will not take more" — the mechanism, not the screen. ## Reading it in practice When someone says the job is slow and one step shows heavy resistance, the honest first answer is an ordered list: confirm the resistance is sustained rather than flapping; walk away from the source to the first step that is *not* blocked; and only then ask what that step is doing with its time. Everything before that last question is locating; everything after it is a different subject, with a different owner.

  • Can a job show no resistance anywhere and still be the reason a downstream consumer is complaining?
    Yes. Resistance only appears where a step has somewhere to push back to, so if the constraint is the very first step — the read from the source, or the source supplying slowly — nothing ahead of it blocks and every figure looks calm. The same is true in a model where each phase materialises its whole output before the next begins: there is no in-flight handover to block.
  • What is the difference between a step reporting resistance and a step simply being slow?
    A resisting step is idle: it has produced a record and is waiting for somewhere to put it. A slow step is busy: it is spending its time computing, reading or waiting on an external call. The bottleneck is always the busy one, which is why the useful reading is the boundary between the two.
  • Is sustained resistance a reason to raise an alert?
    Usually not on its own. Any graph whose steps are not equally fast settles into a state where the earlier ones block, so resistance is the normal shape of a healthy job. What is worth alerting on is resistance that has changed — appeared where there was none, or moved to a different step — and the consumer-facing promise the job exists to keep.

saying these in an interview costs you the question

  • Says the resisting step is the slow one and should be widened
  • Treats any reported resistance as an error a healthy job never shows
  • Reads a blocked step as overloaded rather than idle and waiting
  • Assumes every engine publishes the same per-step blocked-time figure
  • Concludes the input arrives too fast without looking further along the graph
open as a page

Why can the diagnostic lines a worker process printed be unrecoverable once a job run has ended and its machines are released?

level: juniorimportance: must knowfreq 70%

basics

~10 s

A worker's diagnostic lines are written on the machine that ran it, and the cluster releases that machine when the work ends. Only what the run carried off its machines stays readable.

open as a page

A continuous job reports it is four million records behind its input — why does that number alone not say whether it will recover?

level: juniorimportance: must knowfreq 70%

basics

~20 s

Four million records behind is a level, and recovery depends on the trend. Compare two readings taken minutes apart: a distance that is growing means the job is losing ground, a shrinking one means it is already draining.

open as a page

A teammate says a running job 'feels slow' and gives no other detail. Which numbers do you read first, and why those?

level: juniorimportance: must knowfreq 72%

basics

~20 s

Five numbers cover almost every 'slow' report: processed throughput in records and in bytes per second, duration per step of the job graph, distance behind the input, and failed or retried units of work. Read them before changing anything.

open as a page

A job written for a live feed is replayed over six months of stored input - which of its parts depend on the wall clock?

level: juniorimportance: must knowfreq 58%

basics

~20 s

Anything the code reads from the present: the current time written into output, timeouts and inactivity gaps measured in real seconds, rate thresholds per minute, expiry of held state, and relative date filters. Compressed into one hour, none behaves as it did live.

open as a page

Three steps nearest a job's source are each blocked most of the time handing output on; the fourth is not. Which is the bottleneck?

level: middleimportance: must knowfreq 55%

basics

~20 s

The fourth step. Backpressure propagates backwards towards the source, so the place you first notice resistance is not the place that caused it: the cause is the first step, walking away from the source, that is not itself blocked.

open as a page

Why can an optimisation that halves a run's machine-time bill leave a bytes-read bill untouched?

level: middleimportance: must knowfreq 55%

basics

~20 s

Machine-time billing charges for capacity held multiplied by how long it was held, so it falls when a run finishes sooner or holds fewer machines. Bytes-read billing charges only for input touched, which falls only when the run reads less.

open as a page

A job sits twelve million records behind, processes nine thousand records a second and receives six thousand a second — when is it caught up?

level: middleimportance: must knowfreq 62%

basics

~20 s

Divide the backlog by the gap between the rates, not by throughput. Nine thousand processed against six thousand arriving drains three thousand a second, so twelve million records take about four thousand seconds — a little over an hour.

open as a page

Which signals on a running job move first when it degrades, and which only confirm the damage afterwards?

level: middleimportance: must knowfreq 58%

basics

~20 s

Leading signals move while output still looks fine: bytes spilled to local disk, shrinking memory headroom, retried units of work, and the fraction of a step's time spent unable to pass its output on. Distance behind the input and run duration are lagging — they accumulate the damage.

open as a page

A run fails on one step of the job graph and the error names no record. How do you narrow it to a single input piece?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Start from what the failure names: the step and the failed unit of work. If the same unit fails each attempt, map it back to its input range, then run the step's function over that range in one process.

open as a page

A year of history must be pushed through a job whose output lands in a live serving database - what sets the replay's duration?

level: seniorimportance: must knowfreq 52%

basics

~20 s

The write rate the destination can sustain while still serving its live traffic, not the cluster. Beyond that point extra workers only produce rejected writes, retries and contention. Plan the replay backwards from the destination's spare capacity.

open as a page

Why is a counter that every worker adds to and a coordinating process sums better evidence than a line printed on a worker?

level: middleimportance: should knowfreq 58%

basics

~20 s

A carried counter is summed by the one process that outlives the workers and is normally written out with the run's figures, so the evidence survives the machine that produced it. Its limit is that it gives a total, never a record.

open as a page

Records per second is unchanged but each step of the job now takes twice as long — which further measurements separate the causes?

level: middleimportance: should knowfreq 55%

basics

~20 s

Read bytes per second beside records per second, plus bytes per record. A flat record rate with doubled step duration usually means fatter records, more input per piece, or slower machines — three different fixes, separated by the volume figures and by the spread across pieces.

open as a page

A replay enriches each stored record with a customer tier read from a table that is overwritten nightly - what is wrong with the output?

level: middleimportance: should knowfreq 50%

basics

~20 s

Every historical record is labelled with today's tier rather than the one in force at the time, so the replay quietly rewrites history. Reproducing the past needs reference data that retains its own change history, looked up as of each record's timestamp.

open as a page

A job's in-memory buffer between two of its steps keeps filling, so a team doubles that buffer. What does this actually change?

level: seniorimportance: should knowfreq 48%

basics

~20 s

It buys time proportional to the added capacity and nothing else. A larger in-job buffer absorbs a longer burst before the earlier step blocks, but adds no consumption capacity, so a sustained deficit refills it and the resistance returns.

open as a page

A job is slow and its runtime publishes no per-step blocked-time figure. Which other signals identify the bottleneck, and when does backpressure not exist at all?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Substitute buffer occupancy, per-step processed throughput and the read rate from the source for the missing figure, and look for the boundary between idle and busy steps. Where each phase materialises its whole output before the next reads it, there is no in-flight handover and no backpressure to find.

open as a page

Two daily runs write outputs of the same size and one costs ten times the other, so which per-run numbers separate the causes?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Four numbers per run: input bytes read, capacity held times duration, bytes moved between workers, and output bytes written. With equal outputs, a tenfold gap almost always sits in input read or in capacity held while idle.

open as a page

One standing pool runs sixty pipelines for eight teams and the monthly invoice names only the pool. What must be arranged at submission time?

level: seniorimportance: should knowfreq 45%

basics

~20 s

A run tag — a label carrying owner, pipeline and a unique run identifier, attached when the work is submitted — plus per-run usage captured while the run still exists. Without both, an untagged run leaves a remainder nobody can claim.

open as a page

A step rejects a small fraction of records in every run. What do you arrange in advance so the rejected ones can be examined?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Count every rejection, but keep only a bounded sample of the records themselves: a fixed cap per worker process per run, written to shared storage under the run's identity, with the reason attached and sensitive fields left out.

open as a page

A job that fell a day behind is now draining that backlog at full throughput — what does that do to the consumers of its output?

level: seniorimportance: should knowfreq 45%

basics

~20 s

They receive the normal arrival rate plus the drain rate, sustained for as long as the drain lasts. It is a plateau rather than a spike, and it is the load nobody sized for: throttled destinations, alarms keyed to volume, and consumers that fall behind in turn.

open as a page

A job's processed rate has sat below its input rate for a week with no overnight recovery — what kinds of cause produce a permanent deficit?

level: seniorimportance: should knowfreq 55%

basics

~20 s

A permanent deficit has four families of cause: arrivals rose, the work per record rose, the job's own ceiling is fixed and has been reached, or the rate is set outside the job by a destination, a shared pool or capacity lost to repeated re-work.

open as a page

Every health check is green and the process is up, yet consumers say the numbers stopped updating hours ago — what should the alarm have measured?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Liveness is not progress. The useful alarm is a threshold on what the job emits — the freshness and volume of its own output, or its distance behind the input — evaluated by something outside the job, so that silence itself trips the alarm.

open as a page

A job run finished successfully but retried four hundred times as many units of work as usual — why investigate?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A green run hides its retries: the engine re-ran failed units until they passed, so success proves only that the allowance was large enough. The count names a machine, an input piece or a memory limit that is failing today and will exhaust that allowance soon.

open as a page

Six months of stored records are pushed through a job that groups by event time in two hours - how do the results differ from the live run's?

level: seniorimportance: should knowfreq 42%

basics

~20 s

The job's running claim about how far time has advanced is derived from the data, so it races through six months in seconds. Groups close back to back, and the replay can end up either more complete than the live run or less, depending on the order history is read in.

open as a page

How would you prove a corrected job is right over a year of history before any consumer sees its output?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Run it beside the existing output: the same stored input, the same period, written to a side destination nobody reads, then compared against what the live job already produced. The comparison, not the run, is the deliverable.

open as a page

Which evidence should a team carry on every job run, and which should it capture only on a deliberate re-run?

level: principalimportance: should knowfreq 40%

basics

~20 s

Carry always what is small and impossible to reconstruct afterwards: a few totals per step, the run's identity and shape, per-unit status. Raise the expensive evidence on demand, and only where a re-run can genuinely recreate the failure.

open as a page

A daily job has produced the same dataset for three years and nothing reads it, so what makes retiring it a judgment call?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Each day the job runs charges the pool for compute nobody uses, but stopping it risks a consumer nothing has recorded. The call weighs a known recurring bill against an unknown blast radius, and staging it is what makes it safe.

open as a page