A step in a running job spends most of its time unable to hand its output on — what does that indicate?
answer
- a symptom, not a fault
- blocked waiting, not busy working
- the cause lies further along
- a resisting step has spare capacity
- some engines publish no such figure
basics
~20 sBackpressure: a step further along the job graph cannot accept work as fast as this one produces it, so this one blocks. The figure locates a bottleneck ahead of the resisting step, not a fault inside it.
solid answer
~40 sThat is **reported resistance** — the fraction of a step's time spent unable to hand its output on, which some engines publish per step and others leave you to infer. Read it as **backpressure**: a step that cannot accept more work stops the step feeding it, which stops the one feeding that, back towards the source. So a resisting step is transmitting someone else's shortage, not causing one; it is *idle waiting*, not busy, and therefore has spare capacity. Its value is directional — it tells you which way to walk. What it does not tell you is why the constrained step is constrained: uneven work across pieces, a worker process short of memory, and a destination that rate-limits writes all look identical from here.
go deeper
Recall the one sentence: a step reporting that it cannot hand its output on is telling you something after it is slower, not that it is itself the problem. Say the word backpressure and say it is a pointer.
Explain the mechanism — a bounded in-memory queue between two steps, full, so the producer stops, and that stopping propagating back towards the source. Say explicitly that a resisting step is idle and therefore has spare capacity.
Show that you treat the figure as a locator and nothing more, and that you know it is a normal steady state rather than an incident. Then say what you would look at next, and in what order, to turn the location into a cause.
The judgment is what you alarm on. Resistance is a routine steady state, so paging on it produces noise; the defensible position is to alarm on the promise made to consumers and keep resistance as the first diagnostic you open, not the trigger.
## What the figure measures An engine of this class runs a job as a graph of **steps** — one transformation applied to every piece of the input — with each step handing records to the next. Between two adjacent steps sits a **bounded in-job buffer**: a small in-memory queue holding records the earlier step has produced and the later one has not yet taken. *Bounded* is the whole point. When that queue is full, the earlier step has nowhere to put its next record, so it stops. That stopping is **backpressure**: the mechanism by which a step that cannot accept more work halts the step feeding it, which halts the step feeding that, all the way back to the read from the source. **Reported resistance** is the number some runtimes publish for exactly this — the fraction of a step's wall-clock time spent unable to hand its output on. Others publish no such figure and you infer the same thing from buffer occupancy or from per-step throughput. ## Backpressure is a diagnosis, not a fault The single most useful thing to understand about the number is that it is **not an error**. A job can run at exactly the rate its slowest step allows, for months, with several steps reporting high resistance the whole time, and be perfectly healthy: the resistance is simply how the graph settles into one common rate. Nothing is dropped, nothing is retried, no unit of work has failed. What the number does is **locate**. A resisting step is one whose consumer is slower than it is, so the constraint lies somewhere *away from the source* — further along the graph. The resisting step itself is the last place to look for a cause, because by definition it is spending its time idle. ## What it tells you - **Direction.** The bottleneck is ahead of the resisting step, in the direction the records flow. - **Spare capacity.** A step blocked 80% of the time could do roughly five times the work if something accepted it; widening *that* step buys nothing. - **A steady state.** Resistance that is high and flat means the job has already settled to the constrained step's rate. Resistance that appears and clears means an intermittent cause. ## What it does not tell you - **Why** the constrained step is constrained. Uneven work across pieces, a worker process running short of memory, an expensive per-record computation, and a destination that throttles writes all present here as a downstream step that accepts slowly. - **Whether the job is behind.** Resistance is about the relationship between two steps; how far the job sits behind its input — lag in the sense of unread input — is a separate number, and a job can resist heavily while still keeping up with everything that arrives. - **That anything is broken.** See above: it is the normal steady state of any graph whose steps are not equally fast. | Signal | What it says | What it does not say | |---|---|---| | One step blocked most of the time | Its consumer, further along the graph, is the constraint | Which consumer, or why | | Every step blocked | The constraint is at the very end — the write, or the destination | Whether the destination is slow or merely rate-limited | | No step blocked, and the job still behind | The constraint is the read from the source or the source itself | Whether more read parallelism is available | | Resistance flapping on and off | The cause is intermittent | Whether it is a heavy piece, a pause, or a periodic consumer | ## Where the number comes from, and where there is none This is the part candidates most often get wrong by generalising from the one engine they have used. Engines in this family disagree: 1. Some runtimes instrument the handover directly and publish a blocked-time fraction or a busy/idle split per step. 2. Others publish only the occupancy of the in-job buffers, or the per-step processed throughput, and you read resistance off those. 3. In **the two-phase disk-to-disk model** — where every phase writes its whole output to shared storage before the next phase reads it — there is nothing in flight between phases at all, so there is no producer to block and no resistance to report. Slowness there shows up as one phase's elapsed time, not as a blocked step. So the right sentence in an interview is "the figure is called different things and some runtimes do not publish it; what I am looking for is a step that is idle because its consumer will not take more" — the mechanism, not the screen. ## Reading it in practice When someone says the job is slow and one step shows heavy resistance, the honest first answer is an ordered list: confirm the resistance is sustained rather than flapping; walk away from the source to the first step that is *not* blocked; and only then ask what that step is doing with its time. Everything before that last question is locating; everything after it is a different subject, with a different owner.
- Can a job show no resistance anywhere and still be the reason a downstream consumer is complaining?Yes. Resistance only appears where a step has somewhere to push back to, so if the constraint is the very first step — the read from the source, or the source supplying slowly — nothing ahead of it blocks and every figure looks calm. The same is true in a model where each phase materialises its whole output before the next begins: there is no in-flight handover to block.
- What is the difference between a step reporting resistance and a step simply being slow?A resisting step is idle: it has produced a record and is waiting for somewhere to put it. A slow step is busy: it is spending its time computing, reading or waiting on an external call. The bottleneck is always the busy one, which is why the useful reading is the boundary between the two.
- Is sustained resistance a reason to raise an alert?Usually not on its own. Any graph whose steps are not equally fast settles into a state where the earlier ones block, so resistance is the normal shape of a healthy job. What is worth alerting on is resistance that has changed — appeared where there was none, or moved to a different step — and the consumer-facing promise the job exists to keep.
saying these in an interview costs you the question
- Says the resisting step is the slow one and should be widened
- Treats any reported resistance as an error a healthy job never shows
- Reads a blocked step as overloaded rather than idle and waiting
- Assumes every engine publishes the same per-step blocked-time figure
- Concludes the input arrives too fast without looking further along the graph