skip to content

A job is slow and its runtime publishes no per-step blocked-time figure. Which other signals identify the bottleneck, and when does backpressure not exist at all?

level: seniorimportance: should knowfreq 40%

answer

  1. the mechanism outlives the dashboard
  2. full queues behind, empty after
  3. equal rates prove nothing alone
  4. discard the destination and re-measure
  5. materialised phases cannot resist

basics

~20 s

Substitute buffer occupancy, per-step processed throughput and the read rate from the source for the missing figure, and look for the boundary between idle and busy steps. Where each phase materialises its whole output before the next reads it, there is no in-flight handover and no backpressure to find.

solid answer

~50 s

Backpressure — a step that cannot accept work halting the step feeding it — only exists where two steps hand records to each other while both are live. The evidence for it, if the runtime does not publish a blocked fraction, comes from three substitutes: **buffer occupancy** per handover, if exposed (full behind the constraint, empty after it); **processed throughput per step in records per second**, read together with whether the step is computing or waiting; and the **read rate from the source** against the rate at which input arrives, since a throttled read is the far end of the same propagation. A controlled experiment beats all three: run the same job writing to a discard destination, or over a fixed finite input, and see whether the rate moves. And in a model where every phase writes its whole output to shared storage before the next phase reads it, nothing is in flight between phases, so the bottleneck is simply the phase with the longest elapsed time.

go deeper

for a junior

The takeaway is that the mechanism is not the screen. A step stops because the next one will not take its output, and that is true whether or not your tooling puts a number on it.

for a middle

Explain the two workable substitutes — queue occupancy showing full-then-empty across the constraint, and per-step rates read alongside busy time — and why the rates alone are the same everywhere in a settled graph.

for a senior

Show the experiment as well as the observation: a discard destination or a fixed finite input turns an inference into a measurement. Then name a model in which there is no in-flight handover to resist at all.

for a principal

The question behind the question is what you should have arranged in advance. If a job's runtime publishes nothing useful, the durable answer is to emit the handful of numbers the team will actually need from inside the job, and to decide what is not worth measuring.

## Why the figure may be missing The engines in this family instrument themselves differently, and a candidate who has only used one tends to assume the whole class publishes the same thing. Some runtimes measure the handover between steps directly and publish a blocked-time fraction or a busy/idle split per step. Others publish only per-step throughput and the occupancy of the queues between steps. Some publish neither, and you are left with elapsed times and resource counters. None of that changes the mechanism — **backpressure**, a step that cannot accept records stopping the step that feeds it — only the evidence you have for it. ## Substitutes, in the order worth trying 1. **Buffer occupancy per handover.** This is the most direct proxy. Behind the constraint the queues are full; after it they are empty, because the constrained step consumes everything it is given and the steps after it are starved. The boundary between a full queue and an empty one is the constrained step. 2. **Processed throughput per step, in records per second.** In a settled graph every step processes at the same rate, so the numbers alone prove nothing — this is the trap. What distinguishes the constraint is that it achieves that rate while working continuously, whereas its predecessors achieve it while idle much of the time. Pair the rate with any available busy or CPU figure and the boundary appears. 3. **The read rate from the source against the rate at which input arrives.** A throttled read is the last link in the propagation chain. If the job is reading more slowly than input is arriving, and the read step is not itself the expensive one, something further along is holding it back. 4. **Elapsed time per step or phase.** In a model where steps do not overlap, this is not a proxy at all — it is the answer. ## The experiment that settles it Signals are inference; a controlled change is evidence. Two cheap ones: - **Replace the destination with a discard.** Run the same logic writing nowhere. If the rate jumps, the destination or the write step was the constraint; if it does not move, the constraint is inside the graph. - **Run over a fixed finite input.** This removes any variation in what arrives and makes two runs comparable, which is what lets you attribute a change to the thing you altered rather than to the day. Neither is free — a discard run still consumes cluster capacity, and a finite input cannot reproduce everything about a continuous one — but both convert an argument into a measurement. ## Where the question does not apply This is the part that catches people who generalise from one engine. | Execution model | Is there in-flight resistance? | What you read instead | |---|---|---| | Record-at-a-time processing with key-bound state | Yes. Records cross between live steps continuously, so a slow step blocks its producer directly. | Blocked fraction or buffer occupancy per step. | | A continuous job built from repeated small finite runs — the engine slices an endless input into small finite jobs run back to back | Partly. Resistance inside one slice is brief and often not surfaced. | Each slice's duration against the period it covers, and which step dominates that duration. | | The two-phase disk-to-disk model — every phase writes its whole output to shared storage before the next phase reads it | No. Nothing is in flight between phases, so there is no producer to block. | The elapsed time of each phase, and how long the slowest piece within it took. | Saying this out loud is worth more in an interview than any list of signals, because it shows you are describing the class rather than the one engine you happen to know. The generic sentence is: *backpressure is a property of a live handover between two concurrently running steps; where the handover is a write to shared storage that completes before the reader starts, there is no handover to resist.* ## One more place it does not appear Even in a live graph, there is no resistance to see when the constraint is the **first** step. A step only blocks when it has somewhere to push back to, so a slow read from the source — or a source that simply supplies slowly — produces a calm graph and a job that is nonetheless not keeping up. That is why the read rate belongs on the list of substitutes and why a completely quiet set of figures is never, by itself, evidence of health. ## Putting it together The answer an interviewer is listening for has three parts: name the mechanism independently of any screen; name two or three substitutes and say what pattern in them locates the constraint; and name at least one model in which the question dissolves because there is no live handover. A candidate who gives only the first engine's dashboard has answered for one product and not for the class.

  • Every step reports the same processed throughput in records per second. What have you learned?
    Almost nothing on its own — that is the expected steady state of any settled graph, because each step can only process what reaches it. The information is in how each step spends its time at that rate: the constraint is working continuously to achieve it while the steps behind it are idle much of the time. Rates locate nothing without an idle-versus-busy reading alongside them.
  • Resistance appears and clears every few minutes. What does that rule in?
    An intermittent cause at the first step that is not itself resisting. Candidates: a consumer that rate-limits in windows and then releases, a periodic pause in that step for a flush or an internal reorganisation, or an occasional piece of input far heavier than the rest. The period itself is the clue — match it against the cycles of the systems involved before opening anything else.
  • Why is running the job against a discard destination not a complete test?
    Because it changes more than one thing. Removing the write removes serialisation, any ordering or transactional work the write implied, and any waiting on acknowledgements, so a rate that jumps tells you the constraint was somewhere in that bundle rather than naming a part of it. It is a fast way to decide whether to look inside or outside the job, not a diagnosis.

saying these in an interview costs you the question

  • Assumes every engine of this class exposes the same blocked-time figure
  • Reads equal per-step throughput as proof there is no constraint
  • Looks for in-flight resistance in a model where phases materialise to shared storage
  • Treats quiet figures across the graph as evidence the job is healthy
  • Never considers changing the destination to isolate the constraint