skip to content

A job's processed rate has sat below its input rate for a week with no overnight recovery — what kinds of cause produce a permanent deficit?

level: seniorimportance: should knowfreq 55%

answer

  1. which side of the subtraction moved
  2. demand up, or work per record up
  3. a fixed ceiling ignores extra machines
  4. the busiest key bounds a grouped step
  5. re-work consumes throughput without progress

basics

~20 s

A permanent deficit has four families of cause: arrivals rose, the work per record rose, the job's own ceiling is fixed and has been reached, or the rate is set outside the job by a destination, a shared pool or capacity lost to repeated re-work.

solid answer

~50 s

Separate the two sides of the subtraction. On the **demand** side, the input rate has risen past what this job's current shape can process, or the quiet period it used to recover in has disappeared. On the **supply** side there are three shapes. The work per record grew — an added lookup, a larger payload, more state held per key, so the same parallelism yields fewer records a second. The job's ceiling is fixed: where the width of a continuous job is set for its lifetime, or where a key-grouped step's throughput is bounded by its busiest key, more machines change nothing. Or the rate is being set outside the job — a destination that throttles, a shared pool giving this run less than it did, or capacity consumed by repeatedly re-running failed units. Each leaves different evidence, so rule them out in that order rather than guessing.

go deeper

for a junior

Know the two sides: either more is arriving than before, or each record costs more to process than before. Recognise that a deficit surviving the overnight quiet period is the serious kind.

for a middle

Explain how each cause shows up differently — records against bytes per second, worker spread, per-step duration — and why more machines cannot lift a ceiling set by one key or by a throttling destination.

for a senior

Work the causes in order of cheapness to rule out, name which execution model you are describing, and be explicit about whether adding capacity means more processes or more machines.

for a principal

Frame it as a promise rather than a dashboard. A permanent deficit means the freshness the team committed to consumers is no longer deliverable at the current shape, so the real decision is whether to fund capacity, cut the work per record, or renegotiate the promise — and to say which, publicly, rather than letting the queue decide.

## What "permanent" means here A deficit is permanent when the processed rate stays below the input rate across a full cycle of the input's own rhythm — through the overnight trough, through the weekend — so there is no period in which the job recovers what it lost. The distinction matters because the usual pattern is benign: many jobs run a deficit for three hours every afternoon and clear it by midnight. That is capacity sized to the average with a peak absorbed by the queue, which is a design, not a fault. A week of unrecovered ground is something else. ## Where a permanent deficit comes from There are only two sides to the subtraction, so causes sort cleanly. **Demand rose.** - The input rate grew — more producers, a new event added to the same stream, a partner onboarded — and crossed the line the job's current shape could hold. - The trough disappeared. Nothing about the job changed; the overnight quiet period it used to drain in got smaller until it no longer covered the daily peak. **The work per record rose.** - A step was added: an enrichment, a validation, a lookup against an external system whose own latency now bounds the job. - Records got bigger, or a nested field grew, so the same records-per-second is several times the bytes per second it was. - State held per key grew, so each access to it costs more — a per-record cost that rises slowly and is easy to miss because no code changed this week. **The job's own ceiling is fixed and has been reached.** - On engines where a continuous job's width is fixed for the life of the job, throughput cannot rise while it runs; changing the width means a restart with redistribution. On others the width can move, and then the ceiling is elsewhere. - A step grouped by key is bounded by its busiest key: when one grouping key holds most of the records, one worker process — one process on one machine that runs pieces of the job and owns the memory they use — sets the ceiling while the rest of the cluster is idle. Adding machines buys nothing. - An unavoidably serial point, such as a single writer or a lock around a shared resource, does the same thing. **Something outside the job sets the rate.** - The destination throttles. Then the job's processed rate simply equals the destination's accept rate, however much capacity the cluster has. - A shared pool gives this run fewer machines than it had, because a neighbour's run grew. - Capacity is going into re-work: units of work failing and being retried by the engine inside the run, or pieces recomputed after a worker was lost, both consume throughput that never appears as progress. ## Matching cause to evidence | Cause | What you would expect to see | |---|---| | input rate grew | input rate trend up, processed rate flat at its old value | | work per record rose | records a second down, bytes a second flat or up, per-step duration up | | busiest key sets the ceiling | most workers idle, one busy continuously, wide spread in per-piece duration | | destination throttles | processed rate pinned at a suspiciously round value, write errors or waits | | capacity lost to re-work | processed rate below what the machines should give, failed-and-retried counts rising | | shared pool contention | fewer workers than the run used to get, deficit correlated with a neighbour's schedule | ## What varies with the execution model The same deficit shows up in three different units, and saying which model you mean is part of a correct answer: 1. **A job built from repeated small finite runs** — the engine slices an endless input into small finite jobs run back to back — shows it as slice duration exceeding the period between slices, so each slice starts later than the last and the shortfall accumulates. 2. **A record-at-a-time job**, where each record moves through the graph as it arrives, shows it as a sustained gap in records per second, usually with the in-job buffers between steps full behind whichever step is the bottleneck. 3. **An interval-launched two-phase model**, where each phase writes its whole output to shared storage before the next reads it, shows it as a run that no longer finishes inside the window it was given, so the next launch starts against more input than the last. ## Ruling them out in order Ask which side moved before asking what to change. Compare today's input rate with last month's: if it rose, that is the answer and the rest is remedy. If it did not, the supply side moved, and the cheapest discriminator is the spread across workers — a cluster that is mostly idle points at a ceiling or a throttle, a cluster that is fully busy points at work per record or simply not enough machines. Only then is it worth arguing about what to change, and it is worth stating explicitly whether "more workers" means more processes on the same machines, which fixes nothing when the machines are saturated, or more machines.

  • The cluster is mostly idle yet the job still cannot keep up. What does that narrow it to?
    Something is serialising or throttling the work rather than starving it of machines: one grouping key holding most records so a single worker process sets the ceiling, a single-writer point in the job, or a destination whose accept rate the job has been pinned to. Idle capacity rules out the cause that more machines would fix, which is why worker spread is the cheapest first discriminator.
  • Records per second is unchanged but the job started falling behind. What should you look at?
    Bytes per second and per-record work. A payload that doubled, a nested field that grew, or an added lookup all leave the record count flat while the real work rises. State per key growing over the life of the job does the same thing without any code change, which makes it the easiest of these to miss.
  • How do you tell a permanent deficit from an unusually long peak?
    By sampling across the input's full rhythm rather than the last few hours. If the trough still brings the distance behind back to where it was a week ago, the job is absorbing a peak as designed. If each trough leaves it higher than the last, the deficit is permanent even though the line falls every night.

saying these in an interview costs you the question

  • Concludes more machines will fix any deficit without checking whether workers are idle
  • Calls a daily afternoon deficit permanent without looking across a full cycle
  • Assumes a job's parallel width can always be raised while it keeps running
  • Ignores that a throttling destination pins the job's processed rate from outside
  • Overlooks re-work, so retried units are counted as progress
  • Reads an unchanged records-per-second as proof the work per record is unchanged