skip to content

In a continuous job over an endless input, what stands in for per-piece duration when you are looking for uneven work?

level: seniorimportance: should knowfreq 44%

answer

  1. nothing may finish, so nothing to time
  2. how far behind, per piece
  3. spread across pieces, not the total
  4. maximum backlog, never the mean

basics

~20 s

Per-piece backlog — how far behind the newest record each piece is — read as a spread across pieces. A few pieces falling further behind while the rest sit near zero is uneven work; all of them drifting together is not.

solid answer

~50 s

In a finite job I compare per-piece durations, but a job whose input never ends may have no finished unit to time. The substitute is **backlog per piece**: how far behind the newest available record each piece currently is, counted in records or in time, read as a *spread* across pieces rather than as one total. One or two pieces climbing while the rest sit near zero is the uneven-work signature; every piece climbing together is a capacity statement about the job as a whole, which is a different problem. The important caveat is that runtimes differ here. Where continuous work is executed as a rapid succession of small finite jobs, each of those small jobs *does* finish and does report per-piece durations, so the imbalance shows as the same piece being slowest batch after batch. Where the runtime is record-at-a-time, no unit finishes and backlog is what you have.

go deeper

for a junior

Remember that in a job whose input never ends there may be no finished unit to time, so the question becomes how far behind each piece is rather than how long each one took.

for a middle

Explain what backlog measures and why the spread across pieces, not the total, is the reading — a sum climbs whether one piece or all of them are behind.

for a senior

Demonstrate the distinction that decides the conversation: a few pieces diverging is uneven work, every piece climbing together is capacity, and the two lead to entirely different investigations.

for a principal

Take the position seriously that the reading must exist before the incident. If the only backlog figure a team has is aggregated across pieces, no amount of skill during the incident recovers the per-piece picture.

## Why the finite-job reading may not be available The diagnosis in a finite job is a comparison of finished things: each piece ran for some time and read some amount of input, and the spread across those numbers says whether the work was even. A **continuous job over an endless input** — a job whose input never ends — may offer none of that, because a unit that never completes has no duration to report. But the class genuinely splits here, and assuming one model is the single most common mistake on this subject: - **Record-at-a-time runtimes** keep each piece alive for the life of the job, processing records as they arrive. There is no completed unit, no final duration, and no finished output size. Per-piece durations simply do not exist. - **Runtimes that execute continuous work as a rapid succession of small finite jobs** do finish something, repeatedly. Each small job has pieces that start and end, with durations and input sizes, covering the few seconds of data in that batch. The per-piece reading is available — it is just tiny, and it repeats. So the honest answer names both: what replaces duration where nothing finishes, and what the finite-job reading looks like where things finish constantly. ## Backlog, per piece **Backlog** is how far behind the newest available record one piece currently is, counted either in records or in time. It is the closest available analogue of "this piece has more work than its peers", because a piece that is handed more than it can process accumulates exactly that difference. The reading is the **spread across pieces**, and the reference numbers mirror the finite case: - the **median** piece's backlog — is the job broadly current? - the **maximum** piece's backlog — is any part of the result stale? - the trend of each over time — flat, climbing, or draining ## Read the spread, not the total | what the per-piece backlogs look like | reading | |---|---| | one or a few climbing, the rest near zero and flat | uneven work — those pieces are receiving more than their share, or their machines are unwell | | every piece climbing at a similar rate | the job as a whole is consuming more slowly than its input arrives; this is capacity, not imbalance | | every piece flat but all of them high | the job is keeping pace at a constant offset behind the source; unpleasant, but not diverging | | one piece flat and high, the rest near zero | that piece fell behind during an earlier event and has stopped losing ground without recovering it | The line that matters is the first against the second, and it is the reason a single aggregated figure is nearly useless. A sum of backlogs across all pieces climbs in both cases. A mean lets ninety-nine current pieces mask one that is hours behind. The output consumers see is late precisely for the keys that one piece owns, so the **maximum** per-piece backlog is the number that corresponds to what anyone actually experiences; the sum answers a different question, namely how much work is outstanding in total. ## When the source exposes no position Backlog needs two things to subtract: where the piece has read up to, and where the newest record is. A log-shaped source exposes both per piece. Other sources expose neither, and then two substitutes carry most of the weight: 1. **Per-piece processing rate** — records handled per second over a common interval. A piece steadily below its peers on the same kind of records is the same signal in a different coordinate. 2. **The age of the newest record each piece has handled** — the gap between that record's own moment and now. A piece whose gap keeps widening while its peers stay current is behind, with no position arithmetic needed. ## What the repeated-small-jobs case looks like Where continuous work runs as a rapid succession of small finite jobs, uneven work has a distinctive signature: the same piece is the slowest in batch after batch, and each batch's wall time tracks that piece rather than the others. Because a batch does not complete until its last piece completes, a persistent heavy piece sets the batch time, and the backlog then grows as a *consequence* of that rather than as the primary reading. Both readings are available in this model, and they agree; it is worth checking which one the runtime in front of you actually reports before designing a diagnosis around either. ## What the backlog spread does not settle It shows that the work is uneven and which pieces carry it. It does not say whether the cause is the data or the machine, and the evidence separating those is a different reading again. It also says nothing about whether the imbalance is fixable within a running job: in many continuous jobs the number of pieces is fixed for the life of the job, so the reading may be telling you about a shape that cannot be changed without a restart. The reading remains worth having — a job that is uniformly behind and a job with one piece hours behind look identical on a single aggregated figure and call for completely different conversations.

  • The source exposes no per-piece position — what do you read instead?
    Two things the job can report about itself: the records each piece processed per second over a common interval, and the age of the newest record each piece has handled. A piece whose record age keeps widening while its peers stay current is behind, without any position arithmetic.
  • Why watch the maximum per-piece backlog rather than the sum?
    Because a sum or a mean lets a hundred current pieces mask one that is hours behind, and the output is late exactly for the keys that piece owns. The sum answers how much work is outstanding overall; the maximum answers whether any part of the result is stale.
  • Every piece's backlog is flat but large. Is that uneven work?
    No. A flat backlog means each piece is keeping pace with its share of arrivals while sitting at a constant offset behind the source — usually a startup gap or an earlier incident that was never recovered. Uneven work shows as divergence between pieces, not as a shared constant offset.

saying these in an interview costs you the question

  • Watches one aggregated backlog number for the entire job.
  • Averages backlog across pieces and calls the job healthy.
  • Assumes no runtime ever finishes a unit in a continuous job.
  • Reads a backlog rising on every piece as uneven work.
  • Treats a flat but large backlog as the same symptom as a growing one.