skip to content

Your batch jobs lean on runtime re-planning around uneven pieces, and the same logic must now run continuously over an endless input - what coverage is lost, and how do you plan around it?

level: principalimportance: nice to knowfreq 28%

answer

  1. the correction needs completions
  2. an endless input never completes one
  3. small finite jobs are the partial exception
  4. width and key persist across them
  5. size for the worst key

basics

~20 s

Nearly all of it. Re-planning needs a finished step's measured output, and an endless input has none. Uneven-work decisions move to design time and to the author, and changing one means a restart rather than a mid-run correction.

solid answer

~60 s

Runtime replanning - the engine re-deciding unrun steps from what a finished step actually produced - runs on a supply of completed measurements, and a job whose input never ends produces none. Two designs differ here and the difference is the whole answer. Where continuous work is executed as a rapid succession of small finite jobs, each small job does have steps that finish and sizes that can be measured, so corrections can fire inside one of them; but the job's width and the key-to-destination rule persist across all of them, so a key that is heavy is heavy in every one. Where the runtime is record-at-a-time, records are pushed onward as they are produced and nothing is materialised, so there is no measurement at all. Plan around it by moving the decisions earlier: choose width and key before starting, size for the worst key rather than the average, watch per-piece backlog instead of a completed-step report, and treat any change of width or key as a planned restart.

go deeper

for a junior

Recall the dependency: this correction is fed by steps that finish, so a job whose input never ends has nothing to feed it.

for a middle

Explain why the evidence is missing rather than merely switched off, and what replaces a completed-step report as the signal - per-piece backlog, read as a distribution.

for a senior

Distinguish the two continuous designs and name what persists across small finite jobs - width, the key-to-destination rule, key-bound state - so the same key is heavy on every pass.

for a principal

Frame it as a trade the team is making: results that track the input, paid for with an automatic remedy you no longer have, capacity sized for the heaviest key, and shape changes scheduled as restarts.

## Why the correction disappears **Runtime replanning** is the engine re-deciding part of the plan it has not run yet, using the sizes a finished step actually produced. Every word of that depends on completion: a step ends, its output is countable, and a later step has not started. A job over an endless input breaks the first link. There is no last record, so no step finishes in the sense the mechanism needs, and therefore no measured output to re-plan from. This is the leaf's sharpest consequence, and it is the one candidates most often get wrong in the optimistic direction - assuming the runtime notices a heavy piece and deals with it, when in the continuous case it largely cannot. ## But say which continuous design you mean The engines in this class genuinely disagree, and an answer that treats 'continuous' as one thing is wrong about half the market. | design | is there a finished step to measure? | what re-planning can do | |---|---|---| | continuous work run as a rapid succession of small finite jobs | yes, inside each small job | the same corrections can fire within one small job, on that small job's own measured sizes | | a record-at-a-time runtime pushing records onward as produced | no; nothing is materialised between steps | essentially nothing; there is no measurement to act on | Even on the first design the help is thinner than it looks. What persists across the small jobs is what matters for uneven work: the width the job runs at, the key-to-destination rule, and any long-lived state bound to a key. A heavy key is therefore heavy in every single small job, and a correction rediscovered and reapplied each time is not the same thing as fixing the cause. On the second design the situation is starker - the reason its recovery, its latency and its handling of uneven work are all different subjects is that it never materialises an intermediate result at all. ## What moves, and where it moves to Losing the correction does not lose the problem; it relocates the decisions. Three of them move earlier: 1. **Width and shape become design-time commitments.** In a continuous job, the number of parallel instances is usually fixed for the life of the job, and changing it means stopping and restarting with the state redistributed - not a mid-run adjustment. You choose once, in advance, with no measurement. 2. **Sizing is for the worst key, not the average.** With no correction to lean on, the instance that owns the heaviest key sets the job's throughput. Capacity has to be provisioned for that instance rather than for the mean, which is a real and permanent cost. 3. **The key itself becomes the leverage.** If unevenness is structural in the data, the durable fixes are upstream or in the job's own design - and they are the author's, applied deliberately, rather than something the runtime will discover. ## How you see the problem at all The evidence changes too. In a finite job you read what each completed unit produced. In a continuous job nothing completes, so the signal is **backlog** - how far behind the newest available record each piece currently is - watched per piece rather than in aggregate. A mean backlog hides the one instance that is falling permanently behind, which is exactly the instance a heavy key creates. Reading that distribution is its own skill and its own subject; the point here is that your evidence is a live per-piece measure instead of a completed-step report. ## The argument to make to a team When a team proposes moving established finite work onto an endless input, the honest framing is a trade: - **What you buy**: results that track the input instead of arriving on a cadence. - **What you pay**: an automatic correction for uneven work that was quietly doing real operational work for you, and the ability to change the job's shape without stopping it. - **What that implies**: unevenness must be handled explicitly and in advance, capacity is provisioned for the worst key, and any change of width or key is scheduled as a restart from a recovery point rather than absorbed while running. There is no universally right answer - that is why this is a judgment call rather than a fact. The failure mode to avoid is deciding on latency alone and discovering afterwards that the uneven key which was invisible in the finite version, because the runtime kept absorbing it, is now a permanent ceiling on throughput. ## What a strong answer sounds like It states the mechanism ('the correction is fed by completed measurements'), it distinguishes the two continuous designs instead of generalising from one, it names what persists across small jobs, and it converts the loss into concrete commitments - width chosen up front, capacity sized for the heaviest key, changes scheduled as restarts, evidence read as per-piece backlog. A weak answer either asserts that the runtime will handle it anyway, or asserts that continuous jobs get no help at all - both are overstatements of a model that is genuinely split.

  • If each small finite job does have measurable steps, why is the help still thin?
    Because what persists is what hurts. The job's width, the key-to-destination rule and any state bound to a key carry across every small job, so the same key is heavy every time. A correction rediscovered and reapplied on each pass keeps the job moving but never addresses the cause, and it cannot change the parts that are fixed for the job's life.
  • Why can you not simply widen the job when one instance falls behind?
    Changing the width changes which instance owns which key, so any state bound to those keys has to be redistributed. That is a stop-and-resume operation from a recovery point rather than an adjustment while running, and on some engines a width change is not supported without one.
  • What is the evidence you watch instead of a completed-step report?
    Per-piece backlog - how far behind the newest available record each piece is, in records or in time. Watch the distribution rather than the mean: one instance drifting steadily further behind while the rest keep up is the continuous signature of an uneven key.

saying these in an interview costs you the question

  • Expects the runtime to keep re-planning once the input never ends
  • Thinks widening a continuous job mid-run is a free adjustment
  • Says a continuous job's key-to-destination rule can change without stopping the job
  • Treats a correction reapplied per small job as fixing a permanently heavy key
  • Assumes an uneven key stops mattering once the job is continuous
  • Sizes capacity from average backlog rather than the worst piece