skip to content

Your input tripled: the nightly finite job now uses more of the cluster and the continuous job does not — why?

level: seniorimportance: should knowfreq 48%

answer

  1. measured numbers move, stated numbers do not
  2. growth cannot edit a declaration
  3. derived at the read, declared for the rest
  4. check units against lanes first

basics

~20 s

Because the two numbers have different origins. The finite run derives its pieces from the stored bytes, so more data means more pieces and more busy lanes; the continuous job runs at a width the author declared, which no change in the input can move.

solid answer

~50 s

The finite job's width is a *measurement* and the continuous job's is a *declaration*. A **finite job** divides its stored input before it starts, so three times the bytes yields roughly three times the pieces, and the extra pieces occupy lanes that were idle before — up to the lanes that exist. A **continuous job** has no total to measure, so it runs at the **declared operator width** the author stated; tripling the arrival rate gives each existing copy three times as much to do and creates no new copies. Two caveats keep this honest: on runtimes that settle the split count at submission, even the finite job only widens at its next submission, and where continuous work runs as a rapid succession of small finite jobs the *read* side of each small run does widen while the stated width still governs everything after it.

go deeper

for a junior

The takeaway to hold is that one of these numbers is measured from the data and the other is written down by a person. Measurements change when the data changes; written-down numbers do not.

for a middle

Explain the mechanism both ways: more stored bytes means more pieces means more occupied lanes, while more records per second means the same copies each doing more. Then say where the finite job's widening stops — at the lanes available.

for a senior

Demonstrate the caveats: split counts fixed at submission, continuous work implemented as repeated small runs widening only at the read, and a declaration that already exceeds the lanes. Say which evidence you would look at before proposing a change.

for a principal

The fleet-level question is which jobs are allowed to self-size and which carry a declaration someone must revisit. Growth on a declared job is a silent commitment to periodic review; price that review in, or move the workload to a form whose number is derived.

## The same growth, two responses The number that governs how much of a cluster is working is the count of independent units of work: a **piece of the input** is one slice of data that a single **worker thread** takes from start to finish, and the count of those pieces is the ceiling on busy lanes. The observation in the question — one job absorbing the extra capacity and the other not — is the clearest everyday consequence of the fact that this number has two different origins. ## Why the finite job widened A **finite job** is a run over an input that ends, so the stored form can be measured first and cut into pieces: file boundaries, the block boundaries of a **cluster file system** (storage running on the same machines as the compute), or a count the source reports. Triple the stored bytes and, in the ordinary case of many files or many blocks, you get roughly triple the pieces. Pieces are what fill lanes, so the run occupies more of the cluster without anyone touching it — until the piece count passes the number of lanes available, at which point the extra pieces queue and the run simply takes longer in more rounds. That self-widening is worth naming explicitly: nobody made a decision, and the job's parallelism changed anyway, because it was a description of the data all along. ## Why the continuous job did not A **continuous job** is a run over an input with no end. There is no total to measure, so the author supplies a **declared operator width** — how many copies of each operator run. Three times the records per second does not create a fourth copy of an operator that was declared to have three; it gives each of the three copies three times as much to do. The consequences follow mechanically: per-copy throughput becomes the binding constraint, the job's processing falls behind its input, and the cluster may look *less* loaded than you expect because the same number of lanes are simply busier. Raising the declaration is a real option and a separate decision — what value to pick and at which moment it can take effect is owned by the sibling subject on choosing how many pieces — and the work of catching up on what accumulated meanwhile belongs to the operating subject. What this question owns is the reason the two jobs behaved differently at all. ## Side by side | | Finite job | Continuous job | |---|---|---| | Origin of the number | measured from the stored input each time | stated once by the author | | Response to three times the data | roughly three times the pieces | the same copies, each three times busier | | Limit it runs into | the lanes available; beyond that, more rounds | per-copy throughput against arrival rate | | What a human must do | nothing, on runtimes that re-derive | revisit the declaration deliberately | ## Where the rule bends Three caveats stop this becoming a false universal, and naming them is what separates a senior answer from a textbook one: 1. **Finite runs that do not re-derive.** Some runtimes settle the split count when the job is submitted rather than recomputing it per run. There the finite job also fails to widen until something resubmits it, and the observed difference between your two jobs would disappear. 2. **Continuous work as repeated small finite jobs.** A sizeable part of this class implements a continuous run as a rapid succession of small finite runs. There the read side of each small run *is* derived from what arrived, so the reading width does track volume, while the stated width still governs the operators after the read. Such a job partially widens, which is exactly the observation that confuses people who have used only one runtime. 3. **A declaration already above the lanes.** If the declared width exceeds the lanes available, the copies queue rather than run concurrently, and the job was never using its declaration in the first place; adding data changes nothing visible about occupancy because occupancy was already saturated. ## What to check, in order 1. Confirm the finite run's unit count really tripled — if it did not, the input grew inside existing files rather than as new ones, and you have a different problem. 2. Confirm the continuous job's copies are each busier rather than more numerous; that is the signature of a declaration. 3. Decide whether the continuous job needs a larger declaration or the same declaration on faster lanes, remembering that the first is a deliberate change with its own timing rules and the second is a sizing question owned elsewhere. ## The claim to avoid "A streaming system scales itself with load" and "a batch system needs to be re-tuned when data grows" are both folk wisdom, and on this class they are frequently backwards. The defensible statement is the one about origins: a derived number tracks the data by construction, a declared one does not, and which of the two you have depends on whether the input ends — not on which runtime you happen to use.

  • The finite job stopped widening after its unit count passed the lanes available. What happens next?
    The extra units queue. Lanes stay fully occupied and the run works through the units in more rounds, so wall-clock grows roughly in step with the data instead of staying flat. Occupancy is saturated, so from that point the only remedies are more lanes or less work per unit.
  • Why might a continuous job widen partially as data grows?
    Because some runtimes implement continuous work as a rapid succession of small finite runs. Each small run derives its read-side division from what arrived, so reading widens with volume, while the operators after the read keep the stated width. The job therefore tracks growth at the front and not behind it.
  • Does adding machines to the continuous job change how much of it runs at once?
    Not by itself. More machines add lanes, and a declared width says how many copies exist to occupy them, so the extra lanes stay empty until the declaration changes. Where the machines come from, and when capacity can be added, are separate subjects from the width itself.

saying these in an interview costs you the question

  • Says continuous runtimes scale their width with the input rate.
  • Claims a finite job must always be re-tuned when data grows.
  • Assumes adding machines widens a job running at a stated width.
  • Thinks the finite job widens without limit as data grows.
  • Treats the difference as a property of the runtime rather than the input.