How does a finite job arrive at its piece count, and a continuous job at its declared operator width?
answer
- one number, two origins
- measured against stated
- no end means nothing to measure
- derivation reads the stored form
basics
~20 sA finite job derives it: the stored input is measured and cut into pieces before the run. A continuous input has no end to measure, so the author declares how many copies of each operator run.
solid answer
~50 sThe same number has two completely different origins. In a **finite job** — a run over an input that ends — the stored form can be measured first, so the run derives its pieces from it: file boundaries, the block boundaries of a cluster file system, or a count the source itself reports. If that form cannot be divided, the derivation honestly yields one piece. In a **continuous job** — a run over an input with no end — there is nothing to measure in advance, so the author states a **declared operator width**: how many copies of each operator run. Runtimes differ underneath: some re-derive the finite division on every run, others fix the split count when the job is submitted, and where continuous work runs as a rapid succession of small finite jobs the read side of each small job is derived while the stated width still governs the rest.
go deeper
Hold on to the two-case shape: an input that ends can be measured, an input with no end cannot. That alone answers the question, before any detail about how the measuring is done.
Explain the derivation concretely — file boundaries, block boundaries, or a count the source reports — and say what it reports when the stored form will not divide. Then state plainly that a declared width is an input to the run, not a discovery about the data.
Show that you know the derivation varies: re-derived per run on some runtimes, fixed at submission on others, and hybrid where continuous work is repeated small finite runs. That is what tells an interviewer you have used more than one system of this class.
The consequence worth owning is operational: derived numbers track data growth for free, declared ones become a standing review item. Decide which of your fleet's jobs are allowed to self-size and which need a periodic look at their declaration.
## One number, two origins Every run of this class has a number that says how many units of work exist at once: the **piece count**. A **piece of the input** is one slice of the data that a single **worker thread** reads and processes from start to finish, and that number is the ceiling on how many lanes can be busy. What this question is about is where the number comes from — and the honest answer is that it comes from two entirely different places depending on whether the input ends. ## The finite side: derived from the stored form A **finite job** is a run over an input that ends, so the engine can look at the stored data before any work starts and cut it into pieces. The derivation reads the stored form, not the logical dataset: - **file boundaries** — a directory of files gives at least one piece per file; - **block boundaries** — in a **cluster file system**, storage running on the same machines as the compute, a large file is already stored in fixed-size blocks that make natural cut points; - **a count the source reports** — some sources are not files at all and simply state how many independent units they can be read in. The derivation is arithmetic over what is there. It is not a wish: if the stored form cannot be decoded from the middle, the honest derivation is one piece for the whole thing, however large. *Why* a stored form resists division is the subject of the sibling leaf on the layout a reader inherits; what matters here is that the derivation reports it faithfully. ## The continuous side: declared, not derived A **continuous job** is a run over an input with no end. There is no total byte count, because the total is not written yet, and there is no last record to measure back from. So nothing can be derived, and the author states the number instead: a **declared operator width**, meaning how many copies of each operator run side by side. The width is an input to the run, not an output of it. That is the substantive difference. In the finite case the number is a *fact about the data* that the run discovers; in the continuous case it is a *statement by the author* that the data cannot contradict. A source that suddenly carries ten times the records does not widen a declared operator; it simply gives each existing copy more to do. ## Side by side | | Finite job | Continuous job | |---|---|---| | Where the number comes from | measured from the stored input before the run | stated by the author before the run | | What moves it on its own | more or less stored data | nothing about the input | | What it is measured against | the worker threads that exist | the worker threads that exist | | The failure that follows | an undividable form yields one piece and one busy lane | a width far below or far above the available lanes | ## Where runtimes genuinely differ This is the part candidates over-generalise from whichever engine they learned first: 1. **Re-derived, or fixed at submission.** Some runtimes recompute the finite division on every run, so the count follows the data automatically. Others settle the split count when the job is submitted, so the same stored input divides the same way until something resubmits it. 2. **Continuous as repeated small finite jobs.** A sizeable part of this market implements continuous work as a rapid succession of small finite runs. There, the *read* side of each small run is derived from the bytes that arrived, even though the job as a whole is continuous — and the stated width still governs the operators after the read. 3. **Continuous as record-at-a-time.** Where the runtime pushes records through long-lived operator copies instead, nothing is derived at all; the width is the only number, and it holds for the life of the run. 4. **The older two-phase disk-to-disk model** derives a split count for the reading half and takes a stated count for the second half, which is why its two halves can run at very different widths. Any sentence beginning "the engine works out how many pieces to use" is therefore true of some of this class and false of the rest. The safe formulation is the one this question uses: derived in a finite run, declared in a continuous one, with the details of derivation varying. ## What this leaves to others This question owns the *origin* of the number. What value to pick, how to trade bytes-per-piece against available lanes, and at which moment a chosen value can be changed are a separate subject — the sibling on choosing how many pieces. Where a piece then runs, and why one may be heavier than its neighbours, are separate again.
- If the stored input doubles overnight, which of the two numbers moves by itself?Only the derived one, and only on runtimes that re-derive per run: twice the stored bytes gives roughly twice the pieces at the read. A declared operator width is unmoved — it is a statement, and the input has no way to change it. Where a split count is fixed at submission, even the finite job waits for the next submission.
- Does a continuous job implemented as repeated small finite jobs derive anything?Yes, on the read side. Each small finite run measures what arrived in its interval and divides that, so the reading width does track volume. The operators after the read still run at the stated width, so the job as a whole is a hybrid rather than a counter-example.
- What does the derivation report when the stored form cannot be divided?One piece, honestly. The derivation describes the stored form rather than the cluster, so an undividable object becomes a single unit of work regardless of its size or the capacity available. The consequence is a ceiling of one; the reasons a form resists division belong to the layout subject.
saying these in an interview costs you the question
- Says the author always states the number of pieces directly.
- Claims a continuous job derives its width from the incoming record rate.
- Assumes every runtime recomputes the division on every run.
- Thinks stored byte count is what sets a declared operator width.
- Believes the derivation counts logical rows rather than the stored form.