Your team proposes parallelising every batch pipeline by default, so how do you decide whether that pays?
answer
- measure the window before constructing
- the serial share caps the gain
- a fifth of the run gives a fifth at best
- price the permanent costs, not the change
- end-to-end wall clock on real volume
basics
~20 sDecide from measurement, not policy. Find where the wall clock actually goes, bound the achievable gain by the share of work that stays serial, then weigh that ceiling against permanent costs: a bound to tune, order to restore, and failures that no longer reproduce.
solid answer
~50 sStart by measuring where the nightly window is spent. **Amdahl's law** sets the ceiling: if the stage you would parallelise holds a fifth of the wall clock, making it four times faster leaves `0.8 + 0.05` of the original run, about fifteen percent better — a real but modest number that has to justify the construction. Then check whether the pipeline is the constraint at all; batches are often bounded by a single shared dependency or by the write-back, in which case fanning out the scoring stage only queues more work at the same place. Against the ceiling, price the permanent costs: choosing and re-tuning a bound, restoring order at the sink or removing the sink's dependence on it, resuming a run that failed with holes, and a dependency that now sees several concurrent callers. My default is to parallelise where measurement names the stage as the constraint, and to make it a reviewed decision rather than a template.
go deeper
The takeaway is that making one part faster does not make the whole run faster by the same factor. Measure where the time goes before changing anything.
Be able to compute the ceiling: a stage holding a fifth of the wall clock can return at most that fifth, so four times faster on it yields roughly fifteen percent overall. That arithmetic decides whether the work is worth starting.
Argue from end-to-end evidence on realistic volume, including the load the dependency absorbed, and name the permanent costs the construction adds — the bound, the ordering decision, and resuming a run that failed with holes.
Answer the policy, not just the case: publish the conditions under which a pipeline is parallelised, where the bound is configured and reviewed, and what the standard answer is when a run fails halfway, so the second pipeline costs less than the first.
## The question behind the question *Parallelise everything by default* is a policy proposal, and the useful response is not yes or no but the decision procedure that would settle each case. That procedure has three parts: where the time goes, what the ceiling is, and what the construction costs forever. ## Step one — where the wall clock actually goes Before any construction, attribute the run: - How much of the window is the stage in question, measured end to end and not from a stage timer alone? - How much is the read, the write-back, the commit, the retry of a flaky dependency? - What is idle — waiting on something shared with other jobs running at the same hour? - Is the window actually violated, or is this an optimisation with no consumer? The last point closes more of these proposals than any other. A batch that finishes two hours inside its window does not need a construction that changes how failures resume. ## Step two — the ceiling **Amdahl's law** states the limit: the achievable speedup is capped by the fraction of the work that stays serial. If a fraction s of the run cannot be parallelised, the whole run can never be faster than 1/s, no matter how many workers are available. Worked out on the scoring batch: | Scoring stage's share of wall clock | Stage made 4× faster | Stage made free | |---|---|---| | 20% | about 15% shorter | at most 20% shorter | | 50% | about 27% shorter | at most 50% shorter | | 80% | about 38% shorter | at most 80% shorter | The first row is the one people get wrong: four times faster on a fifth of the run leaves `0.8 + 0.05 = 0.85` of the original, not a quarter of it. Compute the ceiling before deciding, because a ceiling of fifteen percent rarely justifies taking on an ordering contract and a tuning parameter. ## Step three — what the construction costs forever Parallel execution inside a pipeline is not a one-off change. It leaves behind: 1. **A bound** that must be chosen from the scarce resource, re-checked when the hardware or the dependency changes, and agreed with whoever owns that dependency. 2. **An ordering decision** — either a reordering step with a window and a straggler policy, or a sink rewritten so position no longer matters. 3. **Harder failure handling**: a failed parallel run leaves a set of results with holes rather than a prefix, so resumption has to be driven by what was written rather than how far the run got. 4. **Non-deterministic reproduction**: interleaving differs between runs, so a defect that appeared last night may not appear when you rerun it. 5. **A new load profile** on everything behind the stage, visible to other teams. ## Step four — the evidence that settles it Evidence that counts: - **End-to-end wall clock**, before and after, on production-shaped volume — not a stage timer, and not a thousand-record fixture. - The **load the dependency saw**, recorded during the same run, because that is the price being paid elsewhere. - A **throughput-against-bound curve** showing where the gain stops, which also identifies whether the resource you assumed was scarce really is. - A **failure drill**: kill the run midway and resume it, to prove the ordering and resumption decisions actually work. Evidence that does not count: a faster stage inside an unchanged window, a profile showing work spread across workers, or a small load test finishing quickly. ## The answer to the policy I would not accept *by default*. I would publish the procedure instead: parallelise when the measured share of wall clock makes the ceiling worth having, the per-element work is genuinely independent, and the sink either tolerates completion order or has an agreed way to restore it. Where those do not hold, the cheaper moves usually are to remove serial work from the window, to batch the write-back, or to run independent partitions of the input as separate runs — which buys the same overlap without putting an ordering contract inside every pipeline. A lead's job here is to make the second pipeline cheaper than the first: one reviewed shape, one place the bound is configured, one documented answer to what happens when a run fails halfway.
- A stage became four times faster but the nightly window did not move. What happened?The stage was not where the wall clock went. Its share of the run bounded the possible gain, and the rest — reading, writing back, waiting on a shared dependency — was untouched. Either the attribution was never done, or it was done with a stage timer rather than end to end, which shows a stage improving while the run does not.
- What ongoing cost does a parallel pipeline carry that a serial one does not?A bound to choose and re-tune, an ordering decision at the sink, resumption logic for a run that failed with holes rather than a prefix, failures that no longer reproduce identically, and a dependency that now sees several concurrent callers from this job. All of them are permanent, and all of them are paid by whoever maintains it next.
- What would you do instead when the ceiling is too low to justify the construction?Attack the serial remainder: remove work from the window, batch the write-back, drop a redundant pass over the data, or split the input into independent partitions run separately. Partitioning gets overlap without putting an ordering contract and a tuning parameter inside the pipeline itself.
saying these in an interview costs you the question
- Parallelises before measuring where the time goes
- Ignores the serial share that caps achievable speedup
- Counts cores but never the shared dependency behind the stage
- Reports a faster stage while end-to-end time is unchanged
- Treats bound tuning and ordering as free follow-up work
- Validates the change on a small fixture only