Your platform's default is one submission per pipeline - which boundaries should become scheduler edges instead, and what does each one cost?
answer
- the cheap edge is the default
- promote only for what crosses it
- who else reads the intermediate?
- every promotion buys an owned artefact
basics
~20 sPromote a boundary only when something other than the next step needs it: a different program or runtime, an intermediate real consumers read, an external wait, or two halves wanting very different capacity. Each promotion costs a materialised named artefact, a second start-up, and a second picture to read.
solid answer
~50 sKeep the default, because the engine's edge is the cheap one: inside one submission the engine owns the handoff and no one has to name, own or delete an intermediate. Promote a boundary to the workflow graph - where an edge means only 'the upstream step finished' - for reasons that are about *what crosses it*, not about how the pipeline is drawn. Four hold up: the two sides are genuinely different programs or runtimes; the intermediate is a dataset other consumers already read; the downstream waits on something that is not a record, such as an external arrival or an approval; or the two halves want such different capacity that holding one allocation across both is wasteful. Each promotion buys you an artefact with an owner, a schema and a lifetime, plus a second start-up and a second place the pipeline's shape lives.
go deeper
Recall which edge is cheap: inside one submission the engine owns the handoff and names nothing, while a scheduler edge forces a stored, named intermediate. That asymmetry is the whole basis of the default.
Explain the cost of a promoted boundary concretely - write, read, second start-up, an artefact with an owner and a lifetime, and the loss of rewriting across it on engines that rewrite.
Show you decide on what crosses the boundary rather than on tooling taste, and that you can price a promotion against the capacity saving that is usually offered as its justification.
Own the standard and its enforcement: a default, four admissible arguments recorded where the pipeline is defined, and an owner plus a deletion rule for every intermediate a promoted boundary creates.
## Why the default is the right default Inside one submission - one whole program handed to the cluster and run end to end - an edge carries records and the engine owns the handoff. Nobody names the intermediate, nobody owns its schema, nobody pays for it after the run, and on engines that rewrite a declarative program into a physical plan the two sides can be reorganised together. In the workflow graph, an edge means only that the downstream step starts after the upstream one reported success, and **nothing travels along it** - so every promoted boundary converts a free, invisible handoff into a durable artefact somebody has to design. That asymmetry is why 'one submission per pipeline' is a sound default and why each departure from it should be argued. ## The four arguments that actually hold 1. **Two genuinely different programs.** Different languages, different runtimes, one of them a declarative statement handed to a service that owns its own storage, plan and capacity. No single submission can contain both, so the boundary exists whether you like it or not; the only question is what crosses it. 2. **The intermediate has consumers.** If other pipelines, other teams or an analyst already read that dataset, you were materialising it regardless. The write is not a new cost, so the boundary is nearly free - it costs one start-up. 3. **The downstream waits on something that is not a record.** A file another organisation delivers, an approval, a signal from a system outside the pipeline. Waiting is precisely what an ordering edge is for, and jamming it inside a submission means a long-running program doing nothing but holding capacity. 4. **The two halves want different capacity.** A short, wide stage and a long, narrow one inside one submission usually hold one allocation shaped for the wider of the two. Splitting lets each ask for what it needs. This is a resources argument and it is real, but it is the weakest of the four, because the saving has to beat the materialisation. ## What each promoted boundary costs - **A named artefact.** A location, a format, a schema somebody owns, a rule for what a rerun does to it, and a deletion policy. This is the cost that outlives the decision. - **A full write and a full read**, unless argument 2 applied and you were writing it anyway. - **A second start-up**, whose size depends on how the machines are supplied. - **No rewriting across the boundary** on engines that would have done so. Concede the variation honestly: an engine that executes per-record functions almost literally, and the oldest disk-to-disk designs, were not going to rewrite much across it in the first place. - **A second picture.** The pipeline's shape now lives in two systems, and a reader has to hold both to know what runs. Two pictures is a real and permanent tax on comprehension. ## Arguments that should not decide it | the argument | why it fails | |---|---| | 'the workflow graph is where the team looks' | legibility of a diagram, paid for with a materialisation per arrow; fix the visibility, not the architecture | | 'the scheduler gives us failure handling and history' | what a scheduler does about a failed step is a real subject, but it is a property of that system rather than a property of *this* boundary - it argues for a scheduler, not for cutting here | | 'splitting makes it faster' | each split adds a write, a read and a start-up; the only speed argument is the capacity one, and it must beat that bill | | 'the intermediate is big' | size argues for fewer boundaries, not more | ## The two failure modes of a platform standard **Over-promotion** is the common one: the workflow graph becomes the place where all structure is expressed, so stages of one computation become steps, data-sized fan-out migrates there, and the pipeline degenerates into a launcher loop with storage between every arrow. The symptom is a graph whose box count tracks the data. **Under-promotion** is rarer and worth guarding against so the standard is honest: one submission containing several teams' unrelated work, where a dataset many people need is invisible inside it, nothing can be run on its own, and the blast radius of any change is the whole thing. A workable standard says three things: the default is one submission per pipeline; a boundary is promoted only on one of the four arguments above, recorded in one line where the pipeline is defined; and every promoted boundary's artefact has a named owner, a schema and a deletion rule from the day it is created. What makes it stick is the last clause - a boundary with no owner for its intermediate is a boundary nobody decided on.
- How do you keep teams from promoting boundaries just to make the pipeline visible?Give them the visibility another way - the engine's own run view, a summary written per run, a lineage record - and make the promoted-boundary standard require a named owner for the intermediate. Most gratuitous boundaries do not survive the question 'who owns this dataset and when is it deleted'.
- Does the four-argument test change if some pipelines target a managed query service?The arguments hold, and argument one fires more often: a declarative statement handed to a service that owns its own storage, plan and capacity cannot live inside a submission to a cluster you sized, so the boundary is forced. What you control is how much crosses it and whether the intermediate is a real product.
- What is the cheapest promoted boundary you can have?One where the intermediate was going to be written anyway because other consumers read it. Then the only new cost is the second start-up, and you gain a dataset with an owner. That is why argument two is the strongest of the four and the one to look for first.
saying these in an interview costs you the question
- Promotes a boundary because the workflow graph is where the team looks.
- Counts only the write and forgets the artefact's owner, schema and lifetime.
- Claims more scheduler steps make a pipeline faster.
- Argues that a large intermediate is itself a reason to split the pipeline.
- Keeps everything in one submission even when other teams need the intermediate.
- Treats where to cut as a tooling preference rather than a question about what crosses.