What does a team lose when its suite's run steps live only in the pipeline definition?
answer
- One invocation, two callers
- The stage should be one line
- Differences become arguments, not a second script
- Print the exact command in the log
basics
~10 sReproducibility. Steps that exist only in the pipeline definition cannot be run locally, so every change to them costs a push-and-wait cycle, and what a developer runs slowly diverges from what the gate runs.
solid answer
~50 sThree things. **Local reproduction**: a failure that only happens in the stage cannot be debugged where the developer is, so diagnosis becomes a series of pushes. **Reviewability**: pipeline definitions are edited by fewer people, reviewed more lightly, and cannot be exercised before merge, so mistakes in them ship. **Convergence**: once the stage knows things the local invocation does not - extra selections, variables, ordering, a cleanup step - “it passes on my machine” becomes structurally true rather than a joke. The alternative is a thin pipeline: one committed entry point in the repository holds the sequence, its defaults and its selections, and the stage's job is to call it with a couple of arguments. The pipeline still owns where the run happens, when it is triggered, and what it may reach. Everything about *how the suite runs* moves into the repository, where it is versioned, reviewed and runnable.
code
pseudocode · 15 lines# committed in the repository; a person and a stage call the same thing
entry_point run_suite(selection, target, output_dir):
defaults:
selection = "smoke"
target = "local"
output_dir = "./run-output"
log("invocation: run_suite selection=" + selection + " target=" + target)
prepare(target)
results = execute(selection)
write_results_document(output_dir, results)
return exit_status_for(results)
# the pipeline stage, in full:
# run_suite selection=full target=staging output_dir=<artefact-dir>go deeper
Be ready to say how you would run your team's suite on your own machine, and whether that is the same command the pipeline runs. If you cannot name the command, that is itself the answer to this question.
Explain the split: the pipeline owns when and where a run happens, the committed entry point owns how the suite runs. Expect to describe the cost of a stage holding selections, ordering or defaults nobody can invoke locally.
Show how you would migrate an accreted stage into an entry point without a freeze, and how you keep one implementation rather than two that drift. Be ready to say which differences legitimately remain as arguments.
Own the trade-off between a thin pipeline and an entry point that grows into an unreadable command with dozens of flags. Be ready to say which axes deserve to vary and how you judge that a stage is doing too much.
There are two ways to describe how a suite runs. One puts the sequence in the pipeline's own definition: a stage lists a dozen steps, each with its arguments, conditions and variables. The other puts the sequence in a single **entry point committed to the repository**, and reduces the stage to one line that calls it. Both produce the same run. Only one of them can be run by a person. ## What the thin shape looks like A thin pipeline draws a line between two kinds of concern: | Concern | Owner | | --- | --- | | When a run is triggered | The pipeline | | Which machine it lands on | The pipeline | | What the run is permitted to reach | The pipeline | | Where artefacts are published afterwards | The pipeline | | Which cases the run selects | The entry point | | Setup, teardown and their ordering | The entry point | | Default deadlines and limits | The entry point | | How results and artefacts are produced | The entry point | The rule of thumb: **the pipeline owns where and when; the entry point owns how.** If a stage contains a conditional about how the suite behaves, that conditional is on the wrong side of the line. ## What you lose when the steps live only in the stage - **Local reproduction.** A failure that only appears in the stage has to be debugged from outside it. The loop becomes edit, push, wait, read the log - minutes per iteration instead of seconds, and every iteration leaves a commit behind. - **Reviewability.** Pipeline definitions attract lighter review than code, are edited by fewer people, and cannot be exercised before they are merged. A mistake in one is discovered by shipping it. - **Convergence.** Once the stage knows things the local invocation does not - an extra selection, a variable, a cleanup step, an ordering - *"it passes on my machine"* stops being a joke and becomes structurally true. The two runs are different runs, and nobody can enumerate the differences. - **Portability.** A sequence written in one pipeline system's vocabulary must be rewritten to run anywhere else: a different provider, a scheduled job, a bisect script, a laptop. A committed entry point moves by being called. - **Attribution.** When a run's duration or cost has to be explained, a list of stage steps is a black box with no internal accounting. An entry point can time and report its own phases. ## Making one invocation serve both callers The goal is one implementation with **arguments**, not two implementations kept in sync. In practice: 1. **Take everything that varies as an argument** - the selection, the target, the output location, the concurrency - and give each a default that works on a developer machine with nothing configured. 2. **Read the surroundings only through named, documented inputs**, so a person can see what they must supply and the list is finite. 3. **Print the exact invocation at the top of the run's log.** Reproducing a stage failure should be a copy, not an archaeology exercise. 4. **Keep the stage to one call plus arguments.** If a second step appears, ask whether it is a where-and-when concern; if it is not, it belongs inside the entry point. 5. **Make the local path the one people use daily.** This is what keeps it honest: a wrapper that exists only for form is never exercised and rots exactly like unused code. ## The limits, honestly Thin does not mean everything is reproducible locally. Some runs need a deployed target only the pipeline can reach, hardware only the fleet has, or a fan-out across many machines a single laptop cannot imitate - and how that fan-out is arranged is the pipeline's concern, not this one's. The claim is narrower and still worth a great deal: **the definition of the run is one thing**, and the differences between contexts are expressed as arguments to it rather than as a second implementation nobody ever diffs. There is a failure mode in the other direction too. An entry point that absorbs everything becomes a single command with thirty flags, several of which interact, and reasoning about it becomes its own job. Keep the arguments to the axes that genuinely vary between callers; anything that has only ever been passed one value is a default, not an argument. When the flag list grows past what one person can hold in their head, that usually means the entry point is running several different things and should be several entry points with distinct promises. Finally, migrating an accreted stage does not need a freeze. Move one concern at a time: pull the sequence into the entry point, have the stage call it with the arguments that reproduce today's behaviour exactly, confirm the two runs match, then delete the old steps. The end state is a stage a newcomer can read in five seconds and a command they can run before lunch.
- Which parts of running a suite should stay in the pipeline rather than move into the entry point?The parts only the pipeline can own: when the run is triggered, which machine it lands on, what it is permitted to reach, and where its artefacts are published afterwards. Everything about how the suite itself runs - selection, ordering, defaults, deadlines, how results are written - belongs in the committed entry point, because that is the part a person needs to reproduce.
- How do you keep the local path from rotting when only the pipeline exercises it?Make it the path people already use daily. If running the suite locally means invoking the same entry point with different arguments, it is exercised constantly and breaks loudly. A local wrapper that exists only for form is never run and decays exactly like an untested code path; the fix is one implementation with arguments rather than two kept in sync by discipline.
- A failure reproduces only in the pipeline, never locally. Is the thin entry point still worth it?Yes, and it is what makes the difference diagnosable. With one invocation the only variables are the arguments and the surroundings, so you can enumerate them: a different target, a different data set, parallel neighbours, a slower machine. With two implementations you first have to prove the two runs were even the same run, which usually costs more than the bug.
It is the difference between a recipe written in the cookbook and one that exists only in a single kitchen's habits: the second is still a recipe, but it cannot be checked, taught or corrected anywhere else.
saying these in an interview costs you the question
- Keeps a second local script and assumes it still matches
- Puts conditionals and case selections in the stage definition
- Debugs a stage failure by pushing commits repeatedly
- Treats the pipeline definition as unreviewable operations configuration
- Assumes anything unreproducible locally is not worth structuring