Some monorepo teams generate the CI pipeline from the affected package set instead of writing every job statically. What does that buy, and what does it cost?
answer
- forty-eight skipped jobs versus none
- the definition becomes program output
- reviewers see the generator, not the pipeline
- new packages picked up automatically
- persist what was generated
basics
~20 sGeneration keeps the pipeline proportional to the change: a small setup job computes the affected packages and emits a child pipeline containing only those jobs. The cost is that the definition no longer exists in the repository — it is program output, so it is harder to review, reason about and debug.
solid answer
~50 sThe static approach writes a job per package and guards each with a condition, so a fifty-package repository has fifty jobs of which forty-eight are skipped, and every new package needs a hand-edit. The generated approach runs one short setup job that computes the affected set and emits a pipeline definition containing exactly the jobs that matter; the platform then executes that as a child pipeline. What you buy is a pipeline whose size tracks the change rather than the repository, plus per-package variation without copy-paste, plus new packages being picked up automatically. What you pay is reviewability: the definition someone reviews is now a generator, and the actual pipeline for a given commit exists only as generated output. Debugging means reproducing the generator, generator bugs fail in ways YAML linting cannot catch, and reporting is spread across a parent and its children.
go deeper
Know the basic shape: instead of listing every job up front, a first job works out which packages are affected and produces the pipeline that then runs.
Compare the two approaches concretely — a run full of skipped jobs versus a run containing only real work — and name the tradeoff that the definition is no longer readable in the diff.
Focus on the new failure class: a generator that emits a valid pipeline missing a job produces a silent skip. Talk about testing the generator, persisting its output, and failing closed.
Own the hybrid decision: which parts of the pipeline stay static and reviewable, which are generated, and what the organization must maintain — generator tests, local reproducibility, reporting that understands parent and child runs.
## Two ways to make a monorepo pipeline proportional Once a monorepo has more packages than anyone wants to build on every commit, the pipeline has to shrink to fit the change. There are two structurally different ways to do that. **Statically, with conditions.** Write a job for each package and attach a condition to each — a path rule, or a check against a precomputed affected list. All jobs exist in the definition; most do not run. **Dynamically, by generation.** Run one small job first. It resolves the diff base, computes the affected set, and writes out a pipeline definition containing only the jobs those packages need. The platform then runs that generated definition as a child pipeline. Most major CI platforms support some form of this — a generated child pipeline, a dynamically produced configuration, an uploaded pipeline definition — under different names. ## What generation genuinely buys **The pipeline is proportional to the change, not to the repository.** With fifty packages, a one-package change produces a run with a handful of jobs, not fifty entries of which forty-eight are skipped. That matters for readability more than for cost: a run page that lists only real work is one a human can actually read. **Per-package variation without duplication.** Packages are not uniform — some need a database service, some need a longer timeout, some need a browser. Expressing that statically means fifty hand-maintained blocks that drift. A generator expresses it once as logic over package metadata. **New packages are free.** Add a package, and the generator emits its jobs on the next run. In the static model, someone must remember to add a block, and the failure when they forget is silent — the package simply has no CI. **Work that is only knowable at run time.** Sharding a test suite across N containers where N depends on how many affected packages there are, or splitting by recorded timings, cannot be written down ahead of time. ## What it costs **The definition leaves the repository.** In the static model the reviewer of a pull request sees the pipeline change in the diff. In the generated model they see a change to the program that produces pipelines. The actual definition for any given commit exists only as generator output, and the reasoning "what will this run do?" becomes "what will this program emit?". **Generator bugs are a new failure class.** A template that emits subtly invalid configuration, an escaping bug in an interpolated package name, an off-by-one in a shard count. Some of these fail loudly at pipeline-parse time; the dangerous ones emit a *valid* pipeline that quietly omits a job. That is the same silent-skip class as a wrong diff base, arriving by a different route. **Debugging is indirect.** "Why did this job not run?" requires reproducing the generator against that revision, which means the generator needs to be runnable locally and its output needs to be persisted with the run. Teams that skip persisting the generated definition make every such investigation archaeology. **Reporting fragments.** Status, timings and logs live across a parent run and one or more children. Anything that aggregates pipeline data — dashboards, flake tracking, duration trends — needs to understand the parent/child relationship, and tooling that assumes one run per commit needs adapting. **Setup latency.** Every run pays for the setup job before anything useful starts: a checkout deep enough to resolve the base, the affected computation, then emitting and starting the child. On a small change that overhead can rival the work itself. ## Choosing between them A reasonable progression: start static while the package count is small and the jobs are uniform. Move to generation when either the number of packages makes the static file unreadable, or per-package variation makes it repetitive, or you need run-time-determined fan-out. A hybrid is common and often best: keep a static skeleton for the things that always run — lint, the repository-wide checks, the security scan, the deploy stage — and generate only the per-package build-and-test fan-out. That keeps the reviewable parts reviewable and confines generation to the part that genuinely varies. Whatever you choose, treat the generator as production code: it is under test, its output is persisted with each run, and it fails closed. A generator that emits an empty pipeline on an internal error must be an error, never a fast green run.
- What is the most dangerous kind of bug in a pipeline generator?One that emits a valid pipeline missing a job. Invalid output fails at parse time and is obvious; a well-formed pipeline that quietly omits a package's tests goes green and merges. It is the same silent-skip failure class as a wrong diff base, so the generator needs unit tests over its output and must fail closed rather than emit an empty pipeline on an internal error.
- When would you keep the static approach instead?When the package count is small enough that the file stays readable, the jobs are near-identical, and nothing about the fan-out depends on run-time information. Static definitions are diffable, reviewable and debuggable with no extra machinery, and that is worth a lot. A common middle ground keeps a static skeleton for repository-wide checks and generates only the per-package fan-out.
- What overhead does the generation step itself add to every run?A setup job that must check out enough history to resolve the diff base, compute the affected set, emit the definition, and hand it to the platform to start a child pipeline. That is fixed cost paid on every change, including trivial ones, and on a small change it can be comparable to the actual work — which is the main argument against generating for small repositories.
saying these in an interview costs you the question
- Calls generated pipelines strictly better than static ones
- Ignores that reviewers no longer see the real definition
- Lets a generator error emit an empty pipeline
- Never persists the generated definition with the run
- Treats the generator as a script rather than production code