skip to content

As the owner of a large monorepo's CI, how far would you trust the affected-target graph as the only pre-merge gate, and what backstop would you run?

level: principalimportance: should knowfreq 33%

answer

  1. the set is a causal claim
  2. inferred edges, inferred confidence
  3. name the classes you cannot see
  4. bound the escape window, don't deny it
  5. building everything is a valid answer

basics

~20 s

Trust it in proportion to how the edges are derived: enforced declarations earn near-total trust, inferred imports earn conditional trust. Either way run a periodic full build so an escape is bounded by that interval rather than discovered by a customer.

solid answer

~50 s

The affected set is a claim about causality, and its credibility is exactly the credibility of the graph's edges. Where dependencies are declared and enforced — a build system that forbids reading undeclared inputs — the claim is close to sound, and gating on it is defensible. Where edges are inferred from static imports, whole classes of coupling are invisible: network calls between services, plugins resolved by name, generated code, environment pinned outside the repository. So I would gate on affected for the common path, and buy back the residual risk three ways. Run a full build on a schedule, which bounds how long an escape can hide. Record what each commit actually tested, so post-incident you can answer the question rather than guess. And treat every unmodelled coupling found in an incident as a graph defect: declare the edge, or accept the class in writing. The judgement is choosing an escape rate you can live with, not pretending it is zero.

go deeper

for a junior

Understand that the affected set is a best guess about what a change could break, and that CI can be green because a test was skipped rather than because it passed.

for a middle

Be able to name specific couplings an inferred graph cannot see — network calls, dynamically loaded plugins, generated code, out-of-repository inputs — and say what each one implies for gating.

for a senior

Show the operational design: fail closed on computation errors, run a scheduled full build to bound the escape window, and persist what each commit actually tested so incidents are lookups rather than investigations.

for a principal

Own it as an explicit exchange rate between feedback time and escape probability, with numbers on both sides. Be willing to say building everything is correct below a certain size, and to tier trust by blast radius rather than applying one uniform policy.

## The claim you are making Gating merges on an affected-only pipeline is an assertion: *nothing outside this set could have been broken by this change*. That is a causal claim about the software, and CI is only entitled to make it as strongly as the dependency graph is complete. The principal-level job is to size that entitlement honestly and to design for the part you cannot close. ## Grade the graph before you grade the risk Edges come from somewhere, and where decides how much they are worth. **Declared and enforced.** A build system that requires every dependency to be written down and prevents a target from reading anything it did not declare produces a graph that is complete by construction. The cost is real — every dependency is someone's explicit work, and the migration is measured in quarters — but the affected set is then a genuine safety property rather than a heuristic. **Inferred from source.** Workspace metadata plus static import analysis. Cheap, usually right, and blind by construction to anything that is not a static import. **Asserted by hand.** Path globs a human wrote. These are correct on the day they are written and drift from then on. Most organizations are at the second tier and reason as if they were at the first. Naming which tier you are on is most of the answer. ## The classes an inferred graph cannot see Write them down, because "we might miss something" is not a plan and this list is what makes it actionable: - **Network coupling.** Two services with a contract and no shared import. - **Dynamic resolution.** Plugins, handlers, or modules loaded from a configuration string. - **Generated and committed artifacts.** A schema and its generated client, where the generator is not modelled as an input. - **Out-of-repository inputs.** A base image tag, a toolchain version, an external dependency resolved at build time. - **Behavioural coupling through data.** Two packages that agree on a serialization format nobody declared. Each class has a specific remedy — declare the edge, add the file as a global input, add a contract test that always runs — and the point of enumerating them is to convert an abstract worry into a defect list someone can close. ## The backstops, in the order I would add them **1. Fail closed on computation failure.** If the diff base cannot be resolved or the graph cannot be built, build everything. An empty affected set must be an error condition, never a fast green run. This one is cheap and removes the most common silent escape. **2. A periodic full build of the main branch.** The most valuable control per unit of effort, because it converts an unbounded escape window into a known one. Nightly is the common cadence. The value is not that it finds much — it is that when it does find something, you learn within hours instead of from a customer, and you get a bisectable window. **3. Recorded evidence.** Persist the change set, the affected set and the task list for every run. After an incident the question "was this package tested on that commit?" should be a lookup, not an investigation. Without this, every escape becomes an argument about what probably happened. **4. Always-on checks that ignore the graph.** Some things are cheap enough to run unconditionally regardless of affectedness: repository-wide lint, the dependency and secret scans, contract tests at service boundaries, the smoke test of the assembled system. These specifically cover the classes the graph is blind to. **5. Graph defects as incident output.** Every escape ends with either a new declared edge or a written acceptance that this class is not covered. Without that discipline the same class recurs and the team quietly stops trusting CI at all, which is a far worse outcome than a known, bounded gap. ## Sizing it as a decision, not a preference The decision is an exchange rate: pull-request feedback minutes against escape probability. Both sides are measurable if you choose to measure them. Full-build wall-clock time, affected-build wall-clock time, changes per day, and the count of escapes attributable to a missing edge over the last two quarters. If escapes are near zero and full builds cost an hour, the affected gate is obviously right. If escapes happen monthly and a full build costs eight minutes, gating on affected is optimizing the wrong thing and you should simply build everything. That last point is worth saying explicitly in an interview, because it is the one candidates skip: **building everything is a legitimate answer.** The affected machinery — graph maintenance, base resolution, generator code, cache infrastructure — is a system with its own failure modes and its own maintenance cost. Below some repository size it is not worth owning. The senior move is knowing where that line sits for your repository rather than adopting the machinery because large repositories elsewhere need it. ## What blast-radius thinking adds Finally, not every package deserves the same confidence. A change to a shared authentication library and a change to an internal developer tool have very different consequences if the affected set is wrong. It is reasonable to apply a tiered policy: for a small number of high-blast-radius packages, always run their consumers' tests regardless of what the graph says. That is a cheap, targeted way to spend confidence where it matters, instead of applying one uniform level of paranoia to fifty packages.

  • Why is a nightly full build the highest-value backstop rather than a stronger pre-merge gate?
    It bounds the escape window. A missing edge means something merges untested; without a full build you find out when a customer does, at an unknown later date. A scheduled full run makes the maximum exposure the interval between runs and gives you a small, bisectable range of commits. It also costs nothing on the pull-request path, so it does not trade away the feedback time the affected gate exists to buy.
  • When would you conclude the affected machinery is not worth owning at all?
    When a full build is cheap relative to change volume. The machinery is a system — graph maintenance, base resolution, generator code, cache infrastructure — with its own failures and upkeep. If everything builds in eight minutes and you merge thirty changes a day, building everything is simpler, has no silent-skip failure mode, and needs no one to maintain it. Adopt affected builds when the full build has become the bottleneck, not before.
  • How would you treat an escape caused by a coupling the graph never modelled?
    As a graph defect with a closing action, not as bad luck. Either declare the missing edge, add the file as a global input, or add an always-on check covering that class — and if none of those is worth doing, write down that the class is accepted and unmonitored. The failure mode to avoid is repeated escapes of the same class, which erodes trust in CI faster than any individual outage.
  • Would you apply the same level of trust to every package in the repository?
    No. Blast radius differs: a shared authentication library that everything depends on is not the same risk as an internal developer tool. For a small set of high-consequence packages it is reasonable to always run consumers' tests regardless of what the graph says. That concentrates the extra cost where a wrong answer actually hurts, rather than applying uniform paranoia across fifty packages.

saying these in an interview costs you the question

  • Treats the affected set as proof rather than a claim
  • Cannot name a coupling class the graph misses
  • Has no plan for bounding the escape window
  • Assumes every large repository needs affected builds
  • Calls building everything an unsophisticated answer

context