In a monorepo containing hundreds of independently deployable projects, why would a CI pipeline choose to only build and test the projects 'affected' by a given code change, rather than running the full build and test suite on every commit?
answer
- cost scales with change size not repo size
- reverse-dependency graph traversal
- false negative = graph misses an edge
- safety net full runs
- Nx / Bazel affected/rdeps
basics
~20 sBecause rebuilding and testing everything every time gets too slow as the codebase grows. Affected detection figures out which projects a change could actually break and only runs those, so a small change gets fast feedback instead of a multi-hour full run.
solid answer
~40 sA monorepo can hold hundreds or thousands of projects, but any single commit usually touches a tiny fraction of them. Running the entire build/test matrix on every push makes CI time grow with total repo size, not with change size — a one-line fix in a leaf library waits behind builds of unrelated services. Affected detection walks the dependency graph outward from the changed files to find every project that depends on what changed, directly or transitively, and scopes the CI run to just that set. This keeps feedback latency roughly proportional to change size, lets a monorepo scale to thousands of projects without CI collapsing, and cuts compute cost. The trade-off is correctness risk: if the dependency graph is incomplete, affected detection can silently skip something that actually broke.
go deeper
Should articulate the basic motivation — smaller, faster CI runs scoped to what changed — even without precise graph-traversal vocabulary.
Should describe the mechanism at a high level: changed files map to owning projects, then dependents are found via the dependency graph, and only that set runs.
Should discuss the false-negative risk from incomplete graphs and name at least one mitigation (safety-net full runs, explicit dependency declarations).
Should connect affected detection to overall monorepo scalability strategy, discuss base-commit selection for PR vs. merge-queue contexts, and reason about when the correctness/speed trade-off is acceptable for a given organization.
## The scaling problem a monorepo creates A monorepo's core selling point — everything in one place, atomic cross-project commits, shared tooling — comes with a scaling problem single-project repos never face: **the codebase keeps growing even though any individual change stays small**. If CI naively rebuilds and retests every project on every push, pipeline duration becomes a function of total repository size, not of the diff. - In a repo with **fifty projects** that might be tolerable. - In one with **two thousand** it turns a one-line fix in a shared utility library into an hours-long, machine-hungry run, and every engineer pays that tax on every commit regardless of what they touched. ## How affected detection works Affected detection (also called 'scoped builds' or 'impacted project detection') solves this by treating the codebase as a **dependency graph** — nodes are projects/packages/targets, edges are 'depends on' relationships derived from imports, build-file declarations, or module manifests. Given a set of changed files (usually from a git diff against a base commit), the tool then: 1. maps each file to the project that owns it; 2. does a **reverse-dependency (upward) traversal** of the graph to find every project that transitively depends on an owning project. That traversal result is the **affected set** — everything whose behavior could plausibly change because of this diff. Build, lint, and test tasks then run only for that set. ## The payoff The payoff is that CI latency and cost scale with change size instead of repo size. - A change confined to a leaf package with no internal consumers might affect only itself. - A change to a widely-shared core library might affect dozens of downstream projects, which is exactly the correct, conservative behavior — you want the CI cost to reflect real blast radius. This is what makes monorepos with tens of thousands of projects (Google's Piper/Blaze ecosystem, or open-source tools like Nx and Bazel used at that scale) operationally viable at all; without scoping, the CI queue time alone would make the monorepo model unusable past a few hundred projects. ## The trade-off The trade-off is that affected detection trades exhaustiveness for speed, and its correctness is entirely bounded by how complete and accurate the dependency graph is. Static analysis of import statements and build manifests captures most dependencies, but it can miss dynamic ones: - reflection-based dependency injection; - string-based dynamic imports; - config files that reference another project by name at runtime; - shared runtime infrastructure (a database schema, a message topic contract) that isn't expressed as a code-level import at all. When the graph misses an edge, affected detection produces a **false negative** — a project that would actually break ships through CI green because the tool never knew it depended on the changed code. This is the central risk teams accept when adopting affected-only CI, and it's why most organizations pair it with a periodic full-suite safety net (nightly or pre-release) rather than trusting affected detection as the sole gate forever. ## What the tools actually do A concrete illustration: Nx computes the affected set by diffing against a base ref (commonly the last common ancestor with main, or the previous successful run on trunk), mapping changed files to their owning `project.json` targets via its project graph, then running `nx affected -- test`/`build`/`lint` only for the reverse-dependency closure. Bazel does the analogous thing with `bazel query 'rdeps(//..., set(changed_targets))'` built from its own build-graph rules. Both systems make the same implicit bet: that the declared graph is a faithful superset of real runtime dependencies. Where that bet breaks down — most commonly with polyglot repos mixing statically-analyzable code with loosely-coupled services communicating over the network — affected detection needs supplementary signals to stay trustworthy: - explicit `implicitDependencies` declarations, - tags marking 'always run', - or manual dependency pinning. Getting this wrong doesn't fail loudly; it fails as a production incident weeks later that nobody's CI run ever caught.
- What base commit should the diff be computed against for a pull request versus for a trunk merge queue?For a PR, you typically diff against the merge-base with the target branch (not the tip of main, which would include unrelated commits that landed after the branch point). For trunk/merge-queue runs, you diff against the last commit that CI verified as green, so the affected set covers exactly what's new since the last known-good state. Getting this wrong either over-scopes (wasted CI time) or under-scopes (missed breakage from unrelated concurrent merges).
- How do teams mitigate the false-negative risk from an incomplete dependency graph?Common mitigations are explicit 'implicit dependency' declarations for anything not visible to static analysis (e.g. a service that reads another project's config schema), tagging certain projects as 'always affected' when they're high-risk, and running a full, untargeted build/test suite on a schedule (nightly) or before releases as a backstop. Some teams also monitor for graph drift by periodically auditing whether declared dependencies match actual runtime coupling.
- Does affected detection help with build time as well as test time?Yes — the same reverse-dependency closure that scopes tests also scopes which projects need to be rebuilt, so unaffected projects reuse their last-built artifacts. This is a separate concern from artifact-level incremental compilation and content-hash caching within a single project, which determine whether a rebuild is actually necessary once a project is in the affected set.
It's like a building inspector who, after an electrician touches one circuit, only re-inspects the rooms wired to that circuit instead of re-inspecting the entire building — fast and usually correct, but only as good as the wiring diagram; if a room was secretly wired off-diagram, it never gets checked.
saying these in an interview costs you the question
- says CI always builds everything and that's fine at any scale
- doesn't mention the dependency graph or reverse-dependency traversal at all
- assumes affected detection is 100% safe with no false-negative risk
- can't explain what 'base commit' the diff is computed against
- conflates affected-project scoping with compiler incremental-build caching