skip to content

Build Scoping & Affected Detection

In a large repo you cannot build everything on every commit, so CI uses the project graph to find what a change actually affects. You will cover affected detection, target and test scoping, and scheduling tasks in parallel to keep pipelines from growing with the repo.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

5

In a monorepo containing hundreds of independently deployable projects, why would a CI pipeline choose to only build and test the projects 'affected' by a given code change, rather than running the full build and test suite on every commit?

level: juniorimportance: must knowfreq 75%

answer

  1. cost scales with change size not repo size
  2. reverse-dependency graph traversal
  3. false negative = graph misses an edge
  4. safety net full runs
  5. Nx / Bazel affected/rdeps

basics

~20 s

Because rebuilding and testing everything every time gets too slow as the codebase grows. Affected detection figures out which projects a change could actually break and only runs those, so a small change gets fast feedback instead of a multi-hour full run.

solid answer

~40 s

A monorepo can hold hundreds or thousands of projects, but any single commit usually touches a tiny fraction of them. Running the entire build/test matrix on every push makes CI time grow with total repo size, not with change size — a one-line fix in a leaf library waits behind builds of unrelated services. Affected detection walks the dependency graph outward from the changed files to find every project that depends on what changed, directly or transitively, and scopes the CI run to just that set. This keeps feedback latency roughly proportional to change size, lets a monorepo scale to thousands of projects without CI collapsing, and cuts compute cost. The trade-off is correctness risk: if the dependency graph is incomplete, affected detection can silently skip something that actually broke.

go deeper

for a junior

Should articulate the basic motivation — smaller, faster CI runs scoped to what changed — even without precise graph-traversal vocabulary.

for a middle

Should describe the mechanism at a high level: changed files map to owning projects, then dependents are found via the dependency graph, and only that set runs.

for a senior

Should discuss the false-negative risk from incomplete graphs and name at least one mitigation (safety-net full runs, explicit dependency declarations).

for a principal

Should connect affected detection to overall monorepo scalability strategy, discuss base-commit selection for PR vs. merge-queue contexts, and reason about when the correctness/speed trade-off is acceptable for a given organization.

## The scaling problem a monorepo creates A monorepo's core selling point — everything in one place, atomic cross-project commits, shared tooling — comes with a scaling problem single-project repos never face: **the codebase keeps growing even though any individual change stays small**. If CI naively rebuilds and retests every project on every push, pipeline duration becomes a function of total repository size, not of the diff. - In a repo with **fifty projects** that might be tolerable. - In one with **two thousand** it turns a one-line fix in a shared utility library into an hours-long, machine-hungry run, and every engineer pays that tax on every commit regardless of what they touched. ## How affected detection works Affected detection (also called 'scoped builds' or 'impacted project detection') solves this by treating the codebase as a **dependency graph** — nodes are projects/packages/targets, edges are 'depends on' relationships derived from imports, build-file declarations, or module manifests. Given a set of changed files (usually from a git diff against a base commit), the tool then: 1. maps each file to the project that owns it; 2. does a **reverse-dependency (upward) traversal** of the graph to find every project that transitively depends on an owning project. That traversal result is the **affected set** — everything whose behavior could plausibly change because of this diff. Build, lint, and test tasks then run only for that set. ## The payoff The payoff is that CI latency and cost scale with change size instead of repo size. - A change confined to a leaf package with no internal consumers might affect only itself. - A change to a widely-shared core library might affect dozens of downstream projects, which is exactly the correct, conservative behavior — you want the CI cost to reflect real blast radius. This is what makes monorepos with tens of thousands of projects (Google's Piper/Blaze ecosystem, or open-source tools like Nx and Bazel used at that scale) operationally viable at all; without scoping, the CI queue time alone would make the monorepo model unusable past a few hundred projects. ## The trade-off The trade-off is that affected detection trades exhaustiveness for speed, and its correctness is entirely bounded by how complete and accurate the dependency graph is. Static analysis of import statements and build manifests captures most dependencies, but it can miss dynamic ones: - reflection-based dependency injection; - string-based dynamic imports; - config files that reference another project by name at runtime; - shared runtime infrastructure (a database schema, a message topic contract) that isn't expressed as a code-level import at all. When the graph misses an edge, affected detection produces a **false negative** — a project that would actually break ships through CI green because the tool never knew it depended on the changed code. This is the central risk teams accept when adopting affected-only CI, and it's why most organizations pair it with a periodic full-suite safety net (nightly or pre-release) rather than trusting affected detection as the sole gate forever. ## What the tools actually do A concrete illustration: Nx computes the affected set by diffing against a base ref (commonly the last common ancestor with main, or the previous successful run on trunk), mapping changed files to their owning `project.json` targets via its project graph, then running `nx affected -- test`/`build`/`lint` only for the reverse-dependency closure. Bazel does the analogous thing with `bazel query 'rdeps(//..., set(changed_targets))'` built from its own build-graph rules. Both systems make the same implicit bet: that the declared graph is a faithful superset of real runtime dependencies. Where that bet breaks down — most commonly with polyglot repos mixing statically-analyzable code with loosely-coupled services communicating over the network — affected detection needs supplementary signals to stay trustworthy: - explicit `implicitDependencies` declarations, - tags marking 'always run', - or manual dependency pinning. Getting this wrong doesn't fail loudly; it fails as a production incident weeks later that nobody's CI run ever caught.

  • What base commit should the diff be computed against for a pull request versus for a trunk merge queue?
    For a PR, you typically diff against the merge-base with the target branch (not the tip of main, which would include unrelated commits that landed after the branch point). For trunk/merge-queue runs, you diff against the last commit that CI verified as green, so the affected set covers exactly what's new since the last known-good state. Getting this wrong either over-scopes (wasted CI time) or under-scopes (missed breakage from unrelated concurrent merges).
  • How do teams mitigate the false-negative risk from an incomplete dependency graph?
    Common mitigations are explicit 'implicit dependency' declarations for anything not visible to static analysis (e.g. a service that reads another project's config schema), tagging certain projects as 'always affected' when they're high-risk, and running a full, untargeted build/test suite on a schedule (nightly) or before releases as a backstop. Some teams also monitor for graph drift by periodically auditing whether declared dependencies match actual runtime coupling.
  • Does affected detection help with build time as well as test time?
    Yes — the same reverse-dependency closure that scopes tests also scopes which projects need to be rebuilt, so unaffected projects reuse their last-built artifacts. This is a separate concern from artifact-level incremental compilation and content-hash caching within a single project, which determine whether a rebuild is actually necessary once a project is in the affected set.

It's like a building inspector who, after an electrician touches one circuit, only re-inspects the rooms wired to that circuit instead of re-inspecting the entire building — fast and usually correct, but only as good as the wiring diagram; if a room was secretly wired off-diagram, it never gets checked.

saying these in an interview costs you the question

  • says CI always builds everything and that's fine at any scale
  • doesn't mention the dependency graph or reverse-dependency traversal at all
  • assumes affected detection is 100% safe with no false-negative risk
  • can't explain what 'base commit' the diff is computed against
  • conflates affected-project scoping with compiler incremental-build caching

context

open as a page

Concretely, how does a build tool determine the 'affected' set of projects for a given code change — walk through the steps from a git diff to a final list of projects to build and test?

level: middleimportance: must knowfreq 70%

basics

~20 s

The tool lists which files changed, figures out which project each file belongs to, then looks at a map of 'who depends on whom' to find every other project that uses those changed projects, directly or through a chain. All of those become the affected set.

open as a page

Once the affected set of projects is known, how do build tools schedule the resulting build/test/lint tasks — what determines which tasks run in parallel versus which must wait, and how does this scale across multiple CI machines?

level: seniorimportance: must knowfreq 65%

basics

~20 s

Tasks form a chain based on which project needs which other project built first. Tasks with no unfinished dependency can run at the same time; tasks that depend on another task's output have to wait for it. Big pipelines split this work across several machines to go faster.

open as a page

What are the most common ways affected-detection-based CI goes wrong in production, and how do these failures typically show up — as a build error, or as something worse?

level: seniorimportance: should knowfreq 55%

basics

~20 s

The biggest danger is the tool missing a real connection between two parts of the code, so it thinks a project is safe when it isn't. CI stays green, but the broken thing ships anyway and only gets noticed later, usually in production.

open as a page

How do you keep an affected-detection CI pipeline fast and reliable as a monorepo grows into the thousands of projects — what specifically starts to break, and what strategies address it?

level: principalimportance: should knowfreq 40%

basics

~20 s

As the codebase gets huge, even the 'figure out what's affected' step and the safety-net full runs get slow, and a few heavily-used shared projects become bottlenecks everyone waits on. You fix this by splitting work across many machines, caching aggressively, keeping the dependency map itself fast to compute, and giving special treatment to the most-depended-on projects.

open as a page