Concretely, how does a build tool determine the 'affected' set of projects for a given code change — walk through the steps from a git diff to a final list of projects to build and test?
answer
- diff -> file-to-project map -> reverse traversal
- seed nodes = directly changed projects
- rdeps query / reverse-dependency closure
- merge-base vs moving tip
- Bazel BUILD deps vs Nx inferred imports + implicitDependencies
basics
~20 sThe tool lists which files changed, figures out which project each file belongs to, then looks at a map of 'who depends on whom' to find every other project that uses those changed projects, directly or through a chain. All of those become the affected set.
solid answer
~40 sIt's a three-stage pipeline. First, compute the changed-file set via a git diff between a base commit and the current one. Second, map each changed file to the project that owns it, using project boundaries defined by manifests (package.json, BUILD files, project.json). Third, take those directly-changed projects as seed nodes and run a reverse-dependency traversal — essentially a graph search from each seed node following incoming edges — over a pre-built project dependency graph, collecting every node reachable that way. The union of seed nodes and everything reachable via reverse edges is the affected set. The dependency graph itself is usually built once by statically parsing import statements or explicit dependency declarations, then cached and incrementally updated so recomputing it on each CI run is cheap.
go deeper
Should be able to describe the basic idea — changed files map to projects, and other projects that use those projects also get included — without needing precise graph-theory terms.
Should walk through the three stages (diff, file-to-project mapping, reverse-dependency traversal) and correctly identify that the traversal follows incoming edges from the changed projects.
Should discuss how project boundaries and dependency edges are declared or inferred in a specific tool (Bazel BUILD deps, Nx project graph/implicitDependencies), and how each stage fails independently.
Should compare declared-dependency (Bazel-style) versus inferred-dependency (import-analysis-style) graph construction trade-offs, and reason about graph staleness/caching strategy at scale.
## A pipeline with three stages Affected detection is a small pipeline with three distinct stages, and understanding each stage separately is what lets you debug it when it produces a wrong answer. ## Detecting what changed Stage one is change detection. The tool computes a git diff between a base ref and the current commit — typically `git diff --name-only <base>...<head>` — producing a flat list of changed file paths. Choosing the base matters: - For a **feature branch** it should be the merge-base (the common ancestor with the target branch), not the target's current tip, otherwise unrelated commits that landed on the target branch after you branched get folded into your 'changes.' - For **trunk-based merge queues**, the base is usually the last commit CI verified green, so the window covers exactly what's new. ## Mapping files onto projects Stage two is ownership mapping: converting a list of file paths into a list of projects. Every project-aware build tool needs an explicit notion of project boundaries — a `BUILD`/`BUILD.bazel` file marking a Bazel package, a `project.json`/`package.json` marking an Nx/Lerna/Turborepo project, or similar. Each changed file is matched to the nearest enclosing project boundary (typically by directory prefix), producing the **'directly changed' project set**. Files outside any project (root config, CI workflow files) are often special-cased to mark everything affected, since they can influence any build. ## The reverse-dependency traversal Stage three is the graph traversal that gives affected detection its name. The build tool maintains (or computes on demand) a **project dependency graph**: nodes are projects, directed edges represent 'depends on,' derived from static analysis of import/require statements, explicit dependency declarations in manifests, or a combination. From the directly-changed projects as seed nodes, the tool performs a reverse-dependency traversal — following edges backwards, i.e., finding every node that has a path of dependency edges leading into a seed node. This is the same operation as Bazel's `rdeps()` query function or Nx's internal project-graph traversal: it answers 'who consumes what I changed, transitively?' The union of the seed set and everything found by that traversal is the final affected set. Only those projects get their build/test/lint targets scheduled. ## Why the separation matters This three-stage separation matters because each stage fails independently and in a different observable way. - **A wrong base commit (stage one)** produces a diff that's too broad or too narrow — visible as CI scope obviously not matching the PR's actual content. - **Wrong project-boundary mapping (stage two)** usually shows up as a change in a shared file (like a root tsconfig) either being ignored, or being over-attributed to a single project when it should mark everything affected. - **An incomplete dependency graph (stage three)** is the most insidious: it silently drops nodes that should be in the affected set because the edge into them was never captured — a project that reads another project's output via a filesystem path rather than an import statement, for instance, has no edge in a graph built purely from static import analysis. ## What the tools actually do A concrete real-world instantiation: - **Nx** builds and caches a project graph from parsing TypeScript/JavaScript imports plus explicit `implicitDependencies` entries in `nx.json` for cases static analysis can't see (e.g., a backend project that depends on a shared database schema project with no code import between them). Running `nx affected --target=test` performs exactly the three stages above — git diff, file-to-project mapping via the cached graph, then `nx affected` internally computing the reverse-dependency closure — and schedules `test` only for the resulting projects. - **Bazel**'s model is similar but graph construction is driven by explicit `deps` attributes in `BUILD` files rather than inferred imports, which trades some ergonomics for a graph that's exhaustive by construction (every dependency must be declared or the build fails to find the symbol), closing off the 'invisible dynamic dependency' failure mode that plagues import-inference-based tools — at the cost of requiring every dependency to be hand-declared.
- Why does the traversal go in the reverse direction — from changed project to its dependents — rather than forward, to its dependencies?Because the question being answered is 'what could this change break,' not 'what does this project rely on.' A project's own dependencies don't need retesting just because the project changed; it's the consumers of the changed code whose behavior might now be wrong, and those are found by following edges backwards from the change.
- What happens when a file changes that doesn't belong to any declared project, like a root-level CI config or shared linter config?Most tools special-case these as 'global' files: any change to them marks the entire project graph as affected, since they can influence how every project builds or runs. This is deliberately conservative — the alternative of silently under-scoping a root config change is a much worse failure mode than occasionally running a full build.
- How is the dependency graph itself kept up to date without recomputing it from scratch on every CI run?Tools typically cache the parsed graph and incrementally update only the parts touched by the current diff, rather than re-parsing every project's imports on every invocation. The specifics of that caching (content hashes, invalidation keys) shade into incremental-build caching territory, which is a distinct concern from the graph-traversal logic itself.
It's like tracing a power outage: first you find which breaker tripped (changed files), then which room that breaker feeds (owning project), then you follow the wiring diagram backwards to find every other room wired downstream of that breaker (reverse-dependency traversal) — those are the rooms you go check.
saying these in an interview costs you the question
- can't distinguish 'directly changed' projects from 'affected via dependents'
- thinks the traversal follows forward/outgoing dependency edges instead of reverse/incoming
- doesn't mention that project boundaries have to be explicitly declared somehow
- assumes the dependency graph is always perfectly accurate with no manual declarations needed