In a monorepo build tool such as Nx, Turborepo or Bazel, how is the "affected" set of targets computed, and why is that stronger than a path filter?
answer
- reverse edges, not forward
- changed files map to owning projects
- then everyone who depends on them
- derived from imports, not from directories
- Bazel spells it rdeps()
basics
~20 sThe tool builds a dependency graph of the repository's projects, maps the changed files to the projects that own them, then walks the graph backwards to include every project that transitively depends on those. Path filters see only the directly edited directories.
solid answer
~50 sAn affected computation has two halves. First, map the diff to owning projects: every changed file belongs to some package, so a set of changed files becomes a set of directly changed projects. Second, walk the *reverse* dependency edges — the projects that import the changed ones, then the projects that import those, transitively. The union is the affected set, and only its targets need to build and test. Nx does this with `nx affected --base=origin/main --head=HEAD`; Turborepo's `--filter` accepts a git range plus a dependents modifier; in Bazel it is literally a query, `rdeps(//..., <changed targets>)`. The strength over path filters is that the graph is derived from the actual build inputs — imports, target dependencies, workspace links — so a change to a shared library correctly pulls in its consumers. The weakness is that anything not expressed in the graph, such as a runtime-only or configuration coupling, is invisible to it.
code
bash · 8 lines# 1. Which projects did this branch change, plus everything that depends on them?
nx affected -t test --base=origin/main --head=HEAD
# 2. Same idea in Turborepo: a git range, plus the dependents modifier
turbo run test --filter='...[origin/main]'
# 3. In Bazel it is an explicit reverse-dependency query over the workspace
bazel query 'rdeps(//..., //packages/auth:lib)'go deeper
Know that these tools build a dependency graph of the repository's projects and use it to run only the work a change could have broken, rather than running everything.
Walk both stages out loud — changed files to owning projects, then the transitive closure of dependents — and name one command that does it. Be clear that this is the reverse of ordinary dependency resolution.
Discuss where the graph's edges come from and what that implies for trust: inferred graphs miss network, runtime and generated-code coupling, so you must know which classes of change your affected set cannot see.
Frame the choice between an inferred graph and an enforced one as an investment decision: what declaring every dependency costs, what it buys in correctness guarantees, and whether your organization's failure rate justifies paying it.
## What "affected" actually means A monorepo build tool models the repository as a directed graph. Nodes are projects (packages, apps, libraries) or, in Bazel's case, individual build targets. Edges are dependencies: `services/api` depends on `packages/auth` because it imports it, or because a build file declares it. The affected set for a change is the set of nodes whose output could differ because of that change. Computing it is a two-stage operation. **Stage one — changed files to changed projects.** Take the diff between the incoming revision and a base, and attribute each changed file to the project whose directory contains it. This is where the graph and the path filter still look alike. **Stage two — reverse-dependency closure.** This is the part a path filter cannot do. For each directly changed project, find everything that depends on it, then everything that depends on *those*, transitively. If `packages/auth` changed and both `services/api` and `services/admin` import it, and `e2e/smoke` imports `services/api`, all four are affected. Only the union gets built and tested. Note the direction. Ordinary dependency resolution walks *forward* (what do I need?). Affected computation walks *backwards* (who needs me?). That inversion is the whole idea, and it is why Bazel expresses it as a query over reverse dependencies. ```bash # Nx: run the test target for changed projects and their dependents nx affected -t test --base=origin/main --head=HEAD # Turborepo: packages touched since origin/main, plus everything depending on them turbo run test --filter='...[origin/main]' # Bazel: reverse dependencies of a target, within the whole workspace bazel query 'rdeps(//..., //packages/auth:lib)' ``` ## Where the graph comes from The accuracy of the whole mechanism rests on how the edges are derived, and the three tools sit at different points on that spectrum. - **Inferred from source.** Nx and Turborepo build the project graph largely from workspace metadata and static imports — the package manifests, the workspace links, the import statements the analyzer can see. Cheap to adopt: the graph mostly already exists in the code. - **Declared explicitly.** Bazel requires every dependency to be written down in the build files, and enforces it — a target cannot read a file it did not declare. That is far more work up front, and it buys a graph that is complete by construction rather than by best effort. That tradeoff — inferred and convenient versus declared and enforced — is the real design decision, and it decides how much you can trust the affected set as a merge gate. ## Task-level granularity A subtlety worth knowing: affected computation happens per *target*, not per project. `nx affected -t lint` and `nx affected -t e2e` can produce different work even for the same change, because the task graph knows that a project's lint task depends only on its own sources while its build task depends on its dependencies' build outputs. Bazel takes this furthest — the node is the target, so changing one file in a large package can affect a single test target rather than the whole directory. ## Global inputs A lockfile, a root compiler configuration, or a shared toolchain file is an input to every project. Good tools model this explicitly: the graph treats such files as inputs of every node, so touching one legitimately marks everything affected. That is the correct answer, not a bug — but it means an automated dependency-update change is a full build, and teams sometimes shave it by declaring narrower input sets. Every narrowing is a claim you are asserting about correctness. ## What the graph still misses The affected set is only as good as the edges: - **Runtime-only coupling.** Service A calls service B over HTTP. Nothing imports anything, so no edge exists, and B's contract test is not affected by A's change. - **Dynamic loading.** A plugin resolved by name at runtime, a module loaded from a configuration string — no static import for the analyzer to see. - **Data and generated code.** A schema file, a fixture, a code generator's output checked in — edges exist only if someone declared them. - **Environment.** A base image or toolchain version pinned outside the repository changes the build without changing any tracked file. ## Why this is the standard answer Path filters approximate the graph by hand, and hand-maintained approximations drift. An affected computation derives the same answer from the artifacts the build already needs to be correct, so it stays right as the code moves. That is the argument to make in an interview: not "it is faster" — both are faster than building everything — but "it is derived rather than asserted, so it stays correct as packages are added and dependencies change."
- Why does the affected computation walk reverse dependencies rather than forward ones?Forward edges answer "what does this project need to build", which is a scheduling question. The question here is "whose behaviour could this change break", and that is answered by the projects that depend on the changed one. So you invert the graph and take the transitive closure of consumers — the changed project plus everyone downstream of it.
- What kind of coupling does an inferred project graph typically miss?Anything without a static import edge: two services that talk over the network, a plugin loaded by name at runtime, generated code whose generator is not declared as an input, or a base image pinned outside the repository. None of these produce an edge, so a change on one side leaves the other side unaffected in the tool's view and untested in CI.
- How does Bazel's approach to the dependency graph differ from Nx's or Turborepo's?Bazel requires dependencies to be declared in build files and enforces them at execution time, so a target cannot silently read something it did not declare. Nx and Turborepo infer most edges from workspace metadata and static imports, which is far cheaper to adopt but means the graph is complete only where the analyzer can see. The tradeoff is up-front cost against how much you can trust the result.
saying these in an interview costs you the question
- Describes forward dependencies instead of reverse dependents
- Thinks affected means only the directly edited packages
- Assumes the inferred graph captures runtime coupling
- Treats lockfile changes as affecting nothing
- Conflates the affected set with a cache hit