A CI pipeline for a 200-package monorepo needs to decide which packages to test on a pull request that touches 3 files. Walk through how a monorepo orchestrator computes that 'affected' set, and what breaks if the dependency graph it relies on is incomplete.
answer
- diff base ref -> changed files -> owning packages
- walk dependents direction, not dependencies
- manifest-declared graph vs inferred-from-imports graph
- undeclared edge = silent CI blind spot
- root/shared package = affected-set explosion
basics
~20 sIt looks at which files changed, maps them to their owning packages, then walks the dependency graph outward to find every package that depends on those, directly or indirectly. If the graph is missing an edge, an affected package can be skipped and ship broken without CI catching it.
solid answer
~50 sThe tool first diffs the PR branch against a base ref (e.g., `main`) to get a raw file list, maps each file to the package that owns it, then treats those as 'directly changed' graph nodes. It walks the graph in the direction of dependents (not dependencies) — from each changed node, find every node with a path leading back to it — because a change to package B can only affect packages that depend on B, never packages B depends on. The union of directly-changed and transitively-dependent packages is the affected set that gets tasks scheduled. If the graph is missing an edge (say, package A silently imports from package B without that being declared), A never shows up as depending on B, so a breaking change to B fails to trigger A's tests — CI reports green while A is actually broken in production. This is why tools increasingly infer the graph from static analysis of actual imports rather than trusting hand-maintained manifests alone.
go deeper
Should understand at a basic level that only packages related to the change get tested, not the whole repo.
Should describe the diff -> owning-package -> graph-walk pipeline and know the walk direction is toward dependents.
Should identify undeclared-dependency graph gaps as a real production risk (false-green CI) and describe at least one concrete mitigation (import linting, Bazel sandboxed builds).
Should reason about graph topology as an architectural concern — recognizing when a centralized shared package is silently defeating affected-detection's value across the org, and driving decisions to decompose it or restructure ownership boundaries.
## What affected detection is for Affected-package detection is the piece of monorepo tooling that decides, for a given code change, exactly which subset of the workspace's build/test tasks actually needs to run — and getting it right is what makes the whole "don't rebuild everything" promise trustworthy rather than merely fast. ## The mechanism, in three steps The mechanism runs in three steps. 1. **First**, the orchestrator computes a raw file diff between the current branch and a base reference — typically `git diff --name-only origin/main...HEAD` or equivalent — producing a flat list of changed file paths with no package awareness yet. 2. **Second**, it maps each changed file to the package that owns it, usually by matching the file's path against each package's declared root directory; a change to `packages/auth/src/login.ts` gets attributed to the `auth` package. This produces the set of "directly changed" nodes in the project graph. 3. **Third** — and this is the step that actually requires the dependency graph — the orchestrator walks the graph in the dependents direction: starting from each directly-changed node, it follows edges backward (toward whatever depends on it, not what it depends on) to find every package that could be affected by that change, transitively. A change to a low-level `utils` package that three services depend on marks all three services affected, even though none of their own files changed, because their build/test output could be influenced by `utils`'s new code. ## Why the direction matters This directionality matters and is a common source of confusion: - **dependencies** — what a package needs; - **dependents** — what needs a package. They are the two ends of the same graph edge, and affected-detection specifically needs the dependents direction. Getting this backward — walking toward dependencies instead — would mark packages as affected because of what they rely on rather than what relies on them, which produces a nonsensical result (marking `utils` affected because `auth` changed, rather than the reverse). ## Why it exists The reason this exists, beyond the caching/speed motivation covered elsewhere, is **correctness under scale**: without it, a team either tests everything on every PR (slow, eventually untenable) or manually decides what to test (error-prone, and exactly the kind of judgment call that erodes under deadline pressure). Automating the graph walk makes the scope of CI verification a deterministic function of the change, not a matter of an engineer's memory of who depends on what. ## The critical trade-off: graph fidelity The critical trade-off, and the place this breaks in production, is **graph fidelity**. The affected calculation is only as trustworthy as the graph it walks, and graphs go stale or incomplete in a few characteristic ways. In ecosystems where dependencies are hand-declared in a manifest (a `package.json`'s `dependencies` list, a Bazel `BUILD` file's `deps` attribute), an engineer can add a new import in code without updating the manifest — the code compiles fine locally because the package is already installed/available in `node_modules` or on the classpath transitively, but the declared graph edge doesn't exist. The affected calculation then never propagates a change through that undeclared edge: package B changes, package A silently depends on B but isn't marked affected, A's tests don't run, and a breaking change to B ships inside A without CI ever exercising it. This is arguably the single most dangerous failure mode in this whole tooling category, because it produces false confidence — a fully green CI run — rather than an obvious loud failure. ## How the tools respond Different tools respond to this risk differently. | Tool | How the graph is established | What happens on an undeclared import | |---|---|---| | **Nx**, for JS/TS monorepos | mitigates it by statically parsing each file's actual `import`/`require` statements to infer the dependency graph rather than trusting only the manifest | catching the common case of "code imports something the manifest doesn't declare" | | **Bazel** | takes the opposite, stricter approach: requires every dependency to be explicitly and exactly declared in `BUILD` files | a sandboxed build will actually fail to compile if code references something not listed as a `dep` — turning a silent graph gap into a loud build error instead | The trade-off is upfront friction (every new import requires a manifest edit) versus long-run graph trustworthiness. ## The opposite extreme: graph explosion The other characteristic failure is graph explosion at the opposite extreme: a package sitting near the root of the dependency graph — a shared `types` or `utils` package many others import — means almost any change to it marks a huge fraction of the workspace as affected, which is correct but defeats the practical benefit of affected detection. Seeing this repeatedly in CI timing data is usually a signal to split that root package into narrower, more independently-versioned pieces rather than to distrust the graph itself. ## Where it shows up A concrete real-world instance: large Nx workspaces at companies with hundreds of packages routinely tune `nx affected` against `main` as their PR-gating CI step specifically because full-workspace test runs became multi-hour and untenable, and graph-inference correctness (via TypeScript's own module resolution) is what keeps that shortcut trustworthy enough to gate merges on.
- Why does the graph walk go in the 'dependents' direction instead of the 'dependencies' direction?A change to package B can only affect the behavior of packages that consume B's output, not the packages B itself consumes — B doesn't retroactively change what it depends on. So affected-detection has to ask 'who relies on this changed thing', which is the dependents (reverse) direction of the graph, not 'what does this changed thing rely on'.
- How would you catch an undeclared dependency edge before it causes a false-green CI run in a manifest-based system like npm/Bazel-style deps?Add a lint/build rule that fails when code imports a module not listed in that package's declared dependencies — Bazel does this by default via sandboxed builds; JS tooling can use ESLint's import-boundary rules or Nx's own dependency-checking lint rule. The key is turning the silent gap into a loud, CI-blocking error at the moment the undeclared import is introduced, rather than relying on someone noticing later.
- If `nx affected` (or equivalent) reports zero affected packages for a PR that clearly changes application behavior, what would you check first?First check that the base ref being diffed against is actually correct (comparing against a stale or wrong branch produces an empty or wrong diff). Then check whether the changed files fall inside a directory excluded from the project graph (e.g., root-level config not attributed to any package), which would mean the change is invisible to the affected calculation even though it clearly matters.
Like tracing a water-main break by following pipes downstream to every building that draws from it, not upstream to the reservoir — and if a building has an illegal, unmapped pipe tapped into the main, the utility company won't know to warn them when the water gets shut off.
saying these in an interview costs you the question
- Doesn't know the walk needs to go toward dependents, not dependencies
- Assumes the dependency graph is always accurate and never mentions manifest/import mismatches as a failure mode
- Can't explain what 'directly changed' vs 'transitively affected' packages means
- Thinks affected detection works by string-matching file paths without any graph traversal
- No awareness that a shared/root package can make affected-detection nearly useless if the graph is too centralized