What are the most common ways affected-detection-based CI goes wrong in production, and how do these failures typically show up — as a build error, or as something worse?
answer
- false negative = graph missing an edge -> silent skip
- green CI, prod breaks later
- wrong base commit (moving tip vs merge-base)
- coarse boundaries = false positives, fine boundaries = more missed-edge surface
- safety-net full runs as backstop
basics
~20 sThe biggest danger is the tool missing a real connection between two parts of the code, so it thinks a project is safe when it isn't. CI stays green, but the broken thing ships anyway and only gets noticed later, usually in production.
solid answer
~50 sThe dominant failure mode is a false negative: the dependency graph misses a real coupling (dynamic imports, reflection, shared runtime contracts like a database schema or message format, config-file references) so a project that actually broke never enters the affected set and CI reports green while shipping a regression — this typically surfaces later as a production incident, not a CI failure. A second failure mode is a wrong or stale base commit for the diff, causing the affected set to be computed against the wrong starting point and either miss recent changes or drag in unrelated ones. A third is overly coarse project boundaries (one giant 'project' covering unrelated code), which defeats scoping by making everything affected together, or conversely overly fine boundaries with undeclared cross-project coupling, which increases the false-negative surface. Mitigations include explicit implicit-dependency declarations, periodic full-suite safety-net runs, and monitoring for graph/reality drift.
go deeper
Should understand that the tool can sometimes miss a real connection between projects, meaning something might break without CI catching it.
Should be able to name at least one concrete cause of a missed dependency (dynamic import, reflection, runtime config) and describe it as producing a false negative.
Should distinguish false-negative failures from base-commit/staleness failures and from coarse-vs-fine project-boundary trade-offs, and name at least one mitigation like implicit-dependency declarations or safety-net full runs.
Should describe an operational strategy combining fast scoped checks with periodic exhaustive checks as the actual reliability model, and reason about how to identify and specially handle high-fan-in, high-risk projects.
## Why these failures are dangerous Affected-detection failures are dangerous precisely because most of them don't look like failures at all — CI reports green, the PR merges, and the breakage surfaces somewhere else, later, and more expensively. Understanding the specific mechanisms behind this is what separates someone who's read about affected detection from someone who's operated it. ## The false negative — a missing edge The most consequential failure mode is the **false negative**: the dependency graph the tool relies on doesn't contain an edge that exists in reality. Static-analysis-based graphs (built by parsing import/require statements) are blind to dynamic and runtime couplings: - reflection-based dependency injection that resolves a class by string name; - a plugin loaded via a path read from a config file at runtime; - a service that reads another project's output file from a shared filesystem location; - two services coupled only through a database schema or message-queue contract with no code-level import between them at all. When one of these hidden edges exists and the upstream side changes, the downstream consumer never enters the affected set, its tests never run, and CI reports green on a change that will break that consumer the moment it deploys. This is the textbook 'CI green, prod broken' incident, and postmortems for it almost always trace back to exactly this gap between the declared/inferred graph and actual runtime coupling. ## Base-commit miscomputation A second class of failure is base-commit miscomputation. If the diff is computed against the wrong reference — the current tip of a fast-moving target branch instead of the merge-base, or a stale cached 'last known good' commit that's actually several merges behind — the resulting changed-file set is wrong, and everything downstream (project mapping, traversal) inherits that error. In a merge-queue setup specifically, if two PRs land close together and each PR's affected computation only considers its own diff without accounting for the other PR's concurrent changes, an interaction bug between the two changes can slip through because neither PR's affected set, computed in isolation, covers the combination. ## Where the project boundaries are drawn A third failure mode sits in how project boundaries are drawn. - **Drawn too coarsely.** If a monorepo defines projects too coarsely — say, one giant 'backend' project covering a dozen logically separate services — affected detection can't discriminate between them, so any change anywhere in that blob marks the whole thing affected, defeating the purpose of scoping (this shows up as 'affected detection barely saves any CI time' rather than as an incident, since it's a false positive problem, not a false negative one). - **Drawn too finely.** Conversely, splitting boundaries too finely without carefully declaring every cross-boundary dependency increases the surface area for missed edges — more, smaller projects means more inter-project relationships that all need to be captured accurately, and any one gap reproduces the false-negative failure above. ## Graph staleness A fourth, more operational failure mode is graph staleness: if the dependency graph itself is cached and the cache-invalidation logic has a bug, the graph used for a given CI run can reflect an older state of the codebase than the diff being evaluated, causing recently added or removed dependency edges to be ignored during traversal. ## Mitigations Mitigations cluster around acknowledging that the graph is a model, not a proof. 1. Teams add explicit **'implicit dependency' declarations** for known-but-invisible couplings (Nx's `implicitDependencies` in `nx.json`, or equivalent constructs elsewhere) whenever they know a hidden edge exists, rather than hoping static analysis finds it. 2. They tag especially risky or foundational projects as **'always run,'** bypassing affected scoping entirely for the highest-blast-radius code. 3. And critically, they don't treat affected-only CI as the sole safety net forever — a nightly or pre-release **full build/test run** across every project acts as a backstop that will eventually catch a false negative that slipped through PR-time CI, even though it does so with much higher latency than the fast-path affected run. Organizations operating monorepos at scale (this pattern is well documented from companies like Google) treat the combination of fast, scoped, imperfect PR-time checks plus slower, exhaustive, periodic checks as the actual reliability story — neither one alone is sufficient.
- Why is the false-negative failure mode specifically hard to detect through normal CI monitoring?Because by definition CI reports success — there's no failing test, no red build, nothing to alert on. The gap only becomes visible when the broken behavior manifests downstream, often in a different system, at a different time, disconnected from the commit that caused it, which makes root-causing it back to a missed dependency edge much harder than diagnosing a normal test failure.
- How would you decide which projects to tag as 'always affected' rather than relying on graph-based detection for them?Good candidates are projects with unusually high fan-in (many other projects depend on them) combined with any known dynamic/runtime coupling that static analysis can't see — shared schemas, widely-used config, security-sensitive core libraries. The cost of running them unconditionally is small relative to total pipeline time, while the cost of a missed dependency on them is disproportionately high given how many consumers they have.
- Besides a nightly full-suite run, what's another way to catch false negatives before they reach production?Some teams run staging/integration environments that exercise real runtime coupling (actual service-to-service calls, actual schema reads) rather than relying purely on unit-test-level affected scoping, since integration environments naturally exercise dependencies that a static import graph can't see. Contract testing between services is another complementary approach that doesn't depend on the build-time dependency graph at all.
It's like a smoke detector wired to only some rooms of a house: as long as every fire starts in a wired room, everyone stays safe and confident the system works — but a fire in the one unwired closet burns just as real, and the alarm's silence gives false confidence right up until it doesn't.
saying these in an interview costs you the question
- thinks affected detection is fully safe with zero risk of missing a real dependency
- can't explain why a false negative looks like a passing CI run rather than a failure
- doesn't distinguish false negatives (missed real dependency) from false positives (over-broad scoping)
- has no mitigation beyond 'trust the tool'