skip to content

How do you keep an affected-detection CI pipeline fast and reliable as a monorepo grows into the thousands of projects — what specifically starts to break, and what strategies address it?

level: principalimportance: should knowfreq 40%

answer

  1. graph construction cost -> cache/incrementally update the graph
  2. high-fan-in core projects = de facto critical path, invest in their speed
  3. historical-duration bin-packing across worker fleet, not round-robin
  4. shard/stagger even the safety-net full run at scale
  5. merge queue: test against queue order, not isolated per-PR base

basics

~20 s

As the codebase gets huge, even the 'figure out what's affected' step and the safety-net full runs get slow, and a few heavily-used shared projects become bottlenecks everyone waits on. You fix this by splitting work across many machines, caching aggressively, keeping the dependency map itself fast to compute, and giving special treatment to the most-depended-on projects.

solid answer

~50 s

At thousands-of-projects scale, several things that were negligible earlier become real bottlenecks: computing/loading the dependency graph itself takes measurable time, a small number of highly-depended-on 'core' projects create critical paths that funnel most affected sets through the same tasks, and periodic full-suite safety-net runs become expensive enough that running them nightly is itself a capacity-planning problem. The standard responses are: caching and incrementally updating the dependency graph rather than rebuilding it from scratch every run; distributing the resulting task DAG across a fleet of CI workers or a remote-execution cluster, load-balanced using historical task-duration data rather than naive round-robin; treating high-fan-in projects as special — extra investment in their own build/test speed pays off across every affected run that touches them, and some organizations shard even their safety-net full runs across days/machines rather than running everything at once. Merge queues at this scale also need to reason about affected-set overlap between concurrently-landing changes, not just per-PR isolation.

go deeper

for a junior

Not typically expected to reason about this; a basic notion that 'huge repos need extra tricks to stay fast' is sufficient.

for a middle

Should recognize that distributing tasks across multiple machines is one lever for scaling, even without deep detail on load-balancing strategy.

for a senior

Should identify high-fan-in projects as critical-path bottlenecks and know that graph computation itself has a cost that needs caching at scale.

for a principal

Should reason across all the strain points together — graph-construction cost, critical-path investment in core projects, duration-aware distributed scheduling, safety-net-run affordability, and merge-queue-order-aware affected computation — and connect them to concrete organizational trade-offs.

## What starts to strain Affected detection is designed to make CI cost scale with change size rather than repo size, but at genuinely large scale — thousands of projects, hundreds of engineers committing continuously — several assumptions that held comfortably at hundreds of projects start to strain, and the mitigations become a distinct engineering discipline in their own right. ## The cost of affected computation itself The first strain point is the cost of affected computation itself. Building or loading the full project dependency graph, then running the reverse-dependency traversal, is not free — at thousands of nodes with dense edges, naive re-parsing of every project's manifest/imports on every CI invocation becomes a meaningful fraction of total pipeline latency. The standard response is to make the graph itself an **incrementally-maintained artifact**: cache the parsed graph, and on each run only re-parse the subset of projects whose manifests changed, patching the cached graph rather than rebuilding it from scratch. Nx's daemon process and cached project graph, and Bazel's persistent analysis cache, both exist substantially to amortize this cost across runs rather than paying full graph-construction cost every time. ## High-fan-in nodes become the bottleneck The second strain point is structural: at scale, dependency graphs in real organizations are rarely uniform — they typically have a small number of extremely **high-fan-in nodes** that a large fraction of all other projects transitively depend on: - shared UI component libraries, - core data models, - authentication SDKs. Because affected detection's reverse-traversal naturally funnels almost every meaningfully-sized change through these nodes, they become de facto critical-path bottlenecks on a huge fraction of CI runs, not just occasionally. The response isn't a build-scoping technique per se but an organizational one: treat build/test speed of these specific high-fan-in projects as a first-class investment, because every second shaved off their task duration multiplies across nearly every affected run in the whole repo, whereas the same investment in a leaf project only helps that project's own runs. ## Distributing the task execution The third strain point is distributing the actual task execution. A single CI machine's core count stops being enough well before a monorepo reaches thousands of projects with meaningful PR velocity, so task execution has to fan out across a worker fleet — via a remote-execution protocol (Bazel's remote execution API) or a distributed-task-orchestration layer (Nx Cloud-style agents, or custom sharding on top of a CI provider's matrix-job primitive). At this point, naive work distribution (round-robin, or splitting tasks evenly by count) performs poorly, because task durations vary wildly — a few large integration-test targets can dominate wall-clock time even if they're a small fraction of the task count. Production-grade schedulers instead **bin-pack tasks across workers using historical duration data** collected from prior runs, aiming to equalize total predicted time per worker rather than task count per worker, which is what actually shortens the long pole. ## Paying for the safety net The fourth strain point is the periodic full-suite safety net itself. Since affected detection accepts some false-negative risk in exchange for speed, most large-scale operations keep an exhaustive nightly or pre-release run as a backstop — but at thousands of projects, even that exhaustive run becomes expensive enough to need its own scaling strategy: - sharding it across many machines just like a normal CI run; - or staggering different portions of the full suite across different nights rather than running literally everything every night, accepting a longer detection window for the rarest false-negative classes in exchange for keeping the safety-net's own cost bounded. ## Merge queues at high velocity Finally, at high commit velocity, merge queues introduce a subtler correctness dimension: affected detection computed per-PR in isolation can miss interaction effects between two changes landing close together, each individually 'affected-clean' but combined causing a break neither change's isolated affected set would have caught. Mature merge-queue implementations address this by re-running affected checks against the queue's actual merge order (each change tested against the tip that includes everything ahead of it in the queue, not against its original branch point), rather than trusting each PR's isolated affected computation as sufficient once several changes are landing concurrently. This is less about the build-scoping mechanism itself and more about correctly sequencing when and against what base each affected computation runs — but at scale, getting that sequencing wrong is a common source of 'green in the queue, broken after merge' incidents distinct from the dependency-graph-completeness failures discussed elsewhere.

  • Why doesn't simply adding more CI workers fully solve the scaling problem at thousands of projects?
    Because the critical path through high-fan-in core projects imposes a floor on wall-clock time that additional workers can't shrink — more workers increase how much can run in parallel at any given moment, but can't make a strictly sequential chain of dependent tasks finish faster. Beyond some point, the bottleneck shifts from 'not enough parallel capacity' to 'the longest dependency chain itself is too slow,' which only gets fixed by speeding up or restructuring that chain.
  • What's a concrete sign that a monorepo's dependency graph construction has itself become a bottleneck, separate from task execution time?
    If CI pipeline time doesn't shrink proportionally even for very small, isolated changes — a one-line fix in a leaf project still takes several minutes just to determine and report the affected set before any actual build/test task starts — that's a signal the graph computation step itself, not the scoped task execution, is now the dominant cost, and needs its own caching/incremental-update investment.
  • Why might a team deliberately shard their nightly full-suite safety-net run across multiple nights rather than running the entire suite every night?
    Because at large enough scale, even the exhaustive nightly run competes for the same finite compute budget as regular CI, and running literally everything every night can become as expensive as regular development traffic. Staggering different portions across nights trades a longer worst-case detection window for a false negative against keeping the safety net's steady-state cost bounded and predictable.

It's like scaling a city's road network: a few major arterial roads (high-fan-in projects) carry most of the traffic regardless of how many side streets you add, so the real capacity investment goes into widening those arteries and smartly timing traffic lights (duration-aware scheduling) rather than just adding more side streets (more workers) and hoping congestion resolves itself.

saying these in an interview costs you the question

  • assumes affected detection scales for free just by adding more CI machines with no other changes
  • doesn't recognize that a small number of high-fan-in projects dominate the critical path at scale
  • thinks the dependency graph itself is free to compute regardless of repo size
  • has no answer for how the safety-net full run itself is kept affordable at scale
  • doesn't distinguish per-PR isolated affected checks from merge-queue-order-aware checks

context