As a monorepo grows to hundreds of contributors and thousands of packages, what specific tooling investments are needed to keep builds, tests, and version control usable -- and what breaks if those investments aren't made?
answer
- affected-target computation
- remote/content-addressable build cache
- sparse checkout / virtual filesystem (EdenFS, VFS for Git)
- Bazel/Buck/Nx/Turborepo
- CI cost scales with change size, not repo size
basics
~20 sBig monorepos need smart tools that only build/test the parts that actually changed, plus tricks so git doesn't have to download or scan the whole giant history. Without those tools, everything gets painfully slow as the repo grows.
solid answer
~50 sAt scale, a monorepo needs three tooling layers: (1) a dependency-aware build system (Bazel, Buck, Nx, Turborepo, Pants) that computes exactly which targets are affected by a change and only rebuilds/retests those, ideally with remote caching so identical inputs never rebuild twice; (2) CI orchestration that maps changed files to affected targets and runs a bounded subset rather than the whole suite; (3) VCS-level scaling -- sparse/partial checkouts, virtual filesystems (Meta's EdenFS, Microsoft's VFS for Git), or a purpose-built VCS (Google's Piper) so clients don't need the full multi-terabyte history on disk. Skip these and you get the classic symptoms: CI times that grow with total repo size instead of change size, git status/clone becoming multi-minute operations, and eventually teams routing around the pain by forking out into separate repos anyway -- defeating the point of the monorepo.
go deeper
Should recognize, in plain terms, that a big repo needs some kind of smart build/CI system so it doesn't rebuild or retest everything on every change; doesn't need to name specific tools.
Should name at least one real build system (Bazel/Nx/Turborepo/Buck) and explain the affected-target idea: only rebuild/retest what a change actually touches.
Should discuss the full stack -- build graph, remote caching, CI sharding by affected targets, and VCS-level scaling (sparse checkout/virtual filesystem) -- and be able to describe the concrete symptoms of skipping each layer.
Should be able to reason about when this investment is and isn't worth making for a given org size/stage, and speak to the migration cost (hermetic build rewrites, infra to run cache servers) as a real trade-off, not just a benefit.
## Why naive tooling breaks at scale The core mechanical problem a large monorepo creates is that most VCS and build tools were designed with an implicit assumption: repository size and 'amount of work for a given operation' scale together, and that's fine because repositories are small. Once a repository holds thousands of packages and history spans years, that assumption breaks -- a one-line change to a leaf utility should trigger a tiny, fast build/test cycle, but naive tools instead do work proportional to the whole repository, not the change. Naive tools here are: - a build script with no dependency tracking - a CI job that just runs the full test command at the root - a `git clone` that fetches full history Fixing this requires deliberately building or adopting tooling at three layers. ## The first layer: a dependency-aware build system The first and most important layer is a **dependency-aware build system**: - `Bazel` (Google's open-sourced Blaze) - `Buck2` (Meta) - `Pants` - JS-ecosystem tools like `Nx` and `Turborepo` These systems require every package/target to declare its dependencies explicitly, which lets the tool construct a full dependency graph of the repository. When a change comes in, the build system computes the exact transitive set of targets whose inputs changed -- the **'affected target'** set -- and only rebuilds/retests those, skipping everything untouched. Layered on top is **remote/content-addressable caching**: if a target's inputs (source files, dependency versions, compiler flags) hash to something the cache has already built, the result is fetched instead of rebuilt, even across different engineers' machines or CI runs. This is what lets Google claim most builds are cache hits rather than full recompiles despite a codebase with billions of lines. ## The second layer: CI orchestration The second layer is **CI orchestration** built on that same affected-target computation: instead of running the entire test suite on every pull request, CI computes which test targets are reachable from the changed files and runs only those, often sharded across many parallel workers. This keeps PR feedback time roughly proportional to the size of the change, not the size of the repository. ## The third layer: VCS and filesystem scaling The third layer is **VCS/filesystem scaling**, because even with smart builds, a full checkout and full history of a multi-million-file repository is itself too large to be practical on a laptop. Solutions here include: - **sparse or partial checkout** -- only materializing the subdirectories a given team touches - **virtual filesystems** that lazily fetch file contents on first access -- Meta's `EdenFS` and Microsoft's VFS for Git (built specifically so the Windows source tree, hundreds of gigabytes, could live in a single Git-compatible repo) are the best-known examples Google's Piper doesn't use Git at all; it's a custom, centralized, massively sharded backend accessed by most engineers through a copy-on-write client (CitC) rather than a local full clone. ## What the three layers cost The trade-off is real and worth naming honestly: all three layers are expensive to build, adopt, and maintain. Adopting `Bazel` typically means rewriting build configuration for every package in an explicit, hermetic style (no implicit classpath scanning, no ambient environment dependencies), which is a significant one-time migration cost and an ongoing discipline cost -- engineers must correctly declare dependencies or the affected-target computation silently misses things. Remote caching and CI sharding require infrastructure that someone has to run and pay for. This is precisely why smaller organizations often shouldn't reach for a 'Google-style' monorepo setup prematurely: the tooling investment only pays for itself once repository size and contributor count are large enough that naive tooling has actually become the bottleneck. ## The failure mode when the investment is skipped The failure mode when this investment is skipped is very recognizable: - CI queue times creep from minutes to tens of minutes to hours as the repo grows, because every job builds/tests everything regardless of what changed - `git clone` and `git status` become multi-minute operations that new hires dread on day one - engineers start hand-maintaining lists of 'which tests to skip,' which is fragile and drifts out of date - eventually, under enough pain, teams begin either forking pieces out into separate repos (silently reintroducing polyrepo problems) or simply stop running the full test suite locally, pushing more bugs into CI or production The Windows source-depot-to-Git migration and Google's decades of Blaze/Piper investment are the two most cited proof points that this class of tooling is not optional cosmetic polish but the actual precondition for a monorepo working at scale -- without it, 'monorepo' just means 'one big slow repository.'
- Why does adopting a build system like Bazel usually require rewriting existing build configuration rather than just installing it?Bazel's speed guarantees depend on hermetic, explicitly-declared dependencies for every target -- it needs to know precisely what each build step reads and produces so it can hash inputs and safely cache/skip work. Most pre-existing build setups (implicit classpath scanning, ambient environment variables, unlisted file reads) violate that assumption, so they have to be rewritten into explicit build targets before Bazel's caching and affected-target logic can be trusted.
- What's the risk if an engineer forgets to declare a dependency edge in a system like Bazel or Nx?The affected-target computation will miss that a change actually impacts a downstream target, so CI can report green while a real breakage ships -- the cache may even serve a stale cached result because it doesn't know the real input changed. This is why these systems usually pair with lint rules or strict/sandboxed build modes that fail loudly on undeclared dependencies rather than silently allowing them.
It's like a library with a card catalog vs. one without: with a catalog (dependency graph), a librarian fetches exactly the three books you need in seconds; without one, they have to walk every aisle in the building every time, no matter how small your request.
saying these in an interview costs you the question
- thinks 'monorepo tooling' just means using Git normally
- assumes CI time is proportional to repo size no matter what (misses affected-target computation)
- can't name any concrete tool (Bazel/Buck/Nx/Turborepo/EdenFS/VFS for Git) when asked how large orgs actually do this
- believes this tooling has no adoption cost