skip to content

questions

6

In a monorepo CI pipeline, what does a changed-file (path) filter do, and what does it fail to notice?

level: juniorimportance: must knowfreq 72%

answer

  1. globs against the diff, nothing more
  2. directory is not a dependency
  3. imported package changed, importer skipped
  4. lockfile and root config touch everything
  5. which base do you diff against?

basics

~20 s

A path filter compares the files a change touched against glob patterns and runs a job only when one matches. It sees file paths only, so it misses packages affected indirectly through dependency edges or shared root files.

solid answer

~50 s

A path filter is the cheapest way to stop a fifty-package monorepo from rebuilding everything on every commit. The pipeline computes the list of files the change touched — normally a diff between the incoming revision and a base branch — and a job runs only if at least one changed file matches its glob, for example `services/api/**`. It needs no extra tooling and the rule is readable right there in the pipeline definition. Its blind spot is that a repository path is not a dependency graph: if `services/api` imports `packages/auth` and only `packages/auth` changed, the api job is skipped even though api's behaviour changed. Shared root files — the lockfile, a base image reference, the root compiler or lint config — have the same problem in reverse. So path filters suit units that genuinely align with directories, and anything with real internal dependencies needs an affected-target graph instead.

go deeper

for a junior

Be able to say plainly that a path filter matches globs against the list of changed files and skips the job when nothing matches, and name the obvious gap: a package you depend on changed but your own directory did not.

for a middle

Explain how the changed-file list is produced, why the merge base is the correct comparison point for a branch, and how shared root files such as the lockfile break simple globs in both directions.

for a senior

Show that you know the failure is asymmetric: over-triggering is expensive but visible, under-triggering ships untested code silently. Talk about how you would detect skipped-but-affected jobs before an incident does it for you.

for a principal

Own the policy question: where the organization draws the line between cheap path rules and a dependency-aware affected computation, and who is accountable when the filter, rather than the code, is what let a defect through.

## The problem a path filter solves A monorepo holds many independently buildable units — packages, services, libraries — in one repository under one commit history. The naive pipeline runs every job on every commit. That is correct, and for a while it is cheap. Once the repository holds fifty packages, a one-line documentation change costs a full-fleet build, and pull-request feedback time becomes the number everyone complains about. The first response is almost always the path filter: attach a set of glob patterns to a job and run the job only when the change touched a matching file. ## How the change set is computed A filter needs two inputs: the list of files this change touched, and the patterns. The list comes from a version-control diff between the incoming revision and a base revision. Choosing the base is the part people get wrong. For a pull request the honest base is the *merge base* — the common ancestor of the branch and the target — not the previous commit on the branch, because a branch with five commits must be evaluated as a whole: ```bash # every file this branch changed relative to where it forked from main git diff --name-only origin/main...HEAD ``` The three-dot form is the merge-base diff. Using `origin/main..HEAD` or `HEAD^..HEAD` instead narrows the set to the last commit and quietly drops work done earlier on the branch. ## What a filter genuinely buys - **No extra tooling.** It is a diff and a glob; every CI platform can express it. - **Reviewability.** The rule lives in the pipeline file and shows up in the diff when someone changes it. - **Good fit for path-aligned units.** A documentation site, an infrastructure directory, a single self-contained service with no internal consumers — for those the directory really is the unit. ## The three things it cannot see **1. Dependency edges.** `services/api` depends on `packages/auth`. Editing `packages/auth` changes what `services/api` does, but no file under `services/api/` changed, so its tests are skipped. This is the classic monorepo escape: a broken change lands because the only suite that would have caught it never ran. **2. Global inputs.** The dependency lockfile, the root compiler configuration, the shared lint rules, the base container image, the CI definition itself. These affect every unit. A filter that lists only source directories skips everything when they change; a filter that adds them to every job's pattern runs everything whenever anyone touches the lockfile — which, with automated dependency updates, is most days. **3. Inputs that are not files in this repository.** A dependency version resolved at build time, an environment variable, a schema pulled from elsewhere. Nothing in the diff reflects them. ## The two symmetric failure modes Filters fail in both directions and the failures feel very different: - **Under-triggering** ships untested code. It is silent — the pipeline is green, because the job that would have failed never ran. - **Over-triggering** is loud but expensive. Patterns get widened after every escape until effectively every job runs on every change, and the mechanism has cost you complexity without buying speed. The pull toward over-triggering is strong precisely because under-triggering is invisible, so teams widen patterns after every incident and never narrow them again. ## Base and history traps Two mechanical traps bite even when the patterns are right. First, a depth-limited clone: if CI clones only the tip commit to save time, the base commit is not in the local history, and the diff either errors or returns a wrong set. Second, the base ref must actually be fetched — a checkout that fetches only the branch has no `origin/main` object to compare against. ## When to graduate Use path filters while the units are genuinely independent and directory-shaped. Move to a dependency-aware affected-target computation the moment packages import each other, which in a real monorepo is almost immediately. A reasonable interim rule: keep filters for the obviously isolated pieces (docs, infrastructure), and treat anything under the shared source tree with a tool that reads the dependency graph. One more habit worth adopting either way: make the *change of the filter itself* trigger the job it guards. Otherwise the only edit guaranteed not to be tested is the edit to the thing that decides what gets tested.

  • Why does diffing against the previous commit instead of the merge base under-trigger jobs?
    A branch is reviewed and merged as a whole, so the pipeline must consider every file the branch changed. Diffing against the previous commit sees only the last commit's files, so work done in earlier commits on the same branch never matches a filter, and the job that should validate it is skipped. Use the merge base with the target branch, which `git diff --name-only origin/main...HEAD` gives you.
  • How should a path filter handle a change to the dependency lockfile?
    Treat it as a global input: a lockfile change can alter the resolved dependency tree of any package, so the safe response is a full build. Adding the lockfile to every job's pattern achieves that but makes automated dependency-update pull requests as expensive as a full run, which is one of the reasons teams move to a tool that resolves which packages a lockfile change actually affects.
  • A CI job's path filter looks right but the diff comes back empty. What would you check first?
    The clone. A depth-limited clone has no common ancestor with the base branch, and a checkout that fetched only the pull-request branch has no base ref at all, so the diff either errors or produces an empty or bogus list. Fetch enough history to resolve the merge base before computing the change set.

saying these in an interview costs you the question

  • Assumes a directory boundary is a dependency boundary
  • Diffs against the last commit instead of the merge base
  • Widens globs after every escape until nothing is skipped
  • Thinks a green pipeline means the relevant tests ran
  • Forgets the lockfile and root config affect everything

context

open as a page

In a monorepo build tool such as Nx, Turborepo or Bazel, how is the "affected" set of targets computed, and why is that stronger than a path filter?

level: middleimportance: must knowfreq 62%

basics

~20 s

The tool builds a dependency graph of the repository's projects, maps the changed files to the projects that own them, then walks the graph backwards to include every project that transitively depends on those. Path filters see only the directly edited directories.

open as a page

In a monorepo that publishes many packages, what is the difference between giving every package one locked version and versioning each package independently?

level: middleimportance: should knowfreq 46%

basics

~20 s

Locked (fixed) mode gives every package the same version number and republishes them together on each release. Independent mode bumps and publishes only the packages that changed, each on its own version line, at the cost of far more bookkeeping.

open as a page

A monorepo pipeline that builds only affected packages let a broken package reach the main branch with its tests never running. What causes would you investigate?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Investigate the change set before the graph: a wrong or unavailable diff base, a shallow clone with no common ancestor, and a squash or rebase that moved the fork point. Then check for a real dependency edge the graph never had.

open as a page

As the owner of a large monorepo's CI, how far would you trust the affected-target graph as the only pre-merge gate, and what backstop would you run?

level: principalimportance: should knowfreq 33%

basics

~20 s

Trust it in proportion to how the edges are derived: enforced declarations earn near-total trust, inferred imports earn conditional trust. Either way run a periodic full build so an escape is bounded by that interval rather than discovered by a customer.

open as a page

Some monorepo teams generate the CI pipeline from the affected package set instead of writing every job statically. What does that buy, and what does it cost?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

Generation keeps the pipeline proportional to the change: a small setup job computes the affected packages and emits a child pipeline containing only those jobs. The cost is that the definition no longer exists in the repository — it is program output, so it is harder to review, reason about and debug.

open as a page