skip to content

In a monorepo CI pipeline, what does a changed-file (path) filter do, and what does it fail to notice?

level: juniorimportance: must knowfreq 72%

answer

  1. globs against the diff, nothing more
  2. directory is not a dependency
  3. imported package changed, importer skipped
  4. lockfile and root config touch everything
  5. which base do you diff against?

basics

~20 s

A path filter compares the files a change touched against glob patterns and runs a job only when one matches. It sees file paths only, so it misses packages affected indirectly through dependency edges or shared root files.

solid answer

~50 s

A path filter is the cheapest way to stop a fifty-package monorepo from rebuilding everything on every commit. The pipeline computes the list of files the change touched — normally a diff between the incoming revision and a base branch — and a job runs only if at least one changed file matches its glob, for example `services/api/**`. It needs no extra tooling and the rule is readable right there in the pipeline definition. Its blind spot is that a repository path is not a dependency graph: if `services/api` imports `packages/auth` and only `packages/auth` changed, the api job is skipped even though api's behaviour changed. Shared root files — the lockfile, a base image reference, the root compiler or lint config — have the same problem in reverse. So path filters suit units that genuinely align with directories, and anything with real internal dependencies needs an affected-target graph instead.

go deeper

for a junior

Be able to say plainly that a path filter matches globs against the list of changed files and skips the job when nothing matches, and name the obvious gap: a package you depend on changed but your own directory did not.

for a middle

Explain how the changed-file list is produced, why the merge base is the correct comparison point for a branch, and how shared root files such as the lockfile break simple globs in both directions.

for a senior

Show that you know the failure is asymmetric: over-triggering is expensive but visible, under-triggering ships untested code silently. Talk about how you would detect skipped-but-affected jobs before an incident does it for you.

for a principal

Own the policy question: where the organization draws the line between cheap path rules and a dependency-aware affected computation, and who is accountable when the filter, rather than the code, is what let a defect through.

## The problem a path filter solves A monorepo holds many independently buildable units — packages, services, libraries — in one repository under one commit history. The naive pipeline runs every job on every commit. That is correct, and for a while it is cheap. Once the repository holds fifty packages, a one-line documentation change costs a full-fleet build, and pull-request feedback time becomes the number everyone complains about. The first response is almost always the path filter: attach a set of glob patterns to a job and run the job only when the change touched a matching file. ## How the change set is computed A filter needs two inputs: the list of files this change touched, and the patterns. The list comes from a version-control diff between the incoming revision and a base revision. Choosing the base is the part people get wrong. For a pull request the honest base is the *merge base* — the common ancestor of the branch and the target — not the previous commit on the branch, because a branch with five commits must be evaluated as a whole: ```bash # every file this branch changed relative to where it forked from main git diff --name-only origin/main...HEAD ``` The three-dot form is the merge-base diff. Using `origin/main..HEAD` or `HEAD^..HEAD` instead narrows the set to the last commit and quietly drops work done earlier on the branch. ## What a filter genuinely buys - **No extra tooling.** It is a diff and a glob; every CI platform can express it. - **Reviewability.** The rule lives in the pipeline file and shows up in the diff when someone changes it. - **Good fit for path-aligned units.** A documentation site, an infrastructure directory, a single self-contained service with no internal consumers — for those the directory really is the unit. ## The three things it cannot see **1. Dependency edges.** `services/api` depends on `packages/auth`. Editing `packages/auth` changes what `services/api` does, but no file under `services/api/` changed, so its tests are skipped. This is the classic monorepo escape: a broken change lands because the only suite that would have caught it never ran. **2. Global inputs.** The dependency lockfile, the root compiler configuration, the shared lint rules, the base container image, the CI definition itself. These affect every unit. A filter that lists only source directories skips everything when they change; a filter that adds them to every job's pattern runs everything whenever anyone touches the lockfile — which, with automated dependency updates, is most days. **3. Inputs that are not files in this repository.** A dependency version resolved at build time, an environment variable, a schema pulled from elsewhere. Nothing in the diff reflects them. ## The two symmetric failure modes Filters fail in both directions and the failures feel very different: - **Under-triggering** ships untested code. It is silent — the pipeline is green, because the job that would have failed never ran. - **Over-triggering** is loud but expensive. Patterns get widened after every escape until effectively every job runs on every change, and the mechanism has cost you complexity without buying speed. The pull toward over-triggering is strong precisely because under-triggering is invisible, so teams widen patterns after every incident and never narrow them again. ## Base and history traps Two mechanical traps bite even when the patterns are right. First, a depth-limited clone: if CI clones only the tip commit to save time, the base commit is not in the local history, and the diff either errors or returns a wrong set. Second, the base ref must actually be fetched — a checkout that fetches only the branch has no `origin/main` object to compare against. ## When to graduate Use path filters while the units are genuinely independent and directory-shaped. Move to a dependency-aware affected-target computation the moment packages import each other, which in a real monorepo is almost immediately. A reasonable interim rule: keep filters for the obviously isolated pieces (docs, infrastructure), and treat anything under the shared source tree with a tool that reads the dependency graph. One more habit worth adopting either way: make the *change of the filter itself* trigger the job it guards. Otherwise the only edit guaranteed not to be tested is the edit to the thing that decides what gets tested.

  • Why does diffing against the previous commit instead of the merge base under-trigger jobs?
    A branch is reviewed and merged as a whole, so the pipeline must consider every file the branch changed. Diffing against the previous commit sees only the last commit's files, so work done in earlier commits on the same branch never matches a filter, and the job that should validate it is skipped. Use the merge base with the target branch, which `git diff --name-only origin/main...HEAD` gives you.
  • How should a path filter handle a change to the dependency lockfile?
    Treat it as a global input: a lockfile change can alter the resolved dependency tree of any package, so the safe response is a full build. Adding the lockfile to every job's pattern achieves that but makes automated dependency-update pull requests as expensive as a full run, which is one of the reasons teams move to a tool that resolves which packages a lockfile change actually affects.
  • A CI job's path filter looks right but the diff comes back empty. What would you check first?
    The clone. A depth-limited clone has no common ancestor with the base branch, and a checkout that fetched only the pull-request branch has no base ref at all, so the diff either errors or produces an empty or bogus list. Fetch enough history to resolve the merge base before computing the change set.

saying these in an interview costs you the question

  • Assumes a directory boundary is a dependency boundary
  • Diffs against the last commit instead of the merge base
  • Widens globs after every escape until nothing is skipped
  • Thinks a green pipeline means the relevant tests ran
  • Forgets the lockfile and root config affect everything

context