A monorepo pipeline that builds only affected packages let a broken package reach the main branch with its tests never running. What causes would you investigate?
answer
- a green run that ran nothing
- two-dot versus three-dot base
- depth-1 clone has no ancestor
- squash moved the fork point
- empty change set must fail closed
basics
~20 sInvestigate the change set before the graph: a wrong or unavailable diff base, a shallow clone with no common ancestor, and a squash or rebase that moved the fork point. Then check for a real dependency edge the graph never had.
solid answer
~50 sSplit the pipeline into its two stages and check them in order. Stage one is the change set: what base was the diff taken against, was it the merge base or just the previous commit, was the base ref actually fetched, and was the clone deep enough to contain the common ancestor? A depth-limited clone or an unfetched base ref typically degrades into an empty or wrong file list, which makes the affected set silently small. Squash merges are a specific trap: after a squash, the branch's commits are not ancestors of main, so a later run's fork point is not where anyone thinks it is. Stage two is the graph: was there a real edge from the changed package to the broken one — a runtime call, a dynamically loaded plugin, generated code, a fixture — that the analyzer could not see? Reproduce by printing the computed change set and affected list for that exact revision, rather than reasoning about it.
code
bash · 26 lines#!/usr/bin/env bash
set -euo pipefail
BASE_REF="origin/main"
# Fail closed: if we cannot resolve a base, build everything rather than nothing.
if ! git rev-parse --verify --quiet "$BASE_REF" >/dev/null; then
echo "base ref $BASE_REF not fetched - falling back to a full build" >&2
echo "AFFECTED_MODE=all"
exit 0
fi
if ! MERGE_BASE=$(git merge-base "$BASE_REF" HEAD 2>/dev/null); then
echo "no common ancestor (shallow clone?) - falling back to a full build" >&2
echo "AFFECTED_MODE=all"
exit 0
fi
CHANGED=$(git diff --name-only "$MERGE_BASE" HEAD)
if [ -z "$CHANGED" ]; then
echo "empty change set at $(git rev-parse HEAD) - treating as suspicious" >&2
echo "AFFECTED_MODE=all"
exit 0
fi
printf '%s\n' "$CHANGED"go deeper
Understand the basic shape of the failure: the pipeline decided nothing relevant changed, so the test that would have caught the bug never ran and the run was green anyway.
Be able to list the mechanical causes in order — wrong diff base, missing history in the clone, an unfetched base ref — and explain why the three-dot merge-base diff is the correct comparison.
Show the diagnostic discipline: reproduce the computation at the failing revision, check change set before graph, then distinguish a missing dependency edge from a task that was simply never wired to run.
Own the systemic answer: an affected-only pipeline converts loud failures into silent ones, so design for fail-closed behaviour, recorded evidence of what each commit tested, and a bounded window in which an escape can hide.
## Diagnose in pipeline order, not in guess order An affected-only pipeline computes three things in sequence: a **change set** (which files), an **affected set** (which projects), and a **task list** (which jobs). A silent escape means one of the three came back smaller than reality. Work forward through them, because the earlier a stage is wrong, the more invisible the failure — a wrong change set produces a small affected set that looks perfectly plausible in the logs. The single most useful first move is to stop reasoning and print. Re-run the computation at the exact merged revision and dump the change set and the affected list. Most investigations end here, because the change set is visibly wrong. ## Cause 1 — the diff base The change set is a diff, and a diff needs a base. Failure modes, in rough order of frequency: - **Previous commit instead of merge base.** A branch of five commits is evaluated as a whole, but the pipeline diffed `HEAD^..HEAD`, so only the last commit's files counted and the earlier changes never marked anything affected. - **Two-dot instead of three-dot.** `origin/main..HEAD` compares tips; commits that landed on main after the fork point pollute the set in one direction and mask it in another. `origin/main...HEAD` is the merge-base form. - **A base branch that moved.** If the pipeline diffs against `origin/main` fetched at job start, and main advanced past the fork point, the computed set drifts. ## Cause 2 — history that is not present CI checkouts are frequently depth-limited for speed. A depth-1 clone contains no common ancestor with the base branch, so the merge base cannot be resolved: ```bash git clone --depth 1 https://example.com/monorepo.git cd monorepo # no shared history and no base ref present: the base cannot be resolved, # and a pipeline that swallows the error proceeds with an empty change set git merge-base origin/main HEAD ``` The dangerous variant is a script that tolerates the failure — a `|| true`, an unchecked exit code — and continues with an empty list. Nothing is affected, every job is skipped, the run is green in seconds. Fast pipelines that got fast overnight deserve suspicion for exactly this reason. ## Cause 3 — the fork point after a squash or rebase Squash merging replaces a branch's commits with one new commit on main. The original commits are not ancestors of main, so the merge base for any later branch cut from that history is not where a human would point. The same applies to rebase-merge workflows. The symptom is an affected set that is inexplicably large or small on the *next* change rather than the current one, and it is easy to misattribute to the graph. ## Cause 4 — a real edge the graph never had Only once the change set is proven correct is the graph the suspect. Ask what connects the changed package to the broken one: - **Runtime-only coupling.** Two services communicate over the network; there is no import to infer. - **Dynamic resolution.** A plugin loaded by name from configuration, a module resolved from a string. - **Generated code or data.** A schema, a fixture, a generator whose output is committed — an edge exists only if someone declared it. - **Undeclared shared inputs.** A tool version file or shared configuration that the tool does not treat as an input to every project. ## Cause 5 — the task list, not the affected set Sometimes the affected set was right and the *task* was not run: the test target was not wired into the job that runs on pull requests, the tool ran `-t build` but not `-t test`, or the job was allowed to fail without failing the run. Check whether the package appeared in the affected list and its test target simply produced no execution record. ## Making the class of failure visible After the fix, add detection, because the next instance will be just as silent: 1. **Record the computation.** Persist the change set and affected list as a run output for every pipeline, so the question "what did we test for this commit?" is answerable after the fact. 2. **Fail closed.** If the base cannot be resolved, build everything rather than nothing. An empty change set should be treated as an error condition, never as "no work to do". 3. **Run a periodic full build.** A scheduled build-and-test of everything on the main branch catches escapes within a bounded window rather than at the next customer report. 4. **Assert the checkout.** Verify the base ref exists and the merge base resolves before the diff runs, and fail loudly if not. The framing to bring to an interview: an affected-only pipeline converts a loud failure (a red build) into a silent one (a build that never happened), so the engineering work is not only computing the set correctly but making a wrongly-empty set impossible to mistake for success.
- Why is an empty change set more dangerous than an over-broad one?Over-broad costs money and time and is immediately visible in the run duration. Empty is indistinguishable from "this change genuinely touched nothing": the pipeline goes green in seconds and the merge proceeds. Because nothing failed, no one investigates. That is why the correct default on a base-resolution error is to build everything, not to skip everything.
- How does squash merging interact with affected-set computation?Squash merge collapses a branch into a single new commit on main, so the branch's original commits are not ancestors of main. Merge-base calculations for later branches therefore resolve to a different point than people expect, and the computed change set for the next change can be far larger or smaller than intended. Teams normally handle it by diffing against the last successfully built commit on main rather than inferring the fork point.
- What would you add to the pipeline so this failure is detectable next time?Persist the computed change set and affected list as a run artifact so the question "what did we actually test for this commit?" is answerable later. Fail closed when the base ref or merge base cannot be resolved. And schedule a periodic full build on main, which bounds how long an escape can hide to the interval between those runs.
- The change set is provably correct but the package still was not tested. Where do you look next?Two places. The graph: was there ever an edge from the changed package to this one, or is the coupling runtime, dynamic, or through generated code the analyzer cannot see? And the task list: the package may have been in the affected set while only the build target ran, or its test job was configured not to fail the run.
saying these in an interview costs you the question
- Blames the graph before checking the diff base
- Treats an empty change set as no work to do
- Assumes a shallow clone is harmless for diffs
- Adds a full build on every commit and calls it fixed
- Reasons about the change set instead of printing it