When static analysis is scoped to only the files a change touched, what escapes the gate?
answer
- Fast because it reads less
- The violation need not be in the diff
- Which commit is the baseline?
- A rule change implicates untouched files
- Schedule the full scan anyway
basics
~20 sAnything whose violation lands outside the diff: a change that breaks an untouched caller, findings that need whole-program context, and every file affected by a newly enabled rule. Diff scoping trades completeness for latency, deliberately.
solid answer
~50 sScoping to changed files buys speed - seconds instead of minutes on every change - and it costs three things. First, violations your change causes elsewhere: removing the last caller of a function reports in the callee's file, and signature or constant changes surface in files you never opened. Second, rules that reason across files; given a subset, some analysers silently weaken their inference, so a subset run is not simply a smaller full run. Third, anything a ruleset change affects, since enabling a rule mostly implicates untouched files. Two mechanics matter: diff against the merge base of the target branch, or most of a long branch is never analysed, and scope by file rather than filtering findings by changed line, because analysers report at the site they consider responsible. The workable answer is hybrid - diff-scoped per change, full scope when the ruleset changes, and a scheduled full scan.
code
pseudocode · 11 linesbase = merge_base(target_branch, head) # not head~1
changed = changed_files(base, head)
if config_changed(base, head) or analyser_version_changed(base, head):
scope = all_source_files() # a rule change touches everything
else:
scope = changed + reverse_dependents(changed)
report = analyse(scope)
block_merge_if(report.errors_attributable_to(changed))
record_trend(report.total_count())go deeper
Be ready to say why teams analyse only changed files at all - a full pass over a large repository is too slow to run on every change - and to name one thing it can miss, such as a violation your edit causes in a file you did not touch.
Explain the mechanics: how the changed set is computed, why the merge base is the right baseline, and why enabling a new rule has to widen the scope because the affected files are precisely the ones nobody edited.
Show that you know a subset run is not a smaller full run - cross-file rules can weaken or change their findings - and describe the hybrid you would actually operate, including where the full scan runs and who triages it. Keep analysis off the pipeline's critical path.
Own the trade explicitly: latency bought with completeness, and the residual risk when the rules that matter most are the cross-file ones scoping disables. Decide what the organisation blocks on per change versus what it tracks as a trend.
### Why anyone scopes analysis to the diff Running a full analysis over a large repository can take longer than the tests. On a quote engine whose whole pipeline budget is 27 minutes, a whole-repository analysis that takes six of them is a real cost paid on every change, most of which touches four files. Scoping the analyser to the files a change touched cuts that to seconds, which is why almost every team eventually does it. The question in an interview is whether you know what you gave up. ### What escapes a diff-scoped run **1. Violations your change causes elsewhere.** Analysis is not always local. Delete the last caller of a function and the unused-symbol finding lands in the *callee's* file, which your change never touched. Change a signature, a constant, an interface or a configuration key and the resulting violations appear in files outside the diff. A diff-scoped run reports nothing, and the repository is now dirtier than the gate believes. **2. Findings that need whole-program context.** Rules that reason across files — reachability, dataflow, dead-code detection, cross-module dependency rules — are not simply "the same rule applied to fewer files". Given a subset, some analysers silently weaken their inference, which changes the result rather than merely reducing it. A subset run can therefore report a finding that a full run would not, and vice versa, and neither is a bug. **3. Anything the ruleset newly cares about.** Enable a rule, upgrade the analyser, or change a per-path override, and the affected files are overwhelmingly files nobody has touched. A pure diff-scoped gate never sees them. The scope rule has to be conditional: if the configuration or the analyser version changed in this change, widen the scope — either to the whole repository or at least to everything the changed configuration governs. **4. Whatever the baseline choice hides.** Diffing against the previous commit rather than the merge base of the target branch means only the last commit of a long-lived branch is ever analysed; the eleven commits before it merge unanalysed. Rebases and force-pushes move the baseline underneath you. Use the merge base, so the scope is everything the change introduces relative to what it will merge into. **5. Line-level filtering artefacts.** A tempting refinement is to keep only findings whose reported line lies inside the diff. But analysers report at the line they consider the site of the problem: a file-level rule reports at line one, a rule about a declaration reports at the declaration even when the offending edit is two hundred lines below, and a formatting rule may report at the first line of a reformatted block. Filter by line and you will drop genuine findings caused by the change. Scope the *analysis* by file; decide attribution separately and generously. **6. Time.** A repository analysed only where it is edited drifts in the parts nobody edits. Files that have not changed in a year have never been seen by any rule added since. ### The shape of a workable answer The pattern that survives contact with a real repository is a hybrid: - **Per-change:** analyse the changed files plus, where the analyser supports it, their reverse dependents. Block the merge on findings attributable to this change. Keep it fast enough to run in parallel with the tests rather than adding to the critical path. - **On configuration or version change:** widen automatically to the full scope, because a rule change is a change to every file. - **On a schedule:** run the full analysis, on a cadence the team will actually triage — nightly or weekly. Route those findings to work items rather than to the author of an unrelated change, and treat a growing count as a signal about the ruleset, not only about the code. How you handle the pile of pre-existing findings a full scan surfaces is a separate mechanism with its own tradeoffs, and it is not what diff scoping is for: diff scoping decides *what the analyser reads*, not *which known findings are forgiven*. Conflating the two is a common interview stumble — a candidate who says "we only analyse changed files so old issues do not block us" has described the wrong tool for that job and has, as a side effect, made the gate blind to their own change's remote consequences. Finally, be honest about the residual risk. Diff-scoped analysis is a deliberate trade of completeness for latency. It is the right trade for most teams. It becomes the wrong one when the rules that matter most — the cross-file ones — are exactly the rules the scoping disables.
- Why is the previous commit the wrong baseline for a diff-scoped run?Because only the last commit gets analysed. On a branch with eleven commits, the first ten merge unexamined, and a rebase or force-push moves the baseline underneath you so the scope changes for reasons unrelated to the code. Diff against the merge base of the target branch, so the scope is everything the change introduces relative to what it will merge into.
- What breaks if you keep only findings whose reported line falls inside the diff?You drop genuine findings. Analysers report at the site they consider responsible: a file-level rule reports at line one, a declaration rule reports at the declaration even when the offending edit is two hundred lines below, and a reformatting rule reports at the start of the block. Scope the analysis by file and decide attribution generously, rather than filtering by line.
- How would you keep the parts of the repository nobody edits from drifting?Run the full analysis on a schedule the team will actually triage - nightly or weekly - and route the findings to work items rather than onto the author of an unrelated change. Widen automatically whenever the ruleset or analyser version changes, since that is when untouched files newly violate, and watch the total count as a trend rather than as a gate.
saying these in an interview costs you the question
- Assumes a subset run finds the same defects as a full run
- Diffs against the previous commit instead of the merge base
- Never re-runs full scope after a ruleset change
- Confuses diff scoping with forgiving known issues
- Filters findings to changed lines and calls it attribution
- Treats faster feedback as equivalent to better coverage