skip to content

How would you use evidence rather than intuition to find Common Closure Principle violations in an existing codebase, and what would you do with the findings?

level: seniorimportance: should knowfreq 22%

answer

  1. mine git log for temporal/logical coupling
  2. coupling across boundary = CCP smell
  3. median components touched per PR
  4. filter reformat/bump/rename noise
  5. fix by moving code or inverting the dependency

basics

~20 s

Mine version-control history: find files that keep changing in the same commits but live in different components. That temporal coupling is direct evidence the boundary cuts across a real axis of change, so move or merge those parts.

solid answer

~50 s

CCP is about change, and change is recorded - so measure it. From the commit log, build a co-change (temporal/logical coupling) matrix: for each pair of files, how often do they appear in the same commit or the same ticket, as a fraction of each one's total changes? High co-change across a component boundary is a CCP smell; low co-change inside a component suggests it bundles unrelated reasons to change and could be split. Complement that with per-change component fan-out - the median number of components touched per merged pull request or per ticket - which is a direct, trackable CCP metric. Filter noise: mass reformatting, dependency bumps, and renames inflate co-change, and always-together files that are simply generated pairs are not violations. Then act: move the co-changing classes together, or if they genuinely belong to different concerns, introduce an abstraction so one side stops changing. Re-measure after, because the axes of change keep moving.

code

bash · 10 lines
bash
# crude temporal-coupling probe: which file pairs change in the same commit?
git log --since=12.months --name-only --pretty=format:%H \
  | awk 'NF==0{next} /^[0-9a-f]{40}$/{c=$0; next} {print c, $0}' \
  | sort -u > commit_file.txt
# pairs within the same commit, ranked by how often they co-occur
awk '{a[$1]=a[$1]" "$2} END{for(c in a){n=split(a[c],f," ");
  if(n>30) continue;                       # skip sweeping mechanical commits
  for(i=1;i<=n;i++)for(j=i+1;j<=n;j++)print f[i]" "f[j]}}' commit_file.txt \
  | sort | uniq -c | sort -rn | head -40
# then: for each hot pair, check whether the two files live in DIFFERENT components

go deeper

for a junior

Say you'd look at the commit history for files that keep changing together but live in different packages, and that this signals the boundary is in the wrong place.

for a middle

Give the coupling ratio (shared commits over the smaller file's commit count), name the noise sources to filter, and propose moving the co-changing classes together.

for a senior

Add the fan-out-per-PR metric and lockstep-release signal, distinguish correlation from causation, and offer the dependency-inversion fix as an alternative to relocation.

for a principal

Turn it into a governance loop: track components-touched-per-change as a standing architecture metric, set a threshold that triggers a boundary review, and budget periodic boundary refactoring since change axes migrate with the roadmap.

## Why evidence beats intuition here CCP is defined in terms of *what changes together for the same reason*. That is an empirical claim about the future, but the past is a good proxy and it is fully recorded in version control. Architectural intuition tends to group by *conceptual similarity* (all things named 'validator'), which is a different and often wrong axis. ## Signal 1: temporal / logical coupling For each pair of files A and B: ``` coupling(A,B) = commits containing both A and B -------------------------------- min( commits containing A, commits containing B ) ``` Use the *min* (or a symmetric variant) so a rarely-changed file that always changes with a hot file still shows up. Interpretation: - **High coupling, different components** -> CCP violation candidate. The boundary is cutting across an axis of change. - **High coupling, same component** -> exactly what CCP wants. Good. - **Low coupling, same component** -> possible CCP-irrelevant bundling; the component may have several unrelated reasons to change and could be split (subject to CRP/REP considerations). Tools that compute this exist (co-change/temporal-coupling analysis over `git log`), but a few lines of scripting over `git log --name-only` is enough to get started. ## Signal 2: change fan-out per unit of work Simpler and more actionable for a team: for the last N merged pull requests (or tickets), record how many components each touched. - Median 1 -> CCP is holding. - Median 3-5, with the *same combination* of components recurring -> those components are really one closure region. - Track it over time as an architecture health metric; regression shows up before it becomes folklore. ## Signal 3: release/deploy coordination How often do two artifacts have to be released together to work at all? Lockstep releases are CCP violations expressed at the deployment level. Count them from the release history or the CI pipeline. ## Noise you must filter - **Sweeping mechanical commits**: reformatting, license headers, lint-rule rollouts, mass dependency bumps, framework migrations. Exclude commits above a file-count threshold or by author/bot. - **Renames and moves**: follow them (`git log --follow`, or rename detection) or a refactor looks like coupling. - **Squash vs merge policies**: squashed PRs give clean per-change granularity; long-lived branches merged wholesale blur it. Prefer PR/ticket granularity over raw commits if the history is messy. - **Legitimate always-together pairs**: a class and its unit test, a proto file and its generated stub. Same-component or by-construction pairs, not violations. - **Age**: history older than the last major restructuring may describe a design that no longer exists. Window the analysis (e.g. last 12 months). ## What to do with a confirmed violation 1. **Move code**: relocate the co-changing classes into one component. Simplest and usually right. 2. **Merge components**: if two components co-change on most changes, they are one closure region wearing two names. 3. **Introduce an abstraction**: sometimes the coupling exists because A knows a detail of B. Inverting the dependency behind a stable interface means B can change without A changing - you fix the coupling instead of relocating it. This is the OCP route, with CCP as fallback. 4. **Accept and document**: some coupling is inherent (a wire contract shared by producer and consumer). Then make it explicit: version the contract, add contract tests, and treat lockstep release as a known cost. 5. **Re-measure.** CCP is dynamic: the axes of change shift with the roadmap, so today's ideal partition is not permanent, and boundary refactoring is normal maintenance rather than an admission of past failure. ## Caveats - Co-change is **correlation, not causation**. Always read the actual commits behind a hot pair before restructuring. - Optimising blindly for co-change would collapse the system into one component; the opposing forces (CRP - don't force consumers to depend on unused code; team autonomy; independent deployability) still apply. - Absence of history (new code, freshly extracted repo) means you have no signal; fall back on domain reasoning about who requests changes and why.

  • Co-change analysis flags two files that always change together but genuinely belong to different concerns. What then?
    Don't blindly merge. Read the commits: if the coupling exists because one file knows a detail of the other, invert the dependency behind a stable abstraction so one side stops changing. That fixes the cause (OCP) rather than relocating the symptom (CCP).
  • What would make co-change data misleading?
    Mass mechanical commits (reformatting, lint rollouts, dependency bumps) inflate every pair; unfollowed renames look like coupling; long-lived branches merged in bulk blur granularity; and history predating a big restructuring describes a design that no longer exists. Filter by commit size, follow renames, use PR/ticket granularity, and window the analysis.
  • Is there a simpler metric a team can track continuously?
    Yes - the median number of components touched per merged pull request or per ticket. It is a direct proxy for CCP, easy to compute from CI, and it trends visibly when boundaries start to erode.

Desire paths in a park. Instead of arguing about where paths should go, you look at the worn tracks in the grass - the routes people actually take - and pave those. Commit history is the worn grass of a codebase.

saying these in an interview costs you the question

  • Treating co-change correlation as proof without reading the underlying commits
  • Forgetting to exclude mass reformatting, bot dependency bumps, and rename commits, which make everything look coupled
  • Concluding that the fix is always to merge components, ignoring the option of inverting the dependency so the coupling disappears
  • Optimising purely for co-change, which drives the system toward a single component and ignores CRP, ownership, and deployability
  • Assuming a partition validated once stays valid - the axes of change move, so the analysis must be repeated

context