skip to content

How do you tell real (knowledge) duplication from coincidental (incidental) duplication, and what happens when you get that judgement wrong in each direction?

level: middleimportance: must knowfreq 68%

answer

  1. same reason + same owner = real
  2. different reasons = coincidence, leave it
  3. wrong abstraction → flag creep → ripple
  4. divergent change is silent
  5. can't unify → derive, else assert agreement

basics

~20 s

Ask whether the copies must always change together for the same reason. If yes, it's real duplication — unify it. If they could evolve separately, it's coincidence — leave them. Merging coincidence creates a shared thing that soon needs flags.

solid answer

~50 s

Real duplication means one rule is encoded in several places, so any change must be applied everywhere or the system becomes inconsistent. Coincidental duplication means several independent rules currently look alike. The discriminator is the *reason to change*: same reason and same owner → real; different owners or different triggers → coincidental. Getting it wrong in the DRY direction (merging coincidence) produces the wrong abstraction: the shared unit accumulates boolean/enum parameters and branches, unrelated features start breaking each other, and unwinding it is harder than the original copy-paste. Getting it wrong in the WET direction (leaving real duplication) produces divergent change: a bug fixed in three of four copies, a rule updated in the service but not the database, contradictory behaviour that no test catches. Practical tactics: look at git history — copies that always change in the same commit are real; consider ownership boundaries; and when you truly cannot unify, add a test asserting the copies agree.

go deeper

for a junior

Show you know duplication isn't automatically bad: give the "must they change together?" test and one example of each kind.

for a middle

Narrate both failure modes concretely — flag creep in an over-merged helper versus a bug fixed in 3 of 4 copies — and mention hidden cross-layer duplication.

for a senior

Add the economics: duplication is a local cost, wrong abstraction is a global coupling cost; cite Metz's re-inline remedy and use ownership/reason-to-change (SRP) as the discriminator.

for a principal

Discuss making drift structurally impossible — contract-first generation, CI consistency assertions — and policy for when teams should be allowed to duplicate across service boundaries to protect autonomy.

## Two things that look the same on screen **Real (knowledge) duplication** — one fact, many encodings. Change the fact and every encoding must move in lockstep, or the system contradicts itself. **Coincidental (incidental) duplication** — many facts that currently have the same shape. Nothing links them except present-day similarity. The screen cannot tell you which you are looking at, because both render as "these two blocks are the same". You must reason about *why the code exists*. ## The discriminating questions 1. **Same reason to change?** If the product owner says "raise the limit", do both copies move? Always, or only sometimes? 2. **Same owner/actor?** Robert Martin's SRP framing: code requested by *different stakeholders* should be allowed to diverge. Two teams' rules that coincide are not one rule. 3. **Same concept name?** If you cannot name the shared abstraction without a vague word ("Helper", "Utils", "Manager", "Common", "processData"), you probably do not have one concept. 4. **What does history say?** `git log` on both files: if edits to A and B land in the same commits repeatedly, that is empirical evidence of coupled knowledge. If they have never moved together in two years, they are independent. 5. **Would a change to one *require* re-testing the other?** If unifying them means feature X's release now needs feature Y's regression suite, ask whether that coupling is justified. ## Failure mode A — merging coincidence (over-DRY) The lifecycle is predictable and worth being able to narrate in an interview: 1. Two similar blocks are extracted into `sharedThing()`. 2. A requirement changes one caller only. Rather than un-extract, someone adds a parameter: `sharedThing(x, forNickname = true)`. 3. More divergence, more flags: `sharedThing(x, strict, legacyMode, skipTrim)`. The body becomes a decision tree, and each branch is exercised by only one caller. 4. The function now has more paths than either original had lines, callers must reason about combinations they never use, and a change for caller A breaks caller B (ripple). 5. Everyone is afraid of it. It becomes a permanent tax. Sandi Metz's summary: **"duplication is far cheaper than the wrong abstraction."** Her recommended remedy once you're in step 4 is counterintuitive but correct: *re-inline* — push the code back into the callers, then re-extract only what is genuinely shared, if anything. Why is the wrong abstraction more expensive than duplication? Because duplication is a *local* problem (fix it where it is, in a file you already understand), while a bad shared dependency is a *global* problem (any fix touches every caller, requires broad regression, and often crosses team boundaries). Duplication costs you edits; wrong coupling costs you coordination. ## Failure mode B — leaving real duplication (under-DRY) Symptoms: - A bug fixed in three of four copies; the fourth resurfaces months later, often in the least-tested path. - A validation rule enforced by the API but not the batch importer, so bad data enters through the side door. - An enum in code and a lookup table in the database that drift, producing rows no branch handles. - Docs/SDKs that describe an older version of the contract. The insidious property is that nothing fails loudly. Divergence is silent; the system just becomes subtly wrong, and each incident is diagnosed as a one-off bug instead of a structural defect. ## Hidden duplication (the expensive kind) Most real duplication does not look duplicated. Common shapes: - **Across layers**: DB constraint + ORM annotation + service validator + client form + API docs, all restating one rule. - **Across representations**: an entity struct plus a hand-written serializer/parser that must list every field. - **Across artifacts**: port numbers, timeouts, and feature flags in code, Dockerfile, CI config, and IaC. - **Between code and comments/docs**: a comment describing behaviour that the code no longer has. Comments are duplication with no compiler to check them, which is why "comment *why*, not *what*" is partly a DRY rule. - **Between production code and test setup**: a hand-rolled fixture that reimplements a construction rule. ## When you cannot unify: make drift loud Sometimes both encodings must physically exist — a database constraint for integrity, a client check for UX latency, a generated SDK on another runtime. The DRY-preserving hierarchy is: 1. **Derive**: one artifact is the source, the others are generated (schema → validators, DTOs, docs, clients). 2. **Verify**: if you cannot generate, write an automated test that asserts the copies agree (e.g. a test that reads the DB constraint and the app-level rule and compares them). Drift becomes a red build, not a production incident. 3. **Annotate**: as a last resort, cross-reference comments at every copy. Weakest, because nothing enforces it — but still better than nothing. ## Deliberate duplication Recognised cases where copying is the *right* call: - **Tests**: readability beats compression; a test that inlines its expected values is a better specification than one that computes them via the same helper the production code uses (which can make the test tautological). - **Across service/team boundaries**: a shared business-logic library couples independent release cycles; small duplication buys autonomy. - **Early in a design**, before the axis of variation is known: waiting is cheap, guessing wrong is not. ## The one-line judgement to give in an interview > "I deduplicate *knowledge*, not text. My test is whether the copies must always change together for the same reason and the same stakeholder. If I'm unsure, I wait — duplication is a local cost, the wrong abstraction is a global one."

  • You inherit a shared helper with five boolean parameters, each used by exactly one caller. What do you do?
    Re-inline it: copy the relevant branch back into each caller, delete the flags, then look again for a genuinely shared core. This is Metz's prescribed cure for a wrong abstraction — go back through duplication rather than adding a sixth flag.
  • What evidence, other than reading the code, helps decide whether two similar blocks encode the same knowledge?
    Version-control history: if the two files repeatedly change in the same commit for the same ticket, the knowledge is coupled. If they have evolved independently for a long time, they are separate concerns. Ownership (CODEOWNERS / team boundaries) is a second signal.
  • Give an example of duplication that grep will never find.
    A length limit stated as a database CHECK constraint, an application-level validator, a regex in a web form, and a sentence in the public API docs. No shared text, one rule, four places to update.

Two neighbours both painted their fence green this year. That doesn't mean they should buy one shared bucket of paint forever — next year one wants blue, and now they're negotiating. But a single property line surveyed twice, with two different numbers on file, is a genuine problem: there is only one boundary, and the records must agree.

context