skip to content

How do you decide when to extract an abstraction from duplicated code versus leaving the duplication in place, and what is meant by 'a wrong abstraction is more expensive than duplication'?

level: middleimportance: should knowfreq 40%

answer

  1. DRY is about knowledge, not identical characters
  2. coincidental duplication changes for different reasons
  3. rule of three / AHA — avoid hasty abstractions
  4. Metz: duplication is cheaper than the wrong abstraction; cure = re-inline
  5. flags inside shared code = wrong abstraction forming

basics

~20 s

Extract when the pieces are the same idea, not merely the same text, and when they change together for the same reason. If they only look alike, keep the duplication — a shared abstraction forced to serve two different reasons grows flags and branches and becomes harder to remove than the copies were.

solid answer

~60 s

DRY is about knowledge, not characters: every piece of *knowledge* should have one authoritative representation. Two identical-looking blocks that change for different reasons are **coincidental duplication** and should stay separate. The wrong-abstraction failure mode (Sandi Metz): someone unifies two similar cases; a third case is almost the same, so a boolean parameter appears; then another, then a special case inside the shared function. Callers pass flags they do not understand, and every change risks unrelated call sites. Unpicking it later is harder than deleting a duplicate would have been, because the code is now shared across teams and tests. Practical rules: wait for the third occurrence (rule of three) so the axis of variation is visible; unify on *reason to change* (Single Responsibility / Common Closure) rather than shape; prefer parameterising data over adding behavioural flags; if a shared abstraction starts sprouting mode flags, inline it back and re-split along the real seam. Duplication across module or team boundaries is usually safer than a shared abstraction that couples them.

code

pseudocode · 6 lines
pseudocode
// Wrong abstraction forming — flags encode two different reasons to change
function applyRate(amount, rate, isTax, skipRounding, legacyMode) { ... }

// Re-split along the real seam
function taxOn(amount, taxRate) { ... }
function discountOn(amount, discountRate) { ... }

go deeper

for a junior

Say DRY is about repeated knowledge, not repeated text; wait until you see the pattern a few times; flags in shared code are a warning sign.

for a middle

Use 'do they change for the same reason?' as the test, cite the rule of three, and describe how flag creep produces a wrong abstraction.

for a senior

Add reversibility reasoning (extracting later is mechanical, un-picking is not), the re-inline cure, distance/ownership considerations, and how to pick the abstraction level (too low = leaky, too high = over-general).

for a principal

Frame it as coupling policy: where deduplication is mandated (wire contracts, security rules), where it is discouraged (across team boundaries), and how shared-library governance and release coupling change the calculus at org scale.

## The two failure modes You can be wrong in two directions: - **Under-abstraction**: the same knowledge is encoded in five places; a rule change requires five edits, one is missed, and behaviour diverges silently. - **Over/wrong abstraction**: unrelated cases are forced into one shape; the shared code accumulates flags, and a change for caller A breaks caller B. The second is the more expensive mistake because it is *coupling* — it grows, spreads across teams and is hard to reverse. Duplication is a local, visible cost; wrong coupling is a distributed, invisible one. ## DRY, correctly stated Hunt and Thomas: *every piece of knowledge must have a single, unambiguous, authoritative representation within a system.* The unit is **knowledge**, not lines. Two functions with identical bodies that implement different rules (VAT calculation and a discount cap that happen to both be `amount * rate`) encode different knowledge; merging them creates a false link. Useful reframing: **the axis of change is what you're deduplicating.** If A and B always change together for the same reason, they are one thing. If they change independently, they are two things that currently look alike. ## Sandi Metz's sequence — how wrong abstractions form 1. Two similar code paths are unified into a shared function. 2. A third caller is *almost* the same → a boolean parameter is added. 3. A fourth is *almost* the same → a second parameter, or an `if` inside. 4. Nobody understands the parameter combinations; each change is risky. 5. New developers, reluctant to break others, add yet another flag. Her conclusion: *duplication is far cheaper than the wrong abstraction* — and when you inherit one, the cure is to **re-inline** the shared code back into each caller, delete the unused branches, and only then look for the real, smaller commonality. ## Heuristics that actually help - **Rule of three.** With two examples you cannot see which parts vary; with three the axis of variation is usually visible. (Related idea: AHA — 'avoid hasty abstractions', prefer duplication over the wrong abstraction.) - **Same reason to change?** Single Responsibility Principle at function level; Common Closure Principle at package level. Group by reason to change, not by shape. - **Data over flags.** Passing a rate, a strategy object or a configuration record is far less corrosive than a `boolean isB` that switches behaviour inside. - **Distance matters.** Duplication inside one small file is cheap and easy to fix later. Duplication across services owned by different teams is *often preferable* to a shared library that couples deploy schedules — this is why microservice guidance frequently tolerates copied code. - **Test duplication is different.** Tests deliberately restate expectations; over-abstracting them hides what is being asserted and makes failures unreadable. - **Reversibility.** Extracting later is a mechanical refactor; un-picking a bad shared abstraction with many callers is not. When uncertain, choose the more reversible option — duplication. ## Choosing the *level* once you do abstract - **Too low** (leaky, mechanism-flavoured): the caller still juggles mechanism details, e.g. an interface exposing HTTP status codes and headers to business code. - **Too high** (over-general): a 'generic pipeline framework' configured by a DSL to do the one thing you needed — huge surface, no user but you. - **Right**: caller vocabulary is domain-shaped, the interface is small relative to what it hides, and the abstraction has at least two genuinely different real implementations or uses. A useful check: could you explain the abstraction's purpose in one sentence without using 'and'? If not, it is probably two abstractions. ## Smells of a wrong abstraction - Boolean or enum parameters that change control flow inside - Callers copying a magic combination of arguments from another call site - A shared helper whose test suite has cases nobody in the owning team recognises - Change requests phrased as 'add a flag to the shared method'

  • You inherit a shared function with five boolean parameters. What is your refactoring plan?
    Re-inline it: copy the body into each call site, then delete the branches that call site never takes, simplifying each copy. Once the true behaviour of each caller is visible, look for the smaller genuine commonality and extract only that. Tests are moved to the call sites first so behaviour is pinned during the split.
  • Does the rule of three mean you should never abstract on the second occurrence?
    No — it is a default, not a law. Abstract early when the knowledge is unambiguous and authoritative (a tax rule, a wire format, a security check), where divergence would be a defect rather than a variation. Wait when the similarity is structural and the future variation is unknown.
  • Why is duplication across service or team boundaries often preferred?
    Because a shared library couples release schedules, upgrade timing and design decisions across teams. A copied 40-line helper costs one occasional re-edit; a shared abstraction can turn every change into a cross-team negotiation. Deduplicate inside a boundary, tolerate it across boundaries unless the knowledge is truly authoritative (e.g. a wire contract, which is better shared as a schema).

Two neighbours' identical front doors do not mean they should share one doorway. Merge them and every time one wants a new lock, both must agree.

saying these in an interview costs you the question

  • Treating DRY as 'no two identical lines anywhere'
  • Extracting a shared abstraction from the very first similarity
  • Adding boolean parameters to keep a shared function serving diverging callers
  • Believing duplication is always technical debt regardless of context
  • Aggressively deduplicating tests until failures no longer explain themselves

context