skip to content

Explain the "rule of three" and AHA ("Avoid Hasty Abstractions"), and describe how you decide the moment at which extracting an abstraction pays off.

level: middleimportance: should knowfreq 56%

answer

  1. 1st: write it; 2nd: wince; 3rd: refactor
  2. 2 samples can't show the axis of variation
  3. AHA = optimise for change, not for sharing
  4. late extraction easy, early extraction sticky
  5. extract day one for security/money/protocol

basics

~20 s

The rule of three says wait until you see the same thing a third time before extracting it — two copies aren't enough to reveal what really varies. AHA adds: prefer some duplication over an abstraction you invented too early.

solid answer

~50 s

With two occurrences you cannot see the axis of variation: any abstraction you extract encodes a guess about what will change. The third occurrence acts as a natural experiment — it shows you which parts are stable and which vary, so the extracted unit is shaped by evidence rather than prediction. AHA (Kent C. Dodds) generalises this: "prefer duplication over the wrong abstraction" and "optimise for change first", i.e. write code that is easy to change rather than code that is already maximally shared. The decision is economic. Extracting costs a permanent coupling between callers plus indirection; leaving duplication costs future synchronised edits and divergence risk. Extract early when the copies are already provably one rule (a correctness or safety concern, a wire contract, a security check) where divergence is a defect. Wait when the similarity is structural, the domain is young, or the copies live in different bounded contexts. And keep the escape hatch: re-inline the moment flags start appearing.

go deeper

for a junior

State both heuristics plainly and give the reason: with only two copies you can't tell what will actually change.

for a middle

Add the cost asymmetry (late extraction easy, early extraction sticky), name the smells of a hasty abstraction, and list the extract-immediately exceptions.

for a senior

Frame it as an economic decision under uncertainty, discuss ownership and blast radius, and distinguish single-sourcing contracts/data (safe) from single-sourcing behaviour (risky).

for a principal

Discuss it as team policy: guidance on when shared libraries are allowed, contract-first generation as the preferred deduplication, and cultural permission to re-inline abstractions rather than accrete flags.

## The two heuristics **Rule of three** — popularised via Martin Fowler's *Refactoring* (attributed to Don Roberts): *the first time you do something you just do it; the second time you wince at the duplication but do it anyway; the third time you refactor.* The number three is not magic; it is the smallest count at which you have enough samples to distinguish the invariant part from the varying part. **AHA — "Avoid Hasty Abstractions"** (Kent C. Dodds, building on Sandi Metz's "prefer duplication over the wrong abstraction"): *optimise for change first*. Do not ask "is this code shared?", ask "is this code easy to change?". Some duplication keeps change local; a premature abstraction makes change global. Both push in the same direction: **delay**. And delay is cheap in one direction only, which is the core insight below. ## Why two samples are not enough An abstraction is a *claim about what varies*. With one example you have no claim. With two you have exactly one difference to look at, and any of these could be the real axis: ``` sendWelcomeEmail(user): template=W, to=user.email, retries=3 sendResetEmail(user): template=R, to=user.email, retries=3 ``` Is the varying axis the template? The recipient resolution? The retry policy? The channel (what if the third case is SMS)? You will pick one and hard-code the rest into the abstraction's shape. If the third case varies along a *different* axis than the one you parameterised, the abstraction cannot express it — so you bolt on a flag, and the decay described under "wrong abstraction" begins. With three samples, the intersection of what is constant is much better evidenced, and the union of what differs suggests the parameter list. Statistically crude, but it works because it replaces prediction with observation. ## The asymmetry that justifies waiting - **Late extraction is easy.** Three copies sitting side by side, with their differences visible, are straightforward to fold into one well-shaped unit. The information you need is all present. - **Early extraction is hard to undo.** Once N call sites depend on a shared unit, un-sharing means touching all N, re-testing all N, and possibly negotiating with other teams. And people rarely do it; instead they add a parameter, which is the cheap local move that makes the global problem worse. So the expected cost of waiting (a few extra synchronised edits, maybe a bug from a missed copy) is usually less than the expected cost of guessing wrong (a permanent coupling with flag creep). That asymmetry — not laziness — is the argument. ## When you should NOT wait The heuristics have important exceptions. Extract at the *first* duplication when: 1. **Divergence is a defect by definition.** Security checks (authorisation, escaping, signature verification), money/tax arithmetic, wire-format encoding, ID generation. Here two copies mean two chances to be wrong, and "they might legitimately diverge" is false. 2. **The knowledge is external and authoritative.** A protocol, a schema, a regulatory rule. You are not inventing an abstraction; you are naming a fact that already has one definition. 3. **You can already name it precisely.** If the shared thing has a real domain name (`TaxRate`, `RetryPolicy`, `SlugFormat`) rather than a vague one (`Helper`, `Utils`, `processData`), the concept exists independently of your code and extraction is safe. 4. **The copies are far apart or invisible to each other.** Two copies in the same file will be noticed; two copies in different services will silently diverge for a year. Distance raises the cost of duplication. Conversely, wait longer than three when: the copies live in different bounded contexts or teams; the domain is new and requirements are still moving; or unification would couple deployment cycles that must stay independent. ## A decision checklist Before extracting, ask: 1. Can I name the concept in domain language, without "and" and without "Util"? 2. Do all current call sites want *all* of the behaviour, with no mode flag? 3. If a requirement changed for one caller only, would the abstraction survive without a boolean? 4. Is divergence a bug, or a legitimate future? 5. Who owns the shared unit once it exists, and are all callers within that ownership boundary? Two or more "no"s → wait, or extract something smaller and more obviously true (a constant, a value type, a pure function) rather than a big shared workflow. ## Signals you extracted too early - A boolean/enum parameter whose only job is to pick a branch. - Callers passing `null`/defaults for parameters they do not care about. - The abstraction's tests enumerate combinations no caller uses. - A change request for one feature requires editing a file named after none of the features. - The name contains `Base`, `Abstract`, `Generic`, `Common`, `Manager` and nobody can explain the concept without listing its callers. **Remedy:** inline it back, then re-extract from evidence. Reversing early is far cheaper than living with it. ## Nuance: DRY is not the same as "extract everything" Some real duplication should be resolved by *derivation* rather than by an abstraction in code — generating clients from a schema, or a CI test that asserts two artifacts agree. Those carry none of the flag-creep risk, because they do not create a shared runtime dependency between callers. When in doubt, prefer single-sourcing *data and contracts* (very safe) over single-sourcing *behaviour and workflows* (where wrong abstractions live).

  • Name a case where waiting for three occurrences would be wrong.
    Anything where divergence is a defect rather than a possibility: an authorisation check, an output-escaping routine, currency rounding, or the encoder for a wire protocol. There is exactly one correct behaviour, so a second copy is already a second chance to be wrong.
  • You extracted after the third occurrence and the fourth caller doesn't fit. Now what?
    Do not add a flag. Either leave the fourth caller duplicated (cheapest, keeps the abstraction honest), or if it reveals a better axis of variation, reshape the abstraction — potentially by inlining and re-extracting along the axis all four now demonstrate.
  • Isn't the rule of three just an excuse to write sloppy code?
    No — it is a claim about information, not effort. Two samples underdetermine what varies, so an early abstraction encodes a guess; and the guess is expensive to reverse while duplication is cheap to fix. You still remove duplication, just when you can see its shape.

Paving a path across a lawn: pour concrete after one person walks across and you'll pave the wrong line. Wait until footprints appear from several directions and the desire path tells you where the path actually belongs. Concrete is easy to pour and hard to move — like a shared abstraction.

context