skip to content

What does the DRY principle ("Don't Repeat Yourself") actually require, and why is it stated in terms of knowledge rather than lines of code?

level: juniorimportance: must knowfreq 82%

answer

  1. knowledge, not lines
  2. single unambiguous authoritative representation
  3. Pragmatic Programmer, Hunt & Thomas
  4. changes together, always, same reason?
  5. coincidental vs real duplication

basics

~20 s

DRY says every piece of knowledge — a rule, a formula, a fact about the system — should live in exactly one place. So when that rule changes, you edit one spot instead of hunting for copies you might miss.

solid answer

~50 s

DRY comes from The Pragmatic Programmer: "Every piece of knowledge must have a single, unambiguous, authoritative representation within a system." The unit is knowledge — a business rule, a tax formula, a validation constraint, a wire format — not text. The failure DRY prevents is divergent change: the same rule encoded in five places, someone updates four, and the system now contradicts itself silently. That is why the definition avoids "lines of code": two identical-looking blocks that encode two independent rules are not a DRY violation, and two textually different blocks that encode the same rule are one. Practically, you apply DRY by asking "if this rule changed tomorrow, how many files would I have to touch, and would I find them all?" If the answer is more than one, the knowledge is duplicated. Deduplicating is usually done by naming the rule — a function, constant, type, schema, or generated artifact — and having every user call that name.

code

pseudocode · 15 lines
pseudocode
// Duplicated KNOWLEDGE: free-shipping threshold in 3 places
checkout():   if total > 100 { shipping = 0 }
banner():     show("Free shipping over $100")
sql:          WHERE order_total > 100  -- promo report
// -> threshold changes to 75, someone updates 2 of 3 -> silent inconsistency

// Single authoritative representation, everything derives from it
FREE_SHIPPING_THRESHOLD = 100
checkout():   if total > FREE_SHIPPING_THRESHOLD { shipping = 0 }
banner():     show("Free shipping over $" + FREE_SHIPPING_THRESHOLD)
report:       parameterised query bound to the same constant

// NOT a violation: same text, different knowledge
usernameLen(s): 3..32   // identity policy
nicknameLen(s): 3..32   // display policy — free to diverge tomorrow

go deeper

for a junior

State the definition in terms of knowledge, give the one-place-to-change benefit, and give a concrete duplication example such as a magic number repeated in several files.

for a middle

Distinguish real from coincidental duplication, name the "changes together, always, for the same reason" test, and list removal techniques beyond extracting a function (constants, schema/codegen, tests that assert agreement).

for a senior

Frame DRY as a trade between divergent change and coupling; cite rule of three / AHA / "duplication is cheaper than the wrong abstraction"; discuss hidden duplication across DB, API contract, client, and docs.

for a principal

Talk about where the authoritative source should live organisationally: contract-first schemas with generated artifacts, ownership boundaries, deliberate duplication across services to preserve deploy autonomy, and enforcing single-source via CI checks rather than convention.

## The definition DRY was coined by Andy Hunt and Dave Thomas in *The Pragmatic Programmer* (1999): > **Every piece of knowledge must have a single, unambiguous, authoritative representation within a system.** Unpack each word: - **Knowledge** — a fact or rule about the problem or the system. Examples: "orders over 100 ship free", "a username is 3–32 characters", "the retry budget is 3 attempts", "the date on the wire is ISO-8601 UTC". Knowledge is *semantic*; it is not the same as text. - **Single** — exactly one place owns it. - **Unambiguous** — you can point at that place and say "this is the rule"; there is no second candidate that might also be the rule. - **Authoritative** — everything else *derives* from it (calls it, imports it, is generated from it) rather than restating it. - **Within a system** — the scope matters. Two separately deployed systems that each know a rule are a different (and often acceptable) situation from two functions in one module that each know it; see the trade-off section. ## Why "knowledge", not "lines" The common misreading is "never write the same characters twice". That leads to two symmetric mistakes. **Mistake 1 — false positives (coincidental duplication).** Two snippets look identical today but encode *different* rules that happen to agree right now: ``` validateUsername(s): return 3 <= len(s) <= 32 validateNickname(s): return 3 <= len(s) <= 32 ``` These are two independent policies. Merging them into one shared `validateShortName` creates *coupling*: the day product says nicknames may be 64 chars, you either break usernames or add a flag parameter, and the "shared" function starts sprouting `if kind == ...` branches. That is a wrong abstraction — the change that DRY was supposed to make cheap has become expensive instead. **Mistake 2 — false negatives (hidden duplication).** The same rule is expressed in places that share no text at all: - a database `CHECK` constraint, a server-side validator, a client-side form regex, and a paragraph in the API docs all restating "3–32 characters"; - an `enum` in code and a lookup table in the database listing the same statuses; - a struct definition and a hand-written JSON parser that must agree field by field; - a build script and a Dockerfile that each hardcode the port. No copy-paste occurred, yet a single change to the rule requires coordinated edits in four repositories. This is the expensive kind of duplication, and grep will not find it. ## The operational test A practical way to decide, without philosophy: > **If this rule changes, must all these places change *together*, *always*, *for the same reason*?** - **Yes, always, same reason** → one piece of knowledge, duplicated. Fix it. - **They might change independently, for different reasons** → two pieces of knowledge that currently coincide. Leave them alone. This is the same test as the Single Responsibility Principle's "one reason to change", applied to duplication. Robert Martin's framing helps: duplication between code owned by *different actors* (different stakeholders who request changes) is usually **false** duplication and should be left duplicated. ## How you remove real duplication Extracting a function is only the most obvious tool. The full menu: | Kind of knowledge | Single-source technique | |---|---| | Behaviour / algorithm | function, method, module | | Value / magic number | named constant, config entry | | Shape of data | one type/schema; generate DTOs, validators, docs, and client SDKs from it (e.g. one interface-definition file) | | Structural repetition across many types | generics/parameterisation, code generation, macros | | Same rule in DB + app | constraint owned in one place, the other derived or tested against it | | Knowledge in prose | generate docs from the code, or test the docs' examples so drift fails the build | Note the pattern: when you genuinely cannot collapse two representations (a database constraint and a client-side check must both physically exist for performance and UX reasons), DRY is preserved by **derivation** — generate one from the other — or, failing that, by an **automated consistency check** that fails loudly when they drift. Duplication you cannot remove should at least be duplication you cannot *silently* break. ## Costs and limits Deduplication is not free; it converts duplication into **coupling**. Every caller of the shared thing is now affected by changes to it. So the cost model is: - **Cost of duplication**: divergent change (a bug fixed in 3 of 4 copies), inconsistent behaviour, more code to read. - **Cost of the wrong abstraction**: a shared component that must serve incompatible needs; parameter creep, conditional branches, ripple failures across unrelated features, and painful un-merging later. Sandi Metz's much-quoted rule of thumb — *"duplication is far cheaper than the wrong abstraction"* — is a reaction to teams that pay cost #2 to avoid cost #1. Related counterweights are the **rule of three** (wait for a third occurrence before extracting) and **AHA** ("Avoid Hasty Abstractions", Kent C. Dodds), which say: prefer to *wait* until the shape of the variation is known, because inlining a premature abstraction is harder than extracting a late one. Also note deliberate DRY exceptions: - **Tests** tolerate more repetition than production code; an explicit, obvious test that restates values is often better than a clever shared fixture, because tests are specifications you read one at a time. - **Across service or team boundaries**, a shared library that encodes business rules couples independent deploy cycles; many architectures deliberately duplicate small amounts of logic to keep autonomy ("prefer duplication over the wrong coupling" is a standard microservice guideline). ## Summary DRY is about *one authoritative source per fact*, enforced ideally by structure (call/derive/generate) and at minimum by an automated check. It is not a mandate to compress text, and it is not unconditional: it trades duplication risk against coupling risk, and a good engineer names which risk they are choosing and why.

  • Give an example of two identical code blocks that should NOT be merged.
    Two validators that both currently allow 3–32 characters but belong to different policies (account identity vs. display nickname). They change for different reasons and different stakeholders; merging them creates a shared function that will need a mode flag the first time one policy moves.
  • If a rule must physically exist in two places — say a database CHECK constraint and a client-side form check — how do you keep DRY?
    Make one authoritative and derive the other (generate the client validator from the schema/IDL), or, if derivation is impractical, add an automated test that asserts the two agree so drift fails the build instead of failing in production.
  • How does DRY relate to coupling?
    Removing duplication creates coupling: every caller now depends on the shared element. DRY is therefore a trade — divergent-change risk against ripple-change risk — and is applied per piece of knowledge, not globally.

Think of a company's official holiday calendar. If every team keeps its own copy in a spreadsheet, one team will still be working on a day the company is closed. DRY is publishing one calendar that everyone subscribes to — not banning the word "holiday" from appearing twice.

context