How does Single Source of Truth differ from the DRY (Don't Repeat Yourself) principle, and how does database normalization relate to both?
answer
- DRY = knowledge in code
- normalization = knowledge in rows
- SSoT = knowledge anywhere
- update/insert/delete anomalies
- denormalize as derived, not as second owner
basics
~20 sDRY is about not duplicating knowledge in code; SSoT generalises the same idea to data, config, schemas and docs. Database normalization is the data-modelling technique that implements SSoT: each fact stored in exactly one place.
solid answer
~60 sThey are the same idea applied at different layers. **DRY**, as originally stated, is "every piece of knowledge must have a single, unambiguous, authoritative representation within a system" — note it says *knowledge*, not *text*; two identical code fragments that change for different reasons are not a DRY violation. **SSoT** applies that rule beyond code: to runtime data, configuration, API contracts, infrastructure definitions and documentation. **Normalization** (1NF/2NF/3NF/BCNF) is the formal, mechanical way to achieve SSoT inside a relational schema: you remove *redundancy* by ensuring every non-key attribute depends on the key of its own table, so a customer's address is stored once and referenced by foreign key rather than copied into every order row. The payoff is identical in all three: an update is a single write, so there is no *update anomaly* — no state where half the copies are new and half are old. The cost is also identical: you must now join, look up, or fetch across a boundary to read the value, which is why denormalization, caching and duplication reappear deliberately for performance — as derived copies with an explicit refresh mechanism.
go deeper
Say DRY is about not repeating knowledge in code and normalization is the database version of the same idea; give the customer-email-in-every-order example.
Quote DRY's actual wording (knowledge, not text), name the update/insert/delete anomalies, and note that coincidental duplication should not be merged.
Discuss the read/write trade-off explicitly: normalization optimises writes and correctness, denormalization optimises reads; keep the owner canonical and make the denormalized form a maintained derivation.
Position all three as one governance question — who owns each fact, at what layer, with what refresh and reconciliation guarantees — and note where over-DRY creates organizational coupling between teams.
## Definitions first - **DRY (Don't Repeat Yourself)**, from *The Pragmatic Programmer*: "Every piece of knowledge must have a single, unambiguous, authoritative representation within a system." The subject is **knowledge**, not characters. Two functions with similar bodies that exist for unrelated reasons are *coincidental duplication*, and merging them is the anti-pattern usually called **false abstraction** or over-DRY. - **SSoT**: the same rule scoped to *anything the system knows* — rows, config keys, feature flags, schemas, infrastructure definitions, runbooks, diagrams. - **Normalization**: a set of relational design rules (normal forms) that eliminate redundancy by decomposing tables so each fact is stored once. ## How they nest ``` SSoT (any knowledge: data, config, contracts, docs) └── DRY (knowledge expressed in code) └── Normalization (knowledge expressed as relational rows) ``` DRY and normalization are *instances* of SSoT in their respective media. This is why the same three failure modes appear in all of them. ## The three anomalies (the concrete reason redundancy hurts) Classic relational theory names them, and they generalize to code and config verbatim: 1. **Update anomaly** — the fact is stored in many rows; an update touches some and misses others, leaving contradictory values. (Code version: the timeout changed in two of three call sites.) 2. **Insert anomaly** — you cannot record a fact because there is no row to hang it on. (Code version: a rule cannot be expressed until some unrelated object exists.) 3. **Delete anomaly** — deleting a row destroys an unrelated fact that only lived there. (Code version: removing a feature silently removes the only definition of a shared constant.) ## Worked example Denormalized: `orders(order_id, customer_name, customer_email, item, price)`. The customer's email is repeated per order. Changing it means an `UPDATE` over an unknown number of rows; miss any and the customer has two emails. Normalized: `customers(customer_id, name, email)` + `orders(order_id, customer_id, item, price)` — one row owns the email; orders reference it. One write, no anomaly. The symmetric code case: a discount rule implemented in the checkout service and again in the invoice renderer. Same fact, two representations, guaranteed to drift. The fix is one implementation both call — or one *definition* both are generated from. ## Where the analogy is imperfect - **DRY has a false-positive mode that normalization does not.** Normalization is driven by functional dependencies — formal and checkable. DRY is driven by human judgement about whether two things represent the *same* knowledge. Deduplicating code that merely looks alike produces coupled modules that must change together for no reason. The heuristic: duplicate is real if a change to the business rule must change both places; otherwise leave them apart. "Rule of three" and "prefer duplication over the wrong abstraction" are the practitioner shorthands. - **Normalization has a performance cost that DRY usually does not.** Fully normalized schemas need joins; at scale you deliberately denormalize, add materialized views, or maintain read models. That is *not* abandoning SSoT — it is keeping the normalized table as the owner and making the denormalized form a **derived copy** with a defined refresh path (trigger, materialized view refresh, CDC stream, cache invalidation). - **SSoT extends where DRY has nothing to say**: which *service* owns a customer record, which repository owns the infrastructure definition, whether the docs are generated from the schema. ## Interview-grade summary Same principle, three media: code (DRY), rows (normalization), everything (SSoT). All three trade *write simplicity and correctness* against *read cost*. When read cost wins, you reintroduce copies — but as generated, read-only derivations from a named owner, never as second authorities.
- Is denormalizing a table a violation of Single Source of Truth?Not if the normalized table stays the owner and the denormalized form is maintained automatically — trigger, materialized view, CDC projection — and is treated as read-only. It becomes a violation the moment application code writes to the denormalized copy directly.
- Give an example of code duplication that should NOT be removed.Two validation snippets that are textually identical today but belong to different rules — say a signup password check and an internal admin token check. They change for independent reasons; merging them creates coupling that forces one to change when the other's rule moves.
One recipe card in the kitchen (SSoT). Photocopying it for each cook is fine while the copies are reprinted from the original; writing corrections on individual copies is how three cooks end up baking three different cakes.