skip to content

What does the "Single Source of Truth" design principle mean, and what problem does it prevent?

level: juniorimportance: must knowfreq 58%

answer

  1. one owner per fact
  2. copies are derived, never hand-edited
  3. divergence is silent, not loud
  4. N copies = N edits forever
  5. regenerate + drift-check in CI

basics

~20 s

Every fact — a piece of data, a setting, a rule — lives in exactly one authoritative place. Everything else reads from that place or is generated from it, so copies can never quietly disagree.

solid answer

~50 s

Single Source of Truth (SSoT) says that for each piece of knowledge there is exactly one canonical, authoritative location that owns it. Other places may still hold the value, but only as *derived* copies produced automatically from the owner — a cache, a read replica, a denormalized column, a generated client, a rendered report. What it prevents is divergence: when two independently editable copies exist, sooner or later one is updated and the other is not, and the system now has two contradictory answers with no rule for which wins. Divergence is expensive because it is silent — nothing crashes, the data is just wrong, and reconciling it later is manual archaeology. In practice SSoT looks like: one config value (not the same timeout hardcoded in code and repeated in the deploy script), one owning table for a fact, one service that owns an entity, one schema from which DTOs and docs are generated.

go deeper

for a junior

Define it in one sentence — one authoritative place per fact — and give a concrete example such as a timeout value duplicated in code and in a deploy script drifting apart.

for a middle

Add the derived-copy distinction: caches, replicas and generated code are fine when generated from the owner; hand-maintained duplicates are the problem. Mention enforcement (codegen, drift checks).

for a senior

Frame it as managing divergence risk, name owners explicitly, and discuss trade-offs: reading from the owner (coupling, latency) vs deriving copies (staleness, invalidation, reconciliation).

for a principal

Talk about ownership as an organizational contract, per-field rather than per-entity ownership, staleness budgets as an explicit SLO, and when centralizing truth becomes a bottleneck worth violating deliberately.

## The terms - **Fact / piece of knowledge**: one thing the system must know — a user's email, a retry timeout, a tax rate, the set of valid order states, the shape of an API request. - **Canonical (authoritative) location**: the place designated *correct by definition*. If it disagrees with anything else, it wins and the other thing is the defect. - **Derived copy**: a value that exists elsewhere but is *produced from* the canonical one by a mechanism (replication, caching, code generation, an export job). Derived copies are fine; hand-maintained duplicates are not. - **Divergence (drift)**: two copies of the same fact holding different values. ## Why divergence is the real enemy Duplication by itself is harmless if nothing ever changes. Systems change. The cost curve looks like this: with one copy, a change is one edit; with N independently editable copies, a change is N edits that must all happen, atomically enough, forever, by every future maintainer — including the one who does not know copy #4 exists. The probability that all N stay in step trends to zero over time. The failure is *silent*. A crash is cheap: you see it, you fix it. Divergent data produces plausible-looking wrong answers: the invoice says one price, the order page another; the app times out at 5s while the load balancer gives up at 3s; the API docs describe a field the server no longer sends. Nobody notices until a customer does. ## What it looks like concretely | Layer | Canonical thing | Derived copies | |---|---|---| | Data | The owning table/row for a fact | caches, read replicas, search index, denormalized reporting tables | | Config | One config source per environment | env vars injected at deploy, rendered templates | | Contract | One schema (OpenAPI/IDL/JSON Schema) | server stubs, client SDKs, docs, validation code | | Domain | One service that owns an entity | events, local read models in other services | | Docs | The code / the schema | generated reference docs, diagrams | ## What SSoT is *not* - **Not "never store a value twice"**. Caches, replicas and denormalized read models exist for latency and availability and are entirely compatible with SSoT — *provided* they are generated from the owner, are labelled derived, and are never written to directly by a human or by application code. - **Not "one giant central database for the whole company"**. SSoT is per *fact*, not per *system*. Ten services can each be the single source of truth for their own facts. - **Not the same as DRY**, though they are cousins: DRY is about knowledge in *code*, SSoT generalises it to data, config, contracts and documentation. ## How you enforce it 1. **Name the owner** for each fact and write it down (ownership table, ADR, `CODEOWNERS`, service catalogue). 2. **Make copies read-only** — the only writer is the derivation mechanism. 3. **Regenerate rather than edit** — generated files carry a "do not edit" header and are rebuilt in CI. 4. **Detect drift** — CI regenerates and fails if the result differs; reconciliation jobs compare derived stores against the owner and report mismatches. 5. **Prefer reading over copying** at low volume; only introduce a derived copy when there is a measured reason (latency, availability, query shape). ## Edge cases - **Two candidate owners** (e.g. billing and CRM both hold an address): pick one owner per *field*, not per entity, if necessary; the loser reads or subscribes. - **The owner is unavailable**: derived copies buy availability, at the price of serving stale data — an explicit, bounded trade (staleness budget/TTL), not an accident. - **Human-entered data in two UIs**: this is the classic generator of divergence; fix it in the workflow, not by adding a nightly reconciliation spreadsheet.

  • Is a cache a violation of Single Source of Truth?
    No, as long as the cache is derived from the owner, is never written directly by application logic, and has a defined invalidation/expiry policy. It is a copy with a known staleness budget, not a second authority.
  • Give an example where duplication is acceptable.
    Two facts that happen to have the same value today but change for independent reasons — e.g. an HTTP client timeout and a batch job timeout that both happen to be 30s. Merging them into one constant would be false coupling, not DRY.

A wall clock and your watch: if you set both by hand they slowly disagree and you never know which to trust. If your watch syncs from a time server, there is one truth and one derived display — you can have a hundred watches and they all still agree.

context