Beyond application code, where else does knowledge get duplicated in a system, and what techniques give you a single authoritative source for it?
answer
- duplication hides in schema/config/IaC/docs/SDKs
- derive > verify > observe
- contract-first + codegen kills drift structurally
- shared test vectors for cross-language rules
- choose the source by who owns and reviews the rule
basics
~20 sThe same rule often lives in the database schema, the API contract, client code, config files, infrastructure scripts, and the docs. You single-source it by generating the copies from one definition, or by adding automated checks that fail when they disagree.
solid answer
~50 sMost expensive duplication is not copy-pasted code — it is one fact restated across artifacts that share no text: a DB constraint, an ORM mapping, a validator, a client form, an API spec, an SDK, a README, and a runbook. Because nothing links them, drift is silent and shows up as production inconsistency. The single-source toolkit, in order of strength: (1) derive — pick one authoritative definition (an IDL/schema, a migration, a rules table) and generate everything else (DTOs, validators, clients, docs, dashboards, infrastructure); (2) verify — where generation is impractical, assert agreement in CI with contract tests, schema-vs-code checks, shared test-vector files, or executable documentation; (3) observe — detect drift in production (config diffing, shadow comparison). Choosing the authoritative artifact is an architectural decision with ownership consequences: it must be readable by the people who own the rule, versioned, and reviewable. And single-sourcing has limits — the generator becomes shared infrastructure, generated code must stay debuggable, and over-generalised templates recreate wrong-abstraction problems at build time.
code
yaml · 14 lines# One authoritative definition; everything downstream is derived, not restated.
User:
properties:
username: { type: string, minLength: 3, maxLength: 32 }
# generated / derived from the above:
# - server request validation
# - typed client SDKs (each language)
# - reference docs for the field
# - DB migration constraint (or a CI test asserting the DB matches)
# - fixtures used by contract tests
#
# If a copy cannot be generated (e.g. a legacy validator in another runtime),
# it runs a shared test-vector file so any disagreement fails the build.go deeper
Point out that the same rule often lives in the database, the code, and the docs, and that changing one and forgetting the others causes bugs.
Give concrete cross-artifact examples and the basic fixes: constants for config values, generating clients from an API spec, and a test that asserts a code enum matches the database's allowed values.
Present the derive/verify/observe hierarchy, use contract-first generation and consumer-driven contract tests as the main tools, and note that unremovable duplication must at least be made loud.
Treat the choice of authoritative artifact as an architecture and ownership decision — reviewability, deployment semantics, who owns the rule — and name the costs: generator ownership, debuggability of generated code, and not turning a build-time single source into a runtime single point of failure.
## The map of non-code duplication Ask "where else does this fact live?" and the list is usually longer than expected. | Fact | Typical duplicate homes | |---|---| | Field is required, max 32 chars | DB `NOT NULL`/`CHECK`, ORM annotation, service validator, client form, API spec, docs, test fixtures | | Set of valid statuses | code enum, DB lookup table or `CHECK`, UI dropdown, analytics dashboard filter, ETL job | | Timeout / retry budget | client config, server config, load-balancer setting, circuit breaker, runbook, SLO doc | | Port / hostname / queue name | app config, Dockerfile, compose file, IaC, CI workflow, monitoring config | | Wire format of an event | producer code, consumer code, schema registry, docs, replay tooling | | A business rule | service code, batch job, report SQL, spreadsheet used by finance, support macro | | Environment inventory | IaC, deploy scripts, on-call docs, dashboards | None of this is findable with a duplicate-code detector. The unifying property: **one change request implies edits in several artifacts, in several repos, owned by several roles** — and no tool tells you when one was missed. ## Technique tier 1 — derive (strongest) Pick one artifact as authoritative, mechanically produce the rest. - **Contract-first APIs**: an OpenAPI/gRPC/GraphQL schema generates server stubs, typed clients, request validation, and reference docs. Producer and consumer cannot disagree about the shape. - **Schema-first events**: an event schema in a registry, with compatibility rules, generates serializers for every language and enforces evolution rules at publish time. - **Types as the source**: define the model once and derive validators, serialization, database mapping, and (with modern toolchains) the client's types from the same declaration. - **Database as source, or migrations as source**: generate model classes from the schema, or the schema from migrations — pick one direction and never hand-edit the other. - **Configuration as data**: one typed config definition generates env-var docs, defaults, validation on boot, and IaC parameters. - **Docs from code**: reference documentation, CLI help, and error-code catalogues generated from annotated source; architecture diagrams generated from the module graph rather than drawn by hand. - **Infrastructure modules**: one parameterised IaC module instantiated per environment, instead of copied scripts that drift. Properties: drift becomes structurally impossible; the copies are artifacts, not sources. Costs: a build-time dependency, generated code in the repo or the build, and a generator that itself becomes shared infrastructure someone must own. ## Technique tier 2 — verify (when you cannot generate) Sometimes both copies must be hand-written (different runtimes, legacy systems, a vendor's format). Then make disagreement *fail loudly*: - **Consumer-driven contract tests**: each consumer publishes expectations; the provider's build verifies them. Prevents "the docs say X, the service does Y". - **Shared test vectors**: a language-neutral file of inputs and expected outputs (e.g. rounding cases, canonicalisation cases) that every implementation of the rule runs as tests. Cheap and extremely effective for algorithms duplicated across languages. - **Schema-vs-code assertions**: a test that reads live DB constraints and compares them to application-level rules; or a test asserting the enum in code matches the lookup table's rows. - **Executable documentation**: run the README's snippets in CI; verify example requests against the real API; snapshot-test the error-code table against the source of truth. - **Config linting**: assert that the port in the app config, the Dockerfile, and the deployment manifest agree. Property: the duplication remains, but it can no longer drift silently — which is most of the harm. ## Technique tier 3 — observe (last resort) - Compare production behaviour of two implementations (shadow traffic / dual-run with diffing) during a migration. - Alert on config divergence between environments. - Cross-reference comments naming the other copies — weakest, since nothing enforces it, but it at least makes the copies discoverable. ## Choosing the authoritative artifact This is the real architectural decision. Criteria: 1. **Who owns the rule?** If compliance owns tax rates, the source should be something they can review — a data table or a rules file, not a constant buried in code. 2. **Reviewability and diffability.** The source must be human-readable and version-controlled; a value in a runtime admin UI is a source of truth with no history and no review. 3. **Deployment semantics.** Code-as-source means changes need a deploy; data-as-source means changes are instant — powerful, and dangerous (no code review, no rollback story). Decide which the rule deserves. 4. **Toolability.** Can everything downstream actually be generated from it, in every language you support? 5. **Stability of the format.** Migrating the source of truth later is expensive; prefer standard formats over bespoke DSLs. ## Limits and failure modes of single-sourcing - **The generator becomes a dependency.** Broken codegen blocks every team; it needs an owner, tests, and a version policy. - **Generated code must be debuggable.** Stack traces through unreadable output, or generated code that people hand-edit "just this once", defeats the purpose. Generated artifacts should be reproducible and either committed with a check that they are up to date, or produced in the build. - **Over-generalised templates are wrong abstractions at build time.** A code generator with twenty switches serving incompatible consumers has all the pathologies of a flag-laden shared function. - **Not every restatement is duplication.** A summary in a design doc is a *different* representation for a different audience; forcing it to be generated makes it useless. Distinguish authoritative statements (must agree) from explanatory ones (may abstract). - **Single source ≠ single point of failure.** If the schema registry or the config service is on the critical path at runtime, you have converted a maintenance problem into an availability problem. Prefer build-time single-sourcing where possible. ## The principal-level framing > "I want each fact to have one owner, one reviewable representation, and a mechanical path to every place it is used. Where I can generate, I generate; where I can't, I add a test that fails when the copies disagree; and where I can't do either, I write down that the duplication exists and who is responsible for it. What I never accept is duplication that can drift silently."
- A validation rule must exist both in a browser form and on the server. How do you keep one source of truth?Make the server-side schema authoritative and generate the client validator from it, so the browser check is a derived artifact. If the runtimes make generation impractical, keep both hand-written but drive them from a shared test-vector file in CI, so any disagreement fails the build rather than reaching users.
- What is the downside of aggressive code generation as a DRY strategy?The generator becomes critical shared infrastructure needing an owner and a version policy; generated code can be hard to debug and tempting to hand-edit; and an over-parameterised generator serving incompatible consumers reproduces the wrong-abstraction problem at build time.
- Are documentation and comments a DRY concern?Yes, when they restate authoritative facts — a comment or doc describing behaviour is an unverified copy that drifts. The fix is to generate reference material from the source or to execute the documented examples in CI. Explanatory prose that gives rationale is a different representation for a different audience and should not be mechanised away.
A building's floor plan: the architect's model is authoritative, and the electrician's, plumber's, and fire-safety drawings are generated views of it. If each trade redraws the plan by hand, the sockets end up in a wall that no longer exists — and nobody notices until the wiring is in.