How do you decide whether something is "architecturally significant" enough to warrant an Architecture Decision Record, and how do you stop the decision log from going stale?
answer
- Hard to change + wide blast radius + quality trade-off
- Deliberate deviations are top-value records
- Tens of records, not hundreds and not three
- ADR reviewed in the same PR as the code
- Confirmation → architecture test / fitness function
basics
~20 sWrite a record when the choice is costly to reverse, affects more than one team or module, or trades off a quality attribute like performance or security. Keep the log alive by writing records at decision time, reviewing them in the same pull request as the code, and superseding rather than ignoring outdated ones.
solid answer
~60 sThe practical filter is Martin Fowler's framing — architecture is "the decisions that are hard to change" — combined with three concrete tests: (1) **cost of reversal** — would undoing this take weeks, a migration, or coordination across teams? (2) **blast radius** — does it constrain code outside the team that made it, or cross a module/service boundary or a public contract? (3) **quality-attribute trade-off** — does it deliberately trade availability, latency, security, cost or operability against something else? Any "yes" earns a record. Cross-cutting conventions (error format, logging, auth scheme), technology adoption or removal, and *deliberate* deviations from an existing standard also qualify. Excluded: internal implementation details, routine dependency bumps, and anything a single reviewer can reverse in an afternoon. Staleness is prevented by process, not intent: write at decision time; review the ADR in the same pull request as its implementation; give the log a generated index; sweep long-lived Proposed drafts; supersede rather than edit; and — most effectively — encode accepted decisions as automated architecture tests so drift fails the build instead of silently accumulating.
go deeper
Say: record it if it would be expensive to undo or if it affects other people's code; write it when the decision is made, not later.
Give the three tests (cost of reversal, blast radius, quality-attribute trade-off), list qualifying categories, and mention reviewing the record in the same pull request as the code.
Add calibration (onboarding test, re-litigation test, tens-not-hundreds), the value of recording deliberate deviations, the generated index, and periodic sweeps of stale drafts.
Treat the log as governance: define the significance bar and acceptance rule organisation-wide, wire accepted decisions to fitness functions and architecture tests so drift fails CI, and use citation habits (PRs, onboarding, code comments) to keep the log load-bearing rather than decorative.
## Part 1 — What deserves a record ### The underlying definition There is no crisp boundary between "architecture" and "design", and chasing one wastes time. Two working definitions do most of the practical work: - **Martin Fowler / Ralph Johnson:** architecture is "the decisions that are hard to change" — the shared understanding that experienced developers have of a system's design, especially the parts they wish they could change later but cannot cheaply. - **The SEI/ISO framing:** an **architecturally significant requirement (ASR)** is one with measurable, broad effect on the system's structure and quality attributes. Both point at the same operational test: *significance is proportional to the cost of being wrong.* ### Three tests, any of which qualifies **1. Cost of reversal.** If reversing this in six months would mean a data migration, a coordinated multi-team release, a rewrite of a module, or customer-visible breakage — record it. If a single engineer can reverse it in an afternoon inside one file, don't. This is the single strongest signal. **2. Blast radius / constraint on others.** Does the choice constrain code that other people write? Module dependency rules, the shape of the HTTP response envelope, the error-code taxonomy, the authentication mechanism, the event schema, the ID strategy — these bind everyone downstream. A choice confined inside one module's private implementation usually does not. **3. Quality-attribute trade-off.** Does the decision deliberately trade one quality attribute against another? Strong consistency versus availability; synchronous call versus event for decoupling at the cost of eventual consistency; caching for latency at the cost of staleness; a managed service for operability at the cost of lock-in. Any conscious trade of a quality attribute is exactly what a future reader will need explained. ### Categories that reliably qualify - **Structural**: service/module boundaries, layering rules, allowed dependencies, sync-call versus domain-event communication. - **Technology adoption or removal**: a database, broker, framework, language, or build system; also *removing* one. - **Cross-cutting contracts**: API response envelope, error taxonomy, pagination scheme, auth/session mechanism, logging and correlation-ID conventions, tenancy model. - **Data**: storage engine per bounded context, schema-migration policy, retention and PII handling, ID generation. - **Operational posture**: deployment topology, multi-region, rollback strategy, SLO targets that constrain design. - **Deliberate deviations**: "this service will *not* follow the standard X, because Y" — arguably the highest-value records of all, because they look like mistakes to anyone reading later. - **Load-bearing non-technical constraints**: licence choice, vendor commitment, a compliance obligation that forces a structure. ### Categories that usually do not - Naming conventions, formatting, code style (these belong in linters and style guides). - Routine dependency version bumps (unless a major migration with real consequences). - Implementation details entirely inside one module with a stable public surface. - Ticket-level task decisions and sequencing. - Anything already codified elsewhere — link to the standard rather than duplicating it. ### Calibration heuristics - **The onboarding test:** would a competent new engineer be *surprised* by this and ask "why?" If yes, record it. - **The re-litigation test:** is this the kind of thing someone will propose changing every six months? Record it once, cite it forever. - **Volume sanity check:** a healthy log for a mature system typically holds tens of records, not hundreds and not three. Hundreds means the bar is too low and nobody reads it; three means real decisions are being made off the record. - **When genuinely unsure, write it.** A short record costs half an hour; a lost rationale costs weeks. But do not lower the bar so far that scanning becomes impossible — that failure mode is just as fatal. ## Part 2 — Keeping the log alive A decision log dies in one of two ways: it stops being written, or it stops being true. Both are process problems. ### Write at decision time Retrofitted records are rationalisation, not rationale, and everyone can tell. The trigger should be the moment of commitment, not the end of the quarter. A useful team rule: *if a discussion changed what we are going to build and it passes the significance tests, someone opens the ADR before the implementation PR.* ### Bind the record to the code change The most effective single practice: the ADR is reviewed in the **same pull request** (or an immediately preceding one) as the code implementing it. This makes the record part of the definition of done, gives reviewers something concrete to argue with, and prevents the wiki-drift problem — the record lives in the repository it governs, versioned with it. ### Make the log navigable - A generated **index** (number, title, date, status, successor) so readers can see what is currently in force without walking supersede chains. - **Bidirectional supersede links**, never in-place edits. - Optionally a published static site (log4brains and MADR-based generators do this) so non-committers can read it. ### Sweep it periodically A short, scheduled review — quarterly, or at the start of a major initiative — that asks only: are there **Proposed** records older than a month (decide or withdraw them)? Are there **Accepted** records whose context has clearly expired (supersede or reaffirm)? Are there decisions we made in the last quarter with no record (back-fill the important ones)? This is a 30-minute exercise, not an audit. ### Close the loop with enforcement The strongest anti-staleness mechanism is to make decisions *executable*. MADR's **Confirmation** section asks exactly this: how will we verify the decision is followed? Concretely — dependency rules checked by an architecture test (ArchUnit, dependency-cruiser, Spring Modulith verification), an API-shape contract test, a lint rule, a CI budget check on bundle size or latency. Neal Ford, Rebecca Parsons and Patrick Kua call these **fitness functions**: objective, automated measures of an architectural characteristic. Once a decision is enforced by a test, drift fails the build, and the record and the system cannot silently diverge. Records that resist automation ("we will prefer X where practical") are worth rewording until they can be checked, or accepting as advisory and marking them so. ### Make it referenced, not just written A log nobody cites decays. Cheap habits that keep it in circulation: link the relevant ADR from pull-request descriptions and design docs; cite it in code comments at the place where the constraint bites (`// ADR-0012: catalog writes go through CatalogApi only`); include "read ADRs 1–10" in onboarding; and when a proposal recurs, answer with the record's number rather than repeating the argument. Citation is what converts a folder of Markdown into working institutional memory.
- Is a decision to deliberately break an existing team standard worth its own record?Yes — those are among the highest-value records. A deviation with no recorded rationale reads to every future maintainer as a mistake, and someone will eventually "fix" it back and reintroduce whatever problem the deviation solved. The record should state the standard being deviated from, the specific reason, the scope of the exception, and ideally the condition under which the service would return to the standard.
- What is the strongest mechanism for stopping a decision log from drifting away from the real system?Making accepted decisions executable. Where a decision can be expressed as an automated check — module dependency rules verified by an architecture test, an API contract test, a lint rule, a CI budget on latency or bundle size — drift fails the build instead of accumulating silently. MADR's Confirmation section exists to capture exactly this, and Ford, Parsons and Kua's term for such automated measures is fitness functions. Decisions that cannot be automated should at least be reworded to be checkable, or explicitly marked advisory.
Deciding what deserves a record is like deciding what goes in a house's survey documents versus its decorating notes. Where the load-bearing walls are, why the extension was built on piles, why the drains run the long way round — that survives every owner. The colour of the hallway does not, because repainting is cheap.
saying these in an interview costs you the question
- Recording every technical choice, producing a log so large nobody reads it — as damaging as recording none
- Recording only successes, so failed experiments and deliberate deviations leave no trace
- Batch-writing records at the end of a quarter, which yields rationalisation rather than rationale
- Keeping the log in a wiki disconnected from the repository it governs, guaranteeing drift
- Chasing a precise definition of "architecture" instead of applying the cost-of-reversal test
- Treating an accepted record as enforcement — without a test or review gate, an accepted decision and the running system diverge quietly