How do you decide, in practice, whether a decision is "architecturally significant" and worth recording — and how do you record it?
answer
- One-way vs two-way door (Bezos Type 1 / Type 2)
- State and published contracts are the sticky parts
- Nygard ADR: Context / Decision / Consequences / Status
- Immutable — supersede, never edit
- Rationale outlives the decision: do the forces still hold?
basics
~20 sAsk: is it hard to undo, does it cross team or module boundaries, does it affect qualities like performance or security, and does it constrain future choices? If yes to any, it's significant — write it down as a short Architecture Decision Record: context, decision, consequences.
solid answer
~50 sI use four tests. (1) **Reversibility** — is this a one-way door (data migration, public contract, framework lock-in) or a two-way door I can walk back next sprint? (2) **Blast radius** — does undoing it need coordination across teams or a coordinated deploy? (3) **Quality-attribute impact** — does it move latency, availability, security, cost, or modifiability? (4) **Constraint on the future** — does it foreclose options for other decisions? Anything hitting these gets an **Architecture Decision Record (ADR)**: a one-page, immutable, numbered document with Title, Status (proposed/accepted/deprecated/superseded), Context (forces, constraints, what we knew), Decision (in active voice: "we will…"), Consequences (both good and bad), and Alternatives considered with why they were rejected. ADRs live in the repo next to the code, are reviewed like code, and are never edited after acceptance — you supersede them with a new one. The value is not the decision, it's the *rationale*: eighteen months later the team needs to know what forces applied, so they can tell whether those forces still hold.
code
markdown · 30 lines# 0007. Publish domain events via a transactional outbox
## Status
Accepted (2026-03-11). Supersedes 0004.
## Context
Services publish events after committing to their own database. Under broker
outage we lost ~0.3% of events (measured over 30 days), causing silent data
divergence downstream. We need at-least-once delivery. XA/2PC is unavailable
on our managed broker. Consumers are already idempotent.
## Decision
We will write events to an `outbox` table inside the same local transaction as
the state change, and relay them to the broker with a separate poller.
## Consequences
+ No lost events; delivery guarantee no longer depends on broker uptime at
write time.
+ Publish path becomes testable without a broker.
- Adds end-to-end latency (poll interval, currently 200ms p50).
- Duplicates are now normal, not exceptional: every consumer MUST be idempotent.
- One more table and one more background process to operate per service.
## Alternatives considered
- Dual write (publish then commit): rejected, loses events on crash between steps.
- Change-data-capture from the DB log: rejected for now, adds an operational
component the team has never run. Revisit if outbox poller lag exceeds 2s.
## Revisit if
Write volume exceeds 5k/s per service, or the platform team ships managed CDC.go deeper
Name the simple test — 'is it hard to undo and does it affect more than my module?' — and show you know ADRs exist and what three sections they contain (context, decision, consequences).
Give the full four-part test, explain one-way vs two-way doors, and describe the ADR lifecycle including superseding rather than editing. Cite a real decision you recorded.
Discuss what actually makes things irreversible (persisted state, published contracts, frameworks that invert control, org structure), and how you keep the process lightweight enough that teams use it voluntarily.
Talk about decision-making at organizational scale: distributing authority (advice process / architecture advisory forum), pairing ADRs with automated fitness functions so intent is enforced, recording expiry conditions, and deliberately converting one-way doors into two-way doors so fewer decisions need the heavy process at all.
## Why a test is needed at all If everything is architecture, nothing is: you drown in ceremony. If nothing is, you wake up with a system whose critical constraints nobody chose and nobody can explain. So you need a cheap, repeatable filter. ## The four-part significance test **1. Reversibility — one-way vs. two-way doors.** Amazon's Jeff Bezos popularized the framing in his 1997/2015 shareholder letters: *Type 1* decisions are near-irreversible one-way doors and deserve slow, careful, consultative process; *Type 2* decisions are two-way doors — walk through, and if you don't like it, walk back — and should be made fast by small teams or individuals. The classic failure is applying Type 1 process to Type 2 decisions, which produces "slowness, unthoughtful risk aversion" — and the mirror failure of treating an irreversible choice casually. What makes something irreversible in software, concretely: - **State.** Code is rewritable; data is not. Any decision that shapes persisted data (schema, partition key, identifier scheme, event-log format) is sticky because terabytes must be migrated live. - **Published contracts.** Once external clients depend on an API shape, a message schema, or a URL structure, you own it roughly forever, or you own a deprecation programme. - **Inversion of control.** A *library* you call can be swapped. A *framework* that calls you (dependency injection container, web framework, ORM, actor runtime, orchestration platform) grows into your code's shape and is far harder to remove. - **Organizational commitment.** Team boundaries, on-call structures, vendor contracts, and hiring profiles all calcify around technical choices — Conway's law working in reverse. **2. Blast radius / coordination cost.** If undoing the decision requires a lock-step deploy of several services, a coordinated release train, or the agreement of three teams, it is architectural regardless of how small the code diff is. **3. Quality-attribute impact.** Quality attributes ("-ilities": performance, scalability, availability, security, modifiability, testability, observability, cost of operation) are what architecture exists to deliver. A decision that measurably moves one of them — especially one you have committed to in an SLA — is significant. Note that these attributes *conflict*: strong consistency trades against availability under partition (the practical reading of the CAP theorem); more caching trades freshness for latency; more decoupling trades runtime simplicity for local autonomy. Architecture is largely the record of which side of each trade you chose and why. **4. Constraint on future decisions.** Some choices are significant not because they are hard to undo but because they *foreclose* things. Choosing eventual consistency forecloses read-your-writes without extra work. Choosing a single shared identity provider forecloses per-tenant isolation models. A useful shorthand for the whole test (from Michael Nygard, who introduced ADRs in 2011): record the decisions that are **"architecturally significant" — those that affect the structure, non-functional characteristics, dependencies, interfaces, or construction techniques.** ## Architecture Decision Records (ADRs) **What they are.** Short markdown files, numbered sequentially (`0007-use-outbox-pattern-for-event-publishing.md`), stored in the repository (commonly `docs/adr/` or `doc/architecture/decisions/`), reviewed via the same pull-request flow as code. **Nygard's canonical sections:** - **Title** — a short noun phrase, numbered. - **Status** — *proposed → accepted → deprecated → superseded by ADR-0012*. Statuses form a chain of history. - **Context** — the forces at play: constraints, requirements, what was true at the time, what was uncertain. This is the section future readers actually need. - **Decision** — stated in active voice: "We will publish domain events via a transactional outbox table." - **Consequences** — everything that becomes easier *and* harder afterwards, positive, negative, and neutral. A record with only benefits listed is a sales pitch, not an ADR. Many teams add **Alternatives considered** (with rejection rationale) and **Assumptions / expiry conditions** ("revisit if write volume exceeds 5k/s"). **The three rules that make ADRs work:** 1. **Immutable.** Never rewrite an accepted ADR. If you change your mind, write a new one and mark the old one *superseded*. The value is the historical chain, not a tidy current-state document. 2. **In the repo, not the wiki.** They version with the code, they're reviewable, and they don't rot in a portal nobody opens. 3. **Short.** One page. If it takes five pages, the decision probably wasn't made yet. **Why the rationale matters more than the decision.** The commonest failure mode in long-lived systems is that a constraint outlives its reason. A team keeps a painful rule for years because "that's how it is", when the original force (a vendor limit, a compliance rule, a traffic profile) disappeared long ago. ADRs let you ask the only useful question: *do the forces recorded in Context still hold?* ## Related lightweight practices - **RFC / design-doc process** — a longer pre-decision document circulated for comment; the ADR is the durable outcome. - **Architecture Advisory Forum** (from Andrew Harmel-Law's "decentralised architecture" model): anyone may take a decision, but must first *seek advice* from those affected and from people with relevant experience; the advice must be considered, not necessarily obeyed. This scales decision-making without a bottleneck architect and without anarchy. - **Fitness functions** — automated checks that a decision is still being honoured (see continuous architecting). An ADR states intent; a fitness function enforces it. ## Edge cases and judgement calls - **Small decisions that are irreversible anyway.** A field name in a public webhook payload is three characters of code and a permanent contract. Significance is not proportional to effort. - **Big decisions that are actually reversible.** Choosing a cloud region, a CI provider, or a metrics backend often looks huge but is genuinely swappable if you kept an interface. Don't over-ceremony these. - **Non-decisions.** Deciding to *defer* a decision is itself worth recording — with the trigger that forces it ("revisit when we onboard a second tenant"). - **Decisions made by omission.** If nobody chose, the first commit chose. Spotting these retroactively and back-filling ADRs for load-bearing accidental choices is high-value archaeology.
- Why must an ADR never be edited after it's accepted?Because its value is the historical record of *why* — the forces that applied at that moment. Editing it retroactively rewrites history and destroys the ability to ask 'do those forces still hold?'. Instead you write a new ADR that supersedes it, leaving a readable chain of how the system's thinking evolved.
- What's the failure mode of an ADR process, and how do you avoid it?Two: ceremony creep (every trivial choice needs a document, so people route around the process) and write-only archives (nobody reads them). Fix the first with an explicit significance test and a hard one-page limit; fix the second by keeping ADRs in the repo, linking them from code and from PR templates, and pairing important ones with automated fitness functions so the decision is enforced, not just described.
- How do you record a decision you deliberately haven't made yet?Write it as an ADR with status 'proposed' or as an explicit deferral: state the options, state why deferring is cheap right now (the door is still two-way), and record the trigger that will force the choice — a traffic threshold, a second customer, a compliance date. This makes 'last responsible moment' an actual practice rather than an excuse.
A two-way door is a restaurant you can walk out of; a one-way door is a tattoo. Both take about the same amount of time to enter. The skill is telling them apart before you sit down — and noticing that a lot of things that feel like tattoos (your CI provider) are really restaurants, while things that feel trivial (a field name in a public API) are really tattoos.
saying these in an interview costs you the question
- Judging significance by effort or code size instead of reversibility and blast radius.
- Writing ADRs that list only benefits — a decision with no negative consequences means the trade-off wasn't analysed.
- Editing accepted ADRs in place to reflect current reality, destroying the rationale trail.
- Keeping decision records in a wiki or slide deck disconnected from the code they govern.
- Treating every technology choice as a Type 1 decision, creating a review bottleneck that teams route around.
- Recording the decision but not the alternatives rejected — leaving the next team to re-litigate them from scratch.