skip to content

Your organisation cannot run a multi-day formal evaluation for every significant decision. How do you design a lightweight, continuous architecture evaluation process (RFCs, decision records, an architecture review board or advice process) that catches real risk without becoming a bottleneck?

level: seniorimportance: must knowfreq 50%

answer

  1. triage by reversibility + blast radius
  2. written RFC with ≥2 options and measures
  3. advice process, not permission
  4. ADR = immutable record, supersede don't edit
  5. automate repeat findings as fitness functions

basics

~20 s

Only review decisions that are expensive to reverse. Use a written proposal template (context, options, trade-offs, chosen option), require the author to consult affected experts, timebox feedback, record the outcome as a decision record, and automate recurring checks instead of re-reviewing them.

solid answer

~50 s

Design it around three ideas. **Triage by reversibility and blast radius**: cheap, reversible, team-local decisions need no review; irreversible, cross-team or compliance-relevant ones do. **Make the artefact a short written RFC**, not a meeting — problem, constraints, quality attribute scenarios it must satisfy, two or three options with trade-offs, the choice and its rationale, and how it can be reversed. Writing forces the analysis; asynchronous comments scale better than calendars. **Route by advice, not permission**: the author must seek feedback from affected teams and named domain experts (security, data, SRE) but keeps the decision, which removes the review-board bottleneck while preserving expertise. Then close the loop: accepted RFCs become immutable ADRs; recurring rules become automated fitness functions, lint and CI checks, or platform defaults so they are never re-discussed; and reviews get an SLA (e.g. comments within three days, silence = consent). Reserve a full ATAM for the rare irreversible bet.

go deeper

for a junior

Say significant decisions get written down with options and trade-offs, reviewed by affected people, and recorded as decision records.

for a middle

Add the triage rule (reversibility, blast radius), the RFC template with measurable scenarios, and ADR immutability/superseding.

for a senior

Design the whole flow: routing (advice process vs narrow ARB), SLAs and silence-is-consent, facilitation rules for the review hour, findings with owners, and automation of repeat findings.

for a principal

Treat it as an operating model: proportional governance, decentralised decisions with centralised guardrails, paved-road defaults, fitness functions and SLOs as continuous evaluation, health metrics on the process itself, and a heavyweight ATAM reserved for rare irreversible bets.

## The tension Heavyweight evaluation (a full ATAM) is high value but costs days of many people's time. Skip review entirely and you accumulate expensive, silent, irreversible mistakes. Review everything centrally and the board becomes a queue: throughput drops, teams route around it, and reviewers approve rubber-stamp-style because they lack context. A working process must be **proportional**. ## 1. Triage: what deserves review Classify the decision: - **Reversibility** — "two-way door" (swap a library next sprint) vs "one-way door" (choose the primary datastore, split the monolith, pick a cloud region topology, expose a public API). - **Blast radius** — inside one team's code vs crossing team, service, data or org boundaries. - **Regulatory/security surface** — personal data, payments, authentication, tenancy isolation. - **Cost commitment** — new vendor, new runtime, new operational burden. Only the one-way, cross-boundary, regulated or expensive decisions justify review. Publish this rule so nobody guesses. ## 2. The artefact: a short RFC A one-to-three-page written proposal beats a meeting because it forces the author to actually do the analysis, is reviewable asynchronously across time zones, and leaves a searchable record. Template: 1. **Context / problem** and what happens if we do nothing. 2. **Drivers**: the two or three quality attribute scenarios this must satisfy, with measures ("p99 < 300 ms at 2k rps", "a new tenant onboarded without code change"). 3. **Options considered** — at least two real ones, each with pros/cons; "do nothing" counts. 4. **Trade-offs and risks** — which attribute is being sacrificed for which, explicitly. 5. **Decision** and rationale. 6. **Reversal cost / exit plan** — what it takes to undo. 7. **Consultations** — who was asked, what they said. A single-option RFC is a red flag: no options means no trade-off analysis happened. ## 3. Routing: advice process vs review board **Architecture review board (ARB)**: a standing group approves proposals. Pros: consistency, one place for cross-cutting concerns. Cons: queue, context starvation, learned helplessness, and a culture where teams optimise for approval rather than outcomes. **Architecture advice process** (as popularised in decentralised-architecture practice): anyone may make a decision, but must first seek advice from (a) everyone materially affected and (b) people with expertise in that area. The decision stays with the person doing the work; the advice and the reasons for accepting or rejecting it are recorded. Pros: no bottleneck, expertise still reaches the decision, ownership stays with the team. Cons: needs a strong writing culture and real consequences for skipping advice. A practical middle ground: advice process by default; a small ARB (or a rotating architecture forum with embedded team representatives) only for the one-way-door class; regular open sessions where in-flight RFCs are discussed live for anyone who wants it. ## 4. Compress the ATAM steps into an hour Even a short review benefits from the ATAM skeleton: business drivers → the design → the two or three top-priority scenarios → how the design handles each → write down risks, sensitivity points and trade-off points. Facilitation rules that make the hour productive: pre-read the RFC (no live presentation), one facilitator keeping the author explaining rather than defending, a scribe capturing findings verbatim, no solutioning in the room, and every risk leaving with an owner and a date. ## 5. Close the loop — the part everyone skips - **ADRs**: accepted decisions become immutable, numbered architecture decision records (context, decision, status, consequences). Superseding is a new record, never an edit — the history is the value. - **Automate recurring findings**: anything you'd say twice becomes a check — dependency/layering rules enforced in CI (module boundary tests, dependency-cruiser-style rules), required security headers, SLO burn alerts, IaC policy checks. These are **fitness functions**: automated tests for architectural characteristics, run continuously. - **Platform defaults**: the cheapest review is the one made unnecessary because the paved road already does the right thing. - **Re-validate assumptions**: non-risks carry assumptions; put a periodic (say quarterly) sweep on the register. ## 6. Health metrics and failure modes Track: median time from RFC opened to decided; share of decisions that skipped required advice; share of findings closed; number of superseded ADRs (healthy — means decisions are revisited). Failure modes: the board as a gate people route around; reviews so late that findings are unaffordable; reviewers without domain context; "approved" with no record of why; findings recorded with no owner; and re-litigating settled ADRs in every new review instead of checking whether their assumptions still hold.

  • How do you keep an architecture review board from becoming a bottleneck?
    Narrow its scope to one-way-door and cross-boundary decisions, make everything else an advice-process decision owned by the team, run reviews asynchronously on written RFCs with an SLA (silence after N days = consent), rotate embedded team members through the board, and convert every repeated ruling into a platform default or automated check so it never queues again.
  • What is a fitness function in this context?
    An automated, objective test of an architectural characteristic that runs continuously — a CI test asserting module dependency rules, a performance budget failing the build, an SLO burn-rate alert, a policy check on infrastructure code. It converts a one-time review finding into a permanently enforced constraint.
  • How do ADRs and RFCs differ?
    An RFC is the proposal-and-debate artefact, mutable while under discussion; an ADR is the resulting record of the decision — immutable, numbered, with context, decision, status and consequences. Changing your mind creates a new ADR that supersedes the old one, preserving the reasoning trail.

Building permits: you don't file paperwork to hang a picture, you do to move a load-bearing wall. The trick is publishing which walls are load-bearing so nobody has to ask.

saying these in an interview costs you the question

  • Reviewing every decision regardless of reversibility or blast radius
  • A board that grants approval but records no rationale
  • RFCs listing a single option — no options means no trade-off analysis
  • Reviewing after the code is built, when findings are unaffordable
  • Findings with no owner, no date, and no follow-up
  • Editing an ADR in place instead of superseding it
  • Re-deciding settled questions manually instead of automating them into CI

context