skip to content

You own architecture documentation standards across many teams. How do you decide what is mandatory, and how do you know the documentation is actually working?

level: principalimportance: nice to knowfreq 18%

answer

  1. Mandate only what costs outsiders if missing
  2. Standardise structure/location/notation, not depth
  3. Tier by criticality and lifespan
  4. Enforce via PR, CODEOWNERS, arch tests, scaffolding
  5. Measure citation, onboarding, incident use; then delete

basics

~20 s

Mandate only what has a named reader and a real consequence: system context, external contracts, decisions, and quality goals. Standardise structure and location, not depth. Measure by outcomes — onboarding time, questions asked, incident-time lookups — not by page counts or compliance ticks.

solid answer

~60 s

Start from decisions and readers, not from artefacts. The mandatory set should be the small slice with cross-team consequences: a **system context** (who calls us, what we call), the **externally visible contracts** (APIs, events, SLOs), an **ADR log** for boundary-affecting decisions, **quality goals** with measurable scenarios, and an **owner plus a last-verified date**. Everything else is per-team discretion. Standardise **structure and location** — one template (arc42 or a trimmed house variant), one place (in the repo), one notation set (C4 plus text-based diagrams) — so readers can move between systems without relearning, while depth stays proportional to system criticality and lifespan. Enforce through the paths that already exist: PR review, CODEOWNERS, architecture tests and fitness functions, plus a lightweight review at natural checkpoints rather than a documentation gate. Measure with **leading indicators** (time to first commit for a new joiner, repeat questions in chat, whether docs are opened during incidents, ADR coverage of significant changes, freshness age) and treat unread documentation as a defect. The two governance failures are the same failure — a mandate nobody reads, and a free-for-all nobody can navigate.

go deeper

for a junior

Say the standard should be small and consistent — same template, same place in the repo, an owner and a date — and that docs are useful only if people actually read them.

for a middle

Add the mandatory core (context, contracts, ADRs, quality goals), enforcement through PR review and templates, and the idea of scaling depth to how critical the system is.

for a senior

Argue from readers and decisions, name the two failure modes (mandate theatre and free-for-all), and give concrete enforcement (CODEOWNERS, architecture tests, fitness functions, scaffolding) plus outcome metrics.

for a principal

Present it as an operating model with an explicit trade-off curve: minimal core with teeth, tiering by criticality, governance fused with the ADR/RFC process, evidence-based measurement (onboarding, repeat questions, incident use, citation), and a deliberate deprecation practice.

## Frame the problem correctly At scale, architecture documentation is not a writing problem, it is a **product and incentive problem**. The document has users; if they don't use it, it has failed regardless of how complete it is. Two symmetrical failure modes dominate: 1. **Mandate theatre** — a heavyweight standard ("every service must have a full 12-section arc42 plus five views"), satisfied by copy-pasted templates that nobody reads and nobody updates. Compliance is 100%; value is near zero; and worse, the existence of a document implies coverage that isn't there. 2. **Free-for-all** — every team invents its own outline, home and notation. Individual documents may be good, but nobody can navigate across systems, and cross-team questions ("who consumes this event?") have no answer anywhere. The cure for both is the same shape: **standardise the cheap things (structure, location, notation, ownership), leave the expensive things (depth, extra views) to judgement.** ## Deciding what is mandatory Apply a single test: *does the absence of this artefact impose cost on someone outside the owning team?* That yields a small mandatory core: - **System context / scope** — the system as a black box with every external dependency and consumer. This is the artefact other teams need most and owning teams need least, so it will not exist unless mandated. - **Externally visible contracts** — APIs, events, data ownership, SLOs, deprecation policy. Ideally generated from schemas and contract tests. - **ADR log** for decisions that cross a boundary, create an external dependency, or trade quality attributes. Not for internal, reversible choices. - **Quality goals** — the top three to five, as measurable scenarios, not adjectives. "Available" is not a goal; "recover to read-only within 5 minutes of an AZ loss" is. - **Ownership and freshness metadata** — a named owner and a last-verified date on every hand-written page. Cheap, and it converts silent staleness into visible staleness. Everything else — runtime views, deployment detail, cross-cutting concepts, glossary depth — is *recommended*, scaled to the system's **criticality and expected lifespan**. A tiered model works well: Tier 1 (regulated / revenue-critical / decade-long) gets full views and periodic review; Tier 3 (internal tool, one team, replaceable) gets a README, a context diagram and an ADR log. ## Standardise structure and location, not volume - **One template** — pick arc42 (or a trimmed house variant with the same numbering) so section 3 means the same thing in every repo. Predictable slots are what make cross-system reading possible. - **One location** — in the code repository, published to a single searchable portal. Findability beats completeness: documentation that exists but cannot be located is indistinguishable from documentation that doesn't exist. A central index of systems, owners and links is often the single highest-value artefact you can create. - **One notation set** — e.g. C4 for static structure, text-based diagram sources (PlantUML/Mermaid/D2/Structurizr DSL) so everything diffs in review and can be rendered by CI. - **Explicit non-goals** — say plainly what you do *not* want (no 100-page documents, no diagrams as binary images, no duplicating a fact in two places). ## Enforcement that doesn't become a gate Use the paths engineers already walk: - **Pull-request review** with a short checklist: boundary changed → update context/contract or write an ADR. - **CODEOWNERS** on documentation directories so the right reviewer is summoned automatically. - **Automated checks**: architecture tests (layering, allowed dependencies), fitness functions (cycles, coupling, latency budgets, licence policy), link/anchor checkers, and a CI step asserting generated views match their source. - **Checkpoint reviews** at natural moments — a new system, a major re-platform, a boundary change — rather than a calendar-driven documentation audit. - **Templates and scaffolding**: a repo template that already contains the trimmed arc42 skeleton and an ADR folder. Defaults beat policies. A standing **architecture forum / RFC process** handles the cross-team decisions, and its output *is* the ADR — governance and documentation become the same act rather than two chores. ## Measuring whether it works Never measure page count or template-compliance percentage; both are trivially gamed and neither correlates with value. Measure **consumption and consequence**: - **Onboarding**: time from joining to first meaningful merged change; ask new joiners which document they wished existed. New joiners are the best instrument you have and their sensitivity expires in about six weeks — harvest it immediately. - **Repeat questions**: recurring questions in team channels are documentation defects. Track the top ones; each should become a doc change or an explicit "we intentionally don't document this". - **Incident behaviour**: were the docs opened during the last five incidents? Did they help or mislead? A doc that misleads during an incident is worse than none. - **ADR coverage**: proportion of boundary-affecting changes that shipped with an ADR — a sample audit of merged PRs, not a self-report. - **Freshness distribution**: age since last verification across the hand-written corpus; watch the tail, not the mean. - **Cross-team lookup success**: can a team answer "who consumes our events?" from the portal in under five minutes? Good leading indicator: the documentation is **cited** — in PR discussions, incident reviews, and design proposals. Uncited documentation is dead weight; schedule it for deletion. ## Deliberate subtraction A corpus only grows unless someone shrinks it. Run periodic **doc deprecation**: pages with no reader, no owner, or an owner who says it is obsolete get archived (moved out of the searchable portal, kept in git history). This is politically harder than writing and more valuable — every stale page taxes the credibility of every fresh one. ## The honest trade-off More mandate buys cross-system consistency and audit readiness at the cost of engineer trust and real staleness risk (people satisfy the letter of a rule they don't believe in). Less mandate buys relevance at the cost of navigability. The stable equilibrium is: **a very small mandatory core with real teeth, a strongly recommended template, generated structure, tested rules, and outcome-based measurement** — plus willingness to delete.

  • A team argues their internal service needs no architecture documentation at all. How do you respond?
    Test it against the outsider rule rather than policy. If nothing outside the team depends on it, it has no compliance exposure and it is genuinely replaceable within a sprint, then a README plus an ADR log for anything irreversible is a legitimate answer, and I would say so — credibility comes from not mandating what has no reader. But I would check the premises: who calls it, what data it owns, what happens when the two people who know it leave, and whether it appears in any incident runbook. Those questions usually surface at least one external consumer, and that consumer's need is what the mandate exists for.
  • How do you keep a mandated template from becoming copy-pasted filler?
    Three moves. First, keep the mandatory core small enough that filling it honestly is faster than faking it. Second, require content that cannot be plagiarised: measurable quality scenarios with actual numbers, a named owner, a last-verified date, and ADRs with real rejected alternatives — filler is obvious in these fields in a way it is not in prose. Third, review for *consumption*, not compliance: sample whether anyone opened or cited the document, and prune sections nobody reads. If a section is filler everywhere across ten teams, that is evidence the mandate is wrong, not that the teams are lazy — remove it.
  • What single artefact would you create first for an organisation with fifty services and no documentation standard?
    A central index: every system, its owner, its repository, its external consumers and dependencies, and a link to whatever documentation exists. It is cheap, it is mostly derivable from repos and service registries, and it answers the highest-frequency cross-team questions — who owns this, who breaks if I change it, where do I look. It also creates the ownership metadata everything else hangs off, and it makes the gaps visible so the next investment is evidence-driven rather than a blanket mandate.

Running documentation standards across many teams is like running a library system rather than writing books. You standardise the catalogue, the shelving scheme, and who owns each section — so any reader can find anything in any branch — while leaving each branch to decide how deep its collection goes. And you weed the shelves every year, because a library nobody trusts to be current is just a warehouse.

saying these in an interview costs you the question

  • Measuring documentation health by page counts or template-compliance percentage — both are gamed and neither implies anyone reads it.
  • Mandating a full heavyweight template for every system regardless of criticality or lifespan.
  • Creating a separate documentation gate or review board instead of using pull requests and existing checkpoints.
  • Assuming a policy alone changes behaviour without scaffolding, defaults, generation and automated checks.
  • Treating documentation volume as progress and never deprecating or deleting anything.
  • Standardising depth and content while leaving structure, location and notation to each team — exactly backwards for navigability.
  • Ignoring the incident and onboarding evidence that tells you which documents actually get used.

context