skip to content

As a technical leader, how do you stop architecture diagrams from becoming stale and misleading across many teams — which diagrams do you maintain, which do you generate, and which do you deliberately throw away?

level: principalimportance: should knowfreq 22%

answer

  1. Stale is worse than missing
  2. Small maintained set; rest disposable
  3. Text model in repo, same PR as the code
  4. Generate volatile levels from ground truth
  5. Owner + date + status; delete unowned

basics

~20 s

Keep few diagrams, store them as text beside the code, and review them in the same pull request as the change. Generate the volatile low-level ones from code and infrastructure; throw away workshop sketches; delete anything nobody maintains rather than leaving it wrong.

solid answer

~60 s

Stale diagrams are worse than none, because readers act on them. The leadership move is to shrink the maintained set and raise its reliability. Concretely: (1) define a **minimum maintained set** per system — context, container, one production deployment view, and one or two dynamic views for critical flows — and explicitly declare everything else disposable; (2) keep the model as **text in the repo** (Structurizr DSL, C4-PlantUML) so a change ships in the same pull request as the code, is diffable, and renders in CI; (3) **generate what can be generated** — component and code views from static analysis, deployment views from infrastructure-as-code or a service catalogue, dynamic views from distributed traces — and hand-author only intent; (4) assign an **owner per diagram** and add lightweight drift checks (parse failures fail the build; a diagram older than its system's last significant change is flagged in review); (5) link every diagram to the **ADRs** that explain the why, since pictures show structure but not rationale; (6) mark status and date on each — current, proposed, or archived — and **delete rather than tolerate** unowned ones.

go deeper

for a junior

Say diagrams should live with the code as text, be updated in the same change, and that a wrong diagram is worse than no diagram.

for a middle

Describe the workflow concretely: model in the repo, rendered in CI, updated in the same pull request, dated with a status, and few in number.

for a senior

Add the generation strategy by volatility (code/component/deployment generated; context/container hand-authored), ADR linkage, ownership, and lightweight drift detection.

for a principal

Set org-level policy: the minimum maintained set, per-team models with shared conventions plus an aggregated landscape, what is generated from ground truth, deletion as a first-class action, and health signals based on usage and correction rate rather than diagram count.

## Framing the problem An out-of-date diagram is not neutral. People make change-impact, incident, security, and hiring decisions from it. A missing diagram makes someone ask a question; a wrong diagram makes them confidently do the wrong thing. So the goal is not "more documentation" — it is **a small, trusted set**. Three forces cause drift: - **Cost of update** is paid by the person changing the code, while the benefit accrues to future readers → classic externality, so reduce the cost and put it in their path. - **Distance**: diagrams stored in a wiki or a slide deck are far from the change; nothing prompts an update. - **Volume**: the more diagrams exist, the less each is maintained; effort per diagram falls as their count rises. ## The playbook ### 1. Define the maintained set (and say what is disposable) Publish a policy, e.g. per system: **System Context**, **Container**, **one production Deployment view**, **1–3 Dynamic views** for the money/auth/failure paths, and Component views only for containers above a complexity threshold. Explicitly declare workshop sketches, review-deck options, and incident whiteboards **disposable** — no one is expected to maintain them, and they must not be filed as documentation. ### 2. Put the source next to the code Text model (Structurizr DSL, C4-PlantUML, Mermaid) in the repository. Consequences: pull-request diffs show architectural change as a reviewable artefact; the reviewer sees a new dependency appear in the model; the barrier to a small update drops to editing two lines. ### 3. Render and validate in CI Build fails if the model does not parse. Rendered SVGs publish to the docs site so nobody hand-exports images. Optional model assertions: every container has an owner tag; nothing reaches the data store except its owning service; no element is tagged `deprecated` without a removal date. ### 4. Generate what can be generated The volatility ranking is roughly: code > components > deployment > containers > context. Invert your effort accordingly: - **Code / component views**: derive from static analysis, module metadata, or annotations; regenerate on demand and never hand-maintain. - **Deployment views**: derive from infrastructure-as-code, service catalogue, or cluster inventory; keep hand-drawn ones at the stable-topology level (tiers, zones), not instance names. - **Dynamic views**: export real flows from distributed tracing as *evidence*; hand-draw only the *intended* flow, and treat divergence between the two as a finding. - **Context / container**: hand-authored, because they encode intent and boundaries that no tool can infer. ### 5. Ownership and lifecycle Every maintained diagram has an owning team, a title stating its abstraction level, a **date**, and a **status** (current / proposed / archived). Archived diagrams move to a clearly marked area — they remain useful history, like the point-in-time reviews teams keep. Unowned diagrams get deleted; deletion is a legitimate and healthy documentation action. ### 6. Pair pictures with rationale A diagram shows *what*, never *why*. Link each maintained view to the **ADRs** (Architecture Decision Records) behind it, so a reader who disagrees with the shape can find the decision, its context, its options, and its consequences — instead of re-litigating it. Structurizr-style tooling can host documentation and decision records alongside the model. ### 7. Make it visible in the process - Include "does the model change?" in the pull-request template for architecturally significant changes. - Use the container/deployment views as standing inputs to threat modelling and to incident reviews — a view that gets *used* gets fixed. - Onboarding is the best drift detector: ask the newest engineer to follow the container diagram and report every discrepancy. ## Trade-offs and failure modes - **Over-modelling**: an exhaustive enterprise model nobody reads consumes real effort and creates a false sense of control. Start from questions people actually ask. - **Automation absolutism**: fully generated diagrams show what *is*, never what was *intended* or what is *deprecated*. Intent must stay hand-authored, or the diagram loses its architectural content and becomes a dependency dump. - **Gatekeeping**: mandatory heavyweight diagram review slows delivery and drives people to draw in private tools. Keep the bar low and the set small. - **Tool centralisation vs team autonomy**: one org-wide model gives consistency but creates a bottleneck and merge contention; per-team models with a shared convention and an aggregated landscape view usually scale better. - **Measuring the wrong thing**: counting diagrams rewards volume. Better signals: how often views are opened, how often they are corrected during onboarding, whether incident reviews cite them. ## The one-line principle Maintain the few diagrams that answer questions people repeatedly ask; generate the volatile ones from ground truth; make everything else explicitly disposable; and delete anything unowned rather than letting it lie.

  • How do you detect drift without a heavyweight review process?
    Cheap signals: the build fails if the model does not parse; the pull-request template asks whether the model changes for architecturally significant work; new joiners are asked to walk the container diagram and log discrepancies; and generated views (from infrastructure-as-code or traces) are diffed against the hand-authored intent, where any divergence is a finding to explain or fix.
  • Why not simply generate every diagram automatically?
    Because generated views show what exists, not what was intended. They cannot express boundaries you are enforcing, elements you are deprecating, a target state, or why a system is split this way. Generate the volatile, factual levels; hand-author intent at context and container level; link both to ADRs for rationale.
  • An org-wide single model or per-team models?
    Per-team models with a shared convention (same DSL, same C4 abstractions, same tags) plus an aggregated landscape view usually wins: it preserves autonomy, avoids merge contention and a central bottleneck, and still gives portfolio-level consistency. A single monolithic model is defensible only in small orgs or where regulation demands one authoritative source.

Treat diagrams like tests: a small suite that runs on every change and is trusted beats a huge suite that is red and ignored. And like tests, deleting the ones nobody maintains improves the signal.

saying these in an interview costs you the question

  • Mandating that every service produce all four C4 levels — volume kills maintenance
  • Leaving an unowned, undated diagram up because 'some of it is still right'
  • Storing diagrams as exported images with the source lost
  • Treating diagrams as a substitute for ADRs — pictures never carry rationale
  • Measuring documentation health by diagram count
  • Assuming full automation solves it, when intent and deprecation cannot be inferred from running systems

context