skip to content

You own engineering standards for a large multi-team codebase. How would you decide which knowledge belongs in inline comments versus other artifacts, and how would you keep TODO markers from decaying into noise?

level: principalimportance: nice to knowfreq 22%

answer

  1. one system of record per knowledge type; link, don't restate
  2. docs-as-code: same repo, same PR, same CI
  3. ADRs: context/options/decision/consequences, superseded not edited
  4. TODO needs owner + issue + removal condition; sweep on closed issues
  5. measure staleness by sampling, not page count

basics

~20 s

Give each kind of knowledge one home: code and tests for behaviour, comments for local why and hazards, commit history for change history, issue tracker for planned work, ADRs for decisions, generated reference for public APIs. Require TODOs to link an issue and expire them.

solid answer

~50 s

Assign a **single system of record per knowledge type** and forbid restating it elsewhere — link instead. Behaviour lives in code and tests; local rationale, hazards and non-obvious constraints live in inline comments; public contracts live in doc comments and generated reference; significant decisions live in dated **ADRs** (context, options, decision, consequences, superseded-by); change history lives in commit messages; planned work lives in the tracker; operational knowledge lives in runbooks near the service. Keep documentation in the repo, reviewed and CI-checked (docs-as-code): link checkers, generated API reference, examples compiled as tests, and doc-comment lint on public symbols. For TODOs, define a lifecycle: a marker must carry an owner or issue ID and a removal condition; CI rejects bare `TODO`; a periodic sweep deletes markers whose issue is closed or unjustified; long-lived ones are converted into tracked debt with an expiry. Measure rot by sampling, not by counting pages — volume is not the goal, trust is.

code

pseudocode · 7 lines
pseudocode
// Enforceable TODO shape: owner + tracked issue + removal condition
// TODO(#4412, @payments): remove this fallback once all v1 clients
// are retired (tracked for Q3; fallback masks a 3% duplicate-charge rate).
if (request.apiVersion == 1) { return legacyCharge(request) }

// Rejected by lint - no owner, no issue, no removal condition:
// TODO: clean this up

go deeper

for a junior

Say each kind of information should have one home — code for behaviour, comments for why, tickets for future work — and TODOs should link an issue.

for a middle

Map knowledge types to artifacts, argue link-don't-restate, and describe a basic TODO rule (owner plus issue ID) enforced in review or lint.

for a senior

Add docs-as-code with CI checks (link checking, generated reference, examples as tests), ADR structure, and a TODO lifecycle with sweeps; explain why comment mandates backfire.

for a principal

Own the whole system: single systems of record, narrow enforceable mandates only, ADR bar and supersession, staleness sampling and incident-driven signals, deletion as a first-class activity, and the organisational trade-off between tooling and review culture.

## The problem at scale In a single service, comment policy is a style question. Across dozens of repositories and teams, it becomes an information-architecture problem with three failure modes: 1. **Duplication** — the same fact stated in a comment, a README, a wiki page and a slide deck. Rot is guaranteed, because only one copy is ever updated. 2. **Homelessness** — knowledge with no assigned artifact (why we picked this queue technology, what to do when the nightly job fails) lives only in people, and leaves when they do. 3. **Ceremony** — mandates ("every public method must have a doc comment") that produce unread, unmaintained text and train engineers to view documentation as compliance rather than communication. ## Step 1: one system of record per knowledge type | Knowledge | Home | Why there | |---|---|---| | What the code does | The code itself | Executable, always current | | Behavioural contracts, edge cases | Tests, named for the behaviour | Fails when the contract breaks | | Invariants, units, nullability | Types and value objects | Enforced by the compiler/runtime | | Local why, hazards, external quirks | Inline comments beside the code | Nothing else can hold them; proximity keeps them alive | | Public API contract | Doc comments → generated reference | Callers must not read bodies; generation avoids hand-copying | | Significant decisions | **ADRs** in the repo | Dated, immutable, superseded rather than edited | | Change history, deleted code | Commit messages, blame, history | Complete, accurate, searchable, free | | Planned/deferred work | Issue tracker | Has owners, priority, lifecycle | | Service operation | Runbooks near the service | Read under incident pressure | | Onboarding orientation | A short README per repo | Entry point that links, not duplicates | The rule that makes this work is **link, don't restate**. A comment saying `// duplicated deliberately to avoid a cross-service join — ADR-014` stays true forever, because it delegates the volatile part. ## Step 2: docs-as-code, with CI Knowledge should live in the same repository as the code it describes, in the same pull request, under the same review. That gives it three properties wikis lack: it changes atomically with the code, reviewers see it in the diff, and CI can check it. What CI can realistically enforce: - **Link checking** — no dangling ADR, ticket or page references. - **Generated API reference** from signatures; fail if it does not build. - **Examples as tests** — doctests, or snippets extracted and compiled so a wrong example breaks the build. - **Doc-comment lint on public/exported symbols only** (documented parameters must exist, etc.) — scoped narrowly so it does not become a comment mandate. - **Commented-out-code detection** at warning level. - **TODO format rules** (below). What it cannot enforce: whether a rationale is still true. That gap is closed by review culture and periodic sweeps, not tooling. ## Step 3: a TODO lifecycle TODO/FIXME markers are useful precisely because they are cheap and co-located with the problem — and dangerous for the same reason: they accumulate silently and become archaeological. A workable policy: 1. **Format rule, enforced.** Every marker carries an owner or issue ID and, ideally, a removal condition: `// TODO(#4412): delete once v1 clients are retired.` A bare `TODO` fails lint. 2. **Tracker linkage.** The referenced issue is the real backlog item; the comment is a pointer at the exact code location. This keeps the tracker authoritative and the comment stable. 3. **Expiry / sweep.** Periodically (a quarterly hygiene pass, or automation that reads issue status) flag markers whose issue is closed — the code either shipped the fix or the marker is stale. Some teams use *expiring* TODOs that fail the build after a date or after a dependency version ships; this is powerful but needs care so it does not break unrelated deploys. 4. **Severity split.** Distinguish `TODO` (nice to do), `FIXME` (known defect), `HACK`/`XXX` (deliberate compromise). Only the defect classes get tracked as debt with owners. 5. **Budget, not zero.** Aiming for zero markers pushes people to delete rather than track; aiming for "every marker is traceable and justified" is achievable. ## Step 4: measure trust, not volume Documentation programmes fail when success is measured in pages or coverage percentages. Better signals: - **Sampling audits**: pick N random comments/docs per quarter, verify against the code, report a staleness rate. - **Onboarding time-to-first-PR** and the questions new joiners ask — repeated questions point at homeless knowledge. - **Incident retrospectives**: was a runbook missing or wrong? That is a documentation defect with a real cost attached. - **Bus-factor mapping** per critical component. ## Trade-offs a principal should articulate - **Mandates versus judgement.** Broad comment mandates reliably manufacture rot; narrow, checkable mandates (public API surface, ADRs for irreversible choices) pay off. Prefer the narrow ones. - **ADR cost.** ADRs are cheap to write and enormously valuable years later, but only if the bar is "significant and hard to reverse". An ADR per trivial choice devalues the set. - **Wikis are not banned, they are demoted.** They suit cross-cutting, non-versioned material (org processes, glossaries). Anything that must match a specific code version belongs in the repo. - **Deleting is a legitimate documentation activity.** Removing a stale page or comment increases the trustworthiness of the rest; treat deletion as a contribution, not a loss. - **Deprecation and machine-readable annotations.** Where the language supports it (`@Deprecated`, attributes, lint annotations), prefer them to prose warnings — they surface in the IDE and can fail the build.

  • How do you stop an ADR set from becoming as stale as the wiki you replaced?
    Make records immutable and dated: a decision is never edited in place, it is superseded by a new record that links back. Staleness then becomes visible (a decision from three years ago with no supersession is a prompt to review) instead of invisible, and the code links to a stable identifier rather than to mutable prose.
  • What is the risk of TODO comments that fail the build after a deadline, and how would you mitigate it?
    They can break unrelated deploys at the worst moment, creating pressure to disable the mechanism entirely. Mitigate by failing a scheduled hygiene job or a warning lane rather than the release pipeline, tying expiry to a concrete condition such as a dependency version rather than a wall-clock date, and giving teams a documented way to renew with justification.

Treat knowledge like a library catalogue: each fact gets one shelf, and everything else holds a call number pointing to it. The moment two shelves hold copies of the same book, one of them quietly becomes the wrong edition.

saying these in an interview costs you the question

  • Mandating a doc comment on every method, which mass-produces unread, rotting prose.
  • Keeping a wiki page that restates code details, then treating it as authoritative during incidents.
  • Measuring documentation success by page count or comment coverage rather than by staleness and time-to-answer.
  • Allowing bare, unowned TODOs and calling the backlog 'in the code'.
  • Editing an old decision record in place, destroying the historical context that made it useful.

context