skip to content

Comments and Documentation

The useful comment explains why, not what, and the rest are noise that rots as soon as the code moves on. You will learn the categories worth writing (intent, warning, legal, TODO) and why commented-out code should just be deleted.

part ofSoftware design & architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In code review, why is a comment that explains WHY the code does something usually more valuable than a comment that restates WHAT the code does?

level: juniorimportance: must knowfreq 72%

answer

  1. code = what, comment = why
  2. comment restating code rots into a lie
  3. no compiler checks comments
  4. unclear what → rename/extract, don't annotate
  5. public API docs are the sanctioned exception

basics

~20 s

The code already shows what it does; a reader can see that. It cannot show the reason, constraint, or bug that forced this approach. "Why" comments add information; "what" comments repeat it and can go stale.

solid answer

~50 s

Source code is an exact, always-current description of WHAT happens — the compiler or interpreter enforces it. A comment restating that (`i = i + 1 // increment i`) adds zero information, costs reading time, and becomes a lie the moment someone edits the line without editing the comment. What the code cannot express is intent and context: why this algorithm instead of the obvious one, which upstream bug or spec clause forced the odd branch, what happens if you "simplify" it. That knowledge lives only in the author's head, so writing it down is the highest-value comment there is. Practical rule: if a reviewer could delete the comment and lose nothing, delete it; if deleting it would let a future maintainer confidently break something, keep it. And prefer fixing an unclear WHAT by renaming or extracting a function rather than annotating it.

code

pseudocode · 7 lines
pseudocode
// BAD - restates the code, will rot
timeout = timeout * 2   // multiply timeout by two

// GOOD - states the reason, cannot be read off the code
// Exponential backoff: the payment gateway rate-limits us at ~5 rps
// and returns 429 without a Retry-After header (vendor ticket #8812).
timeout = timeout * 2

go deeper

for a junior

Say the code shows what, the comment should show why; give one concrete example of each and mention that stale comments mislead.

for a middle

Add the rot mechanism (nothing verifies comments), and note that an unclear 'what' usually signals a naming or extraction problem rather than a missing comment.

for a senior

Frame it as information content per line read, list the legitimate exceptions (public API docs, dense algorithms, external constraints), and treat comment updates as part of code review.

for a principal

Discuss where rationale should live across a codebase — inline comments vs. commit messages vs. ADRs vs. tests — and how to keep documentation from becoming an unverified liability at scale.

## The two kinds of information in a source file A source file carries two very different things: 1. **Mechanics — WHAT the program does.** This is fully encoded in the code itself. It is *executable*, so it is *verified*: if the code says `total = total + item.price`, that is definitively what happens. Nothing can drift out of sync with it, because it *is* the thing. 2. **Rationale — WHY it does it that way.** This is *not* encoded anywhere. Why did we retry three times and not five? Why do we skip validation for records created before a certain date? Why is the loop written backwards? The reason existed in the author's head at 3pm on some Tuesday and, unless written down, it evaporates. Comments are the only place category 2 can live inside the code. Spending them on category 1 wastes the one channel you have. ## Why restating WHAT is actively harmful, not merely useless - **Zero information gain.** `// increment the counter` above `counter++` tells a reader nothing they did not already read. It is noise, and noise trains people to skim past *all* comments — including the rare important one. - **It rots.** Code is compiled/executed and covered by tests; comments are not checked by anything. When someone changes `counter++` to `counter += batchSize`, the comment stays. Now the file contains a *false statement*, and a false statement is worse than no statement: a reader who trusts it makes a wrong decision. This drift is called **comment rot** or **stale comments**. - **It hides a naming problem.** A comment explaining what an obscure line does is usually a symptom: the variable is named `d`, or the function is 200 lines. Deleting the comment and renaming `d` to `elapsedDays` fixes the problem permanently and for every reader, including tooling and search. ## What good WHY comments look like - **Intent / decision:** `// Binary search, not linear: this list is on the hot path and can hold ~1M entries.` - **Constraint from outside the code:** `// Vendor API returns 200 with an error body; we must parse the body to detect failure.` - **Warning of consequence:** `// Not thread-safe: callers must hold the session lock.` - **Non-obvious correctness:** `// The +1 is because the API's end date is exclusive.` - **Reference to a decision record or ticket:** `// See ADR-014: we deliberately duplicate this field to avoid a cross-service join.` Each of these is information a reader genuinely cannot obtain from the code. ## The nuance: "what" comments are not *always* wrong The rule is a heuristic, not a law. Legitimate summarizing comments exist: - **Public API documentation** (Javadoc/docstrings/XML docs) intentionally describes what a function does and what its contract is, because callers should not have to read the body. Here the audience is external, and "what" is the point. - **A one-line summary above a dense but irreducible block** — a hand-tuned numeric routine, a regular expression, a bit-twiddling trick — can save minutes even though it restates behavior, because the code's "what" is genuinely hard to decode. - **Very high-level orientation:** `// --- Phase 2: reconcile local and remote state ---` in a long procedure. Better still is to extract that phase into a named function, but the comment beats nothing. So the real test is not the literal word "why" — it is **information content per line of reading**. Ask: *does this comment tell the reader something the code cannot?* If yes, keep it. If no, either delete it or, better, improve the code so the question never arises. ## Edge cases and trade-offs - **Over-correcting.** Some teams ban comments entirely, which loses genuine rationale. The goal is fewer, better comments — not zero. - **Comments as apology.** `// sorry, this is a mess` is not a why-comment; it is a TODO in disguise. Either fix it or file a tracked item with a link. - **Cost of maintenance.** Every comment you keep is a line you must update on every relevant change. WHY comments survive refactoring far better than WHAT comments, because rationale changes less often than mechanics — another reason to prefer them. - **Reviewers should treat a comment as code.** If a pull request changes a line but not the comment above it, that is a review finding.

  • If a comment is needed to explain what a confusing line does, what should you usually do instead of writing it?
    Improve the code: rename the variables to say what they hold, or extract the line/block into a well-named function so the name carries the explanation. That fixes it for every reader and cannot drift out of sync.
  • Are documentation comments on public functions (Javadoc, docstrings) an exception to 'don't say what'?
    Yes. Their audience is callers who should not read the body, so describing the contract — parameters, return value, thrown errors, thread-safety, side effects — is exactly the job. They are also often tooling-checked and published, which reduces rot.

A road sign that says "this is a road" is useless — you can see that. A sign that says "bridge freezes before road surface" tells you something the road itself cannot, and that is why it is worth maintaining.

saying these in an interview costs you the question

  • "Every function needs a header comment" — mandated comment quotas produce noise and rot.
  • "More comments always means more maintainable code."
  • "The comment is right, the code must be wrong" — comments are unverified; trust the code and investigate.
  • Adding a comment to explain a bad name instead of fixing the name.
  • Believing a comment is safe because it was accurate when written — nothing prevents it from drifting.

context

open as a page

Clean-code guidance splits comments into a small set of justified categories and a longer list of harmful ones. Name several of each and state the single criterion that decides which side a comment falls on.

level: middleimportance: must knowfreq 58%

basics

~20 s

Good: legal/licence headers, intent, clarification of code you can't change, warnings of consequences, TODOs, and public API docs. Bad: redundant restatement, journal/changelog entries, commented-out code, noise, banners, attributions. Criterion: does it tell the reader something the code cannot?

open as a page

What does "self-documenting code" mean in practice, which specific refactorings let you replace a comment with code, and where does that technique stop working?

level: seniorimportance: must knowfreq 55%

basics

~20 s

It means naming and structuring code so it explains itself: rename vague identifiers, extract a commented block into a well-named function, and put a complex condition into a named variable. It fails for reasons outside the code — external constraints, hazards, and decisions.

open as a page

A teammate leaves a 40-line block of commented-out code in a pull request, arguing "we might need it back". What concrete harms does that cause, and is there any situation where keeping it is defensible?

level: middleimportance: should knowfreq 48%

basics

~20 s

It is dead text nothing checks: no compiler, no tests, no refactoring tool touches it, so it silently goes stale. Readers won't delete it because they assume it matters. Version control already stores it — delete it and recover from history if needed.

open as a page

Comments and documents drift out of sync with the code they describe. Explain the mechanism behind this drift and the techniques that keep explanatory material trustworthy over time.

level: seniorimportance: should knowfreq 38%

basics

~20 s

Code is verified by compilers and tests; prose is verified by nobody. So edits update the code and skip the comment, and the comment quietly becomes false. Fixes: fewer comments, put facts in code/tests/types, and review comments as code.

open as a page

You own engineering standards for a large multi-team codebase. How would you decide which knowledge belongs in inline comments versus other artifacts, and how would you keep TODO markers from decaying into noise?

level: principalimportance: nice to knowfreq 22%

basics

~20 s

Give each kind of knowledge one home: code and tests for behaviour, comments for local why and hazards, commit history for change history, issue tracker for planned work, ADRs for decisions, generated reference for public APIs. Require TODOs to link an issue and expire them.

open as a page