What makes a code review checklist effective, and why do review checklists decay over time?
answer
- It is a memory aid, not a specification
- Where should the items come from?
- Only one direction of growth is easy
- Machine-decidable items dilute human ones
- Give it an owner and an item budget
basics
~20 sAn effective review checklist is short, drawn from defects this team actually let through, and made of items only a human can judge. Checklists decay because they only ever grow, the system's failure modes move on, and long lists get ticked ritually instead of read.
solid answer
~50 sA checklist earns its place when it encodes *this* team's escaped defects rather than generic advice: each item should trace back to something that once reached production here. Keep it short — roughly the number of items a reader can hold without consulting the list — phrase items as questions the reader must answer rather than commands to obey, and keep out anything a machine can decide, because mechanical items train reviewers to skim. Decay has three causes: the list only ever grows, since adding an item after an incident is easy and removing one feels like removing a safeguard; the system changes, so items describe risks that no longer exist while new failure modes go unrepresented; and once a list is long, reviewers satisfy it rather than think with it. The remedy is treating the checklist as a living artefact with an owner, a review cadence and a hard item budget.
go deeper
Know what a review checklist is for and that shorter is generally better. Be able to say why a list of things the build already verifies wastes a reviewer's attention rather than protecting it.
Explain the design rules and the decay mechanisms: provenance from real escapes, question form, an item budget, and why lists only ever grow unless somebody is accountable for removals.
Show that you can run the lifecycle — turning an escaped defect into the question that would have caught it, retiring items that catch nothing, and splitting one bloated list into a default plus a high-risk supplement.
Own the limits of the instrument. Be ready to argue when a checklist is the wrong answer entirely, and what you would change instead when reviews are missing defects because of thin domain understanding rather than forgetfulness.
## What a checklist is for A review checklist is a memory aid for a reader under time pressure. It exists because human attention is unreliable in a predictable way: reviewers notice what they are currently thinking about and miss whole categories of problem simply because those categories never came to mind. A checklist restores the categories. It is not a specification of a correct change and it is not a scoring rubric — treat it as either and it stops helping. That framing already implies most of the design rules. ## What makes one effective **Provenance.** The strongest items come from defects this team has already let through review. After an escape, the question worth asking is not "who missed it" but "what question, asked during review, would have surfaced it?" — and that question, not the incident, becomes the item. A checklist adopted wholesale from a published list is generic by construction: it encodes someone else's failure history, which overlaps yours only by coincidence. **Brevity.** A checklist a reader can hold in their head is one they will actually apply; a long one is a document they consult once and then approximate. Different sources put the practical ceiling in the range of roughly seven to twelve items; the exact number matters less than the discipline that adding an item usually means retiring one. **Judgement-shaped items.** Every item should require a human to think. "Does the error path leave stored data consistent if the second write fails?" needs a reader; a property a machine can decide should be decided by the machine, and its presence on a human list is worse than useless because it teaches the reader that most items can be ticked without looking. **Question form.** "Are the boundary values covered by an assertion?" produces a search; "Cover boundary values" produces agreement. Questions are harder to satisfy dishonestly. **Scope.** A single list for every change is either too generic to help or too long to use. Most teams do better with a small default list plus a short supplementary list for a narrow category of high-risk change — data migrations, changes to money handling, changes to an interface other teams depend on. ## Why checklists decay **They only grow.** Adding an item is the natural response to an incident and costs nothing at the moment of adding; removing one requires arguing that a named risk is now acceptable, which nobody wants to be on record for. The result is monotonic growth, and growth is decay: at twelve items a reviewer reads, at forty they scan, and attention spreads so thin that the three items that actually matter get the same half-second as the rest. **The system moves and the list does not.** Items describe the architecture that existed when they were written. As components are replaced, whole categories of risk disappear — while genuinely new ones, belonging to whatever replaced them, are represented nowhere. A list can be fully satisfied and still miss the current failure modes entirely. **Automation hollows it out.** Over time, more of a checklist's mechanically-decidable items become things the build already decides. Those items stay on the list, get ticked instantly every time, and act as filler that dilutes the items still requiring thought. **Ritual replaces reading.** Once a checklist is attached to an approval, satisfying it becomes the goal. Reviewers learn the shape of a compliant answer and produce it. The list then measures conformance to the list, which is not the thing anyone wanted to know. **Ownership evaporates.** Most lists are written once by whoever cared and then belong to nobody. With no owner and no cadence, no one is responsible for the removals that would keep it short. ## Keeping one alive Give it an owner and a standing review — quarterly is a common cadence. At each review ask, per item: has this caught anything since we last met, and could a machine decide it now? Items that catch nothing get retired; items a machine can decide get moved to the build and removed from the human list. Enforce a hard budget so adding requires removing. Feed it from incidents rather than from opinion, and phrase each addition as the question that would have caught the escape. And be honest about what a checklist cannot do. It restores categories a reader might forget; it does not supply the domain understanding that lets someone see that a correctly-written calculation is computing the wrong thing. A team whose reviews depend on the checklist rather than being reminded by it has a knowledge problem the list will not fix.
- Your checklist has thirty items and every one was added after a real incident. How do you shorten it?Judge items by what they have caught since they were added, not by the incident that created them. Retire the ones with no recent catches, move anything mechanically decidable into the build, and merge near-duplicates into one broader question. Then split what remains: a short default list for every change, and a supplementary list that only applies to the narrow categories of change where the retired risks actually lived.
- How do you tell whether a checklist is being read or just satisfied?Look at what review produces, not at whether the list was completed. If findings cluster on the same two or three items while whole sections never generate a comment, those sections are being ticked rather than answered. Reviewers completing the list in less time than reading the change would take is another signal. Asking a reviewer which item was hardest to answer on a recent change gets an honest read quickly.
A checklist is a packing list, not an inventory of your house: it works because it is short enough to run through at the door, and it stops working the moment somebody adds every object you own.
saying these in an interview costs you the question
- Adopts a published checklist unchanged and calls it done
- Believes a longer checklist catches more defects
- Keeps items a machine could decide on the human list
- Treats the checklist as a specification of a correct change
- Adds an item after every incident and never removes one
- Assigns the checklist no owner and no review cadence