skip to content

What belongs on a review checklist for AI-assisted changes that the ordinary checklist does not?

level: seniorimportance: should knowfreq 44%

answer

  1. Add lines, not a document
  2. Start from what the build cannot decide
  3. Invariants living outside the edited file
  4. A machine-enforceable line gets deleted

basics

~20 s

Only what the build and the ordinary checklist cannot already cover. In practice a handful of lines: invariants the generator could not see, the caller's authority, provenance for a distinctive block, and whether the author can explain every line.

solid answer

~50 s

The ordinary checklist does not change: the change still has to work, be tested, be readable, and fit the design. What an assisted-change section adds is short, and follows from one fact: the change was drafted by something that could not see your whole system. So it asks whether the change honours an invariant enforced in a file nobody opened, whether the caller's authority is checked or only the operation performed, whether a large distinctive block needs a provenance answer, and whether the author can walk through every line they are asking to merge. It deliberately leaves out anything the compiler, the linter or the suite already decides, because a line a machine can enforce belongs in the pipeline and not on a list a tired reviewer skims. Keep it to a handful of lines, give it an owner, and delete a line the day tooling starts catching it.

code

text · 17 lines
text
REVIEW CHECKLIST - the lines we added for assisted changes
(everything above this section is unchanged)

[ ] Invariants: does this honour a rule enforced somewhere
    the author's editor was not showing? Name the rule.

[ ] Authority: does the change check who is asking, or only
    perform the operation correctly?

[ ] Provenance: is any block large and distinctive enough
    that we should ask where it came from before it ships?

[ ] Author's model: can the author explain every line in
    review, without re-reading the whole change first?

Owner: the platform team. Revisited each quarter.
A line moves into CI the day a check can decide it.

go deeper

for a junior

Know that the ordinary review bar does not drop for assisted code, and that the extra lines are about what the tool could not see rather than about the tool. Be ready to name two of them.

for a middle

Explain why a line belongs on a list rather than in the pipeline: a pipeline decides things without knowing intent, and the list holds what needs a person who knows this system.

for a senior

Show how you would derive the lines from a defect that actually shipped, keep the list short enough to be used under pressure, and retire a line once a check can enforce it.

for a principal

Own who maintains the standard, what it costs in review time across teams, and how you would tell whether it is working rather than merely being filled in.

## What the artefact is for A review standard for AI-assisted contributions is a **short, written list of questions a reviewer asks that the machinery cannot ask for them**. It is not a policy document, not a statement of values, and not a taxonomy of everything that can go wrong. Its whole value is that it is read at the moment of review, by someone with limited attention, and changes what they look at. That framing settles most of the arguments about its contents. Every line competes for the same scarce attention, so a line has to earn its place against the line it displaces. ## The situation it usually gets written in A team ships a bug in an assisted change, and the week after, someone volunteers to write the review standard. The pull is toward a long document that answers that one incident in detail. What survives contact with a busy Tuesday is about five lines and a named owner. The useful post-incident question is not *what else could go wrong* but a pair: 1. **Which line would have caught this one?** Write that line, in the words of the actual defect. 2. **Could a machine have caught it instead?** If yes, buy the check and do not write the line at all. ## What the ordinary checklist already covers Nothing about assistance lowers the existing bar, and restating it wastes the list. The ordinary review already asks whether the change works, whether it is tested, whether it reads clearly, and whether it fits the design the team chose. Those questions are unchanged, and a reviewer who was asking them before keeps asking them. ## The lines that are genuinely specific Each of these follows from the same mechanism: **what the tool could see determines what it could get right**, and everything outside its view arrives looking finished. - **Invariants enforced elsewhere.** A rule that lives in another file, a guard the team applies by convention, an ordering constraint that is true in this repository and in no general codebase. The change can be locally perfect and still break it. - **The caller's authority.** Whether the code checks who is asking, or merely performs the operation correctly. These are separate concerns and only the second is visible locally. - **Provenance.** Whether any block is large, distinctive and self-contained enough that someone should ask where it came from before it ships, and who that question goes to. - **The author's own model of the change.** Can the author explain, in review, why each part is there? This is the highest-value line on the list, because the failure that is specific to assisted work is a *plausible* change whose author never formed a model of it. - **Disclosure, if the team has decided it wants one.** What the change records about how it was produced, and what that record is and is not used for. ## What must stay off it | Tempting line | Why it does not belong | Where it lives instead | |---|---|---| | "Every referenced name exists" | Fails the moment anything resolves it | The build | | "Dependencies are declared and permitted" | Decided mechanically, with no intent needed | The dependency check | | "Style matches the house style" | Settled before review opens | The formatter | | "The suite is green" | The pipeline reports it already | CI | | "The design is right" | True of every change, not this class | The ordinary checklist | Two further subjects sit next to this list and are deliberately not on it: reading a large multi-file change produced by an unsupervised session, and judging whether generated tests assert anything, are each their own review standard with their own discipline. ## Keeping it alive A checklist with no owner rots quietly, and a rotten checklist is worse than none because it is still being ticked. Three habits keep it honest: 1. **Name an owner and a review date.** Someone has to be able to delete a line. 2. **Retire a line when tooling catches it.** The list should shrink as the checks improve; a line whose job a machine can now do is pure cost. 3. **Make it a submission list too.** The author runs it before asking for review. Most of its value is spent before a reviewer ever opens the change, and it stops the list from reading as suspicion of the author. ## What an interviewer is listening for Candidates who answer this well do three things. They keep the list short and can say what they left off. They can justify each line by naming what makes it undecidable by a machine. And they treat the standard as something with an owner and an expiry, not a document. A candidate who recites a long list of risks has described the subject; a candidate who names five lines and the reason each survived has described an artefact their team could actually use tomorrow.

  • How do you tell whether the checklist is working rather than just being ticked?
    Look at what it changes, not at compliance. Record what review actually catches and which line caught it. A line that never catches anything is a prompt to ask whether it is badly worded or the risk has moved, not an automatic deletion. And watch incidents: a repeat of the defect that prompted the list is the honest signal that the line was written for the wrong thing.
  • A reviewer wants to add ten more lines after a second incident. What do you say?
    Ask which single line would have caught it, and whether a machine could have caught it instead. One precise line earns its place; ten do not, because length is what makes a list skimmed under time pressure. Add the one, retire something that tooling now covers, and keep the rest of the reasoning in the incident write-up where it belongs.

A cockpit checklist is short because it is read under pressure, while the thick manual it was distilled from stays on the shelf. A review list works the same way: it holds only what a person must decide in the moment, and everything mechanical belongs in the machinery.

saying these in an interview costs you the question

  • Thinks a longer checklist catches more defects
  • Believes the list should restate the build's own checks
  • Says a green pipeline means no checklist is needed
  • Treats the list as review-only, never used before submitting
  • Expects the same lines to stay correct as tooling improves