skip to content

Which defects in a generated change does the build already catch, and which need a human reviewer?

level: middleimportance: must knowfreq 48%

answer

  1. Two columns, one sorting question
  2. Can a check decide it alone?
  3. Cheap checks measure surface plausibility
  4. Generators produce plausible text by construction

basics

~20 s

The build decides whatever has a deterministic oracle: resolution, types, style, declared dependencies, the existing suite. A person is needed wherever correctness depends on a rule the generator could not see, and that is where review attention belongs.

solid answer

~50 s

Sort defects by one test: can a check decide this without knowing what the change was for? Resolution, types, declared dependencies, style and the existing suite all pass that test, they fail loudly and early, and a name that does not exist fails the moment anything resolves it. What is left needs a person - whether the caller's authority is checked or only the operation performed, whether an invariant enforced in a file nobody opened still holds, what the failure path returns when the thing it assumes is missing, and whether a large distinctive block needs a provenance answer before it ships. There is an asymmetry worth saying out loud: the cheap checks all measure surface plausibility, and plausibility is precisely what a generator is good at producing, so they tend to filter a generated change less than one typed by a tired human. Spend the review on the second list.

code

text · 16 lines
text
WHAT DECIDES THIS DEFECT?

  machine, before review          person, during review
  ----------------------          ---------------------
  the name does not resolve       the caller's authority is
  the types disagree                never checked
  a dependency is undeclared      a rule enforced in another
  the style is not house style      file is broken
  an existing test now fails      the failure path returns the
                                    wrong thing to the caller
                                  a large distinctive block has
                                    no provenance answer

  Left column:  fails loudly, costs nothing, needs no intent.
  Right column: can pass everything on the left, and is the
                reason review exists at all.

go deeper

for a junior

Know that a passing build shows a change is well-formed, not that it is right, and be able to name one thing a compiler decides and one thing only a person can.

for a middle

Explain the sorting test - can a check decide this without knowing what the change was for - and use it to classify a defect rather than reciting two fixed lists.

for a senior

Show where you would convert a recurring human check into a build failure, and be honest about which checks can never move because they need the intent behind the change.

for a principal

Own the investment question: which checks the team buys, what each one removes from human review for good, and what remains that no purchase will remove.

## The sorting question Every defect in a change belongs in one of two columns, and the test that sorts them is a single question: **can a check decide this without knowing what the change was for?** If yes, a machine should own it and a reviewer should never spend attention on it again. If no, it needs a person who knows the system - and that is the entire justification for a review checklist that is not just a second copy of the pipeline. ## The column a machine owns These share one property: the standard they are measured against is written down somewhere a program can read. - **Resolution.** Whether a referenced name exists at all. This fails the moment anything resolves it, which is why it is the cheapest class there is and why it does not deserve reviewer attention. - **Types.** Whether the values agree at every boundary they cross. - **Declared dependencies.** Whether the change pulls in something undeclared, unavailable, or outside what policy permits. - **Style and shape.** Formatting, ordering, naming conventions, banned constructs. - **Existing behaviour.** Whether the suite that passed yesterday passes today. All of these fail loudly, early, and without argument. They are also, crucially, **checks on surface plausibility** - whether the text hangs together as code. ## The column a person owns - **The caller's authority.** Whether the change checks who is asking, or merely performs the operation correctly. Performing correctly is local; who may ask is a rule from outside the file. - **Invariants enforced elsewhere.** The guard every other handler applies by convention, the ordering that is true in this repository and nowhere else, the field that must be written before another is read. - **The failure path.** What happens when the thing this code assumes is absent - and whether what it returns then is what this system needs, rather than what a generic implementation would return. - **The degenerate case.** Empty, zero, one, and the case nobody wrote a test for because nobody thought of it. - **Provenance.** Whether a large, distinctive, self-contained block needs an answer about where it came from before it ships. ## Why the split bites harder on generated code Here is the mechanism, and it is the part worth being able to say out loud in an interview. The cheap checks all test whether a change *looks* like well-formed code against rules a machine can apply. A generator is optimised to produce text that looks like well-formed code. The two are measuring nearly the same property, so the cheap column tends to filter a generated change less than it filters a hand-typed one, where fatigue produces exactly the typos and mismatches those checks were built for. This is not a claim that generated code is more often wrong. It is a claim about **which** wrongness survives to review: the kind that arrives looking finished. | Question about the change | Who can decide it | Why | |---|---|---| | Does every referenced name exist? | Machine | Fails as soon as anything resolves it | | Do the types agree at each boundary? | Machine | The standard is written in the code itself | | Does it match house style? | Machine | The rule is a configuration file | | Did the existing suite stay green? | Machine | Yesterday's behaviour is recorded | | Is the caller allowed to do this? | Person | The rule lives outside the edited file | | Does the invariant enforced elsewhere still hold? | Person | Nothing in view states it | | Is the failure path right for this system? | Person | "Right" depends on what the system promises | | Does this block need a provenance answer? | Person | Judgement about shape and risk, not syntax | ## What follows in practice 1. **Move a line out of review the moment it can move.** A recurring human check that can be made mechanical - every request-path read must go through a scoping helper, say - is worth more as a build failure than as a habit, because it then applies to changes nobody reviews carefully. 2. **Do not re-check the machine's column.** A reviewer confirming that the names resolve is spending the scarcest resource on the cheapest question. 3. **Say what the green run actually proves.** It proves the change is well-formed and did not break recorded behaviour. It does not prove the change is right, and on generated code that gap is wider than usual. Two neighbouring judgements sit just outside this sort and have their own discipline: whether a set of generated tests asserts anything real, and how to read a large multi-file change produced in an unsupervised session. ## What an interviewer is listening for The sorting rule, stated as a rule rather than as two memorised lists. A candidate who can say *"ask whether a check could decide it without knowing the intent"* can classify a defect they have never seen before, which is the skill being tested. The follow-up that separates answers is the asymmetry: why the same pipeline filters less of a generated change than of a hand-typed one.

  • Can the team automate the question of whether the caller is allowed to do this?
    Partly, and it is worth doing. A check can require that every request-path read goes through a scoping helper, and fail a change that adds a raw one. What it cannot decide is whether the scope applied is the right one for this operation, because that depends on the rule the team intended. Automate the presence of the check and keep the judgement.
  • Does a strong type system move much from the human column into the machine one?
    It moves real work. Making an illegal state unrepresentable turns a class of reviewer questions into compile errors, which is the best kind of automation because it applies to every change. It does not move intent: a change can be perfectly typed and still do the wrong thing, and well-typed is something a generator is good at.

saying these in an interview costs you the question

  • Says a green pipeline means a generated change is correct
  • Believes a scanner can decide whether the scope a change applied is correct
  • Thinks review should re-check what the compiler already rejected
  • Treats plausible-looking code as evidence that it was checked
  • Says nothing needs a human once static analysis is thorough enough