skip to content

When is a mechanical rewrite across every automated case unsafe, compared with editing each case by hand?

level: middleimportance: nice to knowfreq 24%

answer

  1. Does every call site mean the same?
  2. Syntax moves cleanly; judgement does not
  3. A script that decides is guessing
  4. Green is what a bad sweep looks like
  5. Sweep the majority, hand-edit the refusals

basics

~20 s

Mechanical rewrites are safe when the change is purely syntactic and every call site means the same thing. They turn unsafe the moment a site needs judgement: a value to choose, or an intent the old code expressed badly.

solid answer

~50 s

Ask one question: **does every call site mean the same thing?** If the transformation is a rename, a parameter reorder, or wrapping an argument, and the result is correct at every site without anyone looking at it, a scripted rewrite is both cheaper and safer than hundreds of hand edits, because it cannot get bored. It turns unsafe when the rewrite has to *decide*: pick a value the old call did not carry, choose between two replacements, or preserve an intent the old code expressed badly. Applied there, a script produces hundreds of plausible-looking cases that compile, run and assert something slightly wrong — the worst failure a suite has, because the pack stays green. In practice you split the change: sweep the mechanical majority as its own reviewable change, then hand-edit the sites the rule refused to touch.

code

pseudocode · 10 lines
pseudocode
rewrite rule "add region to open_account":
    match   call open_account(a, p)
    when    enclosing_case has tag "single-region"
    then    replace with open_account(a, p, "home")
    else    record_refusal(call.location, reason = "region not inferable")

# the sweep reports both halves, and the second half is the valuable one
sweep_report:
    rewritten: 412        # landed as one change, nothing else in it
    refused:    37        # each gets a human edit, case in front of them

go deeper

for a junior

Recall the two routes and the obvious tradeoff: a scripted rewrite is fast and uniform, a hand edit is slow and considered. Be able to say that a script does exactly what it was told, everywhere, including everywhere that is wrong.

for a middle

Explain the deciding test — whether every call site means the same thing — and give an example of each side: a rename that moves cleanly, and a new parameter whose value depends on what the case was doing.

for a senior

Show how you make a large sweep reviewable: one change with nothing else in it, review of the rule plus a sample of shapes, an explicit refusal list, and proof that a rewritten case still fails when its behaviour breaks.

for a principal

Own the risk shape. A bad hand edit is local and loud; a bad sweep is uniform and silent and leaves a pack that is green while covering less than it claims. Budget review effort against that asymmetry, not against the file count.

## The test is decidability, not scale The instinct is to choose by size — "four hundred call sites, obviously script it" — and that is the wrong axis. The right question is: **does every call site mean the same thing, such that the correct result can be computed from the call itself?** Where the answer is yes, a scripted rewrite is not merely faster than hand-editing, it is *safer*. A script applies the identical transformation at site four hundred that it applied at site one. A person will not: attention degrades, and the mistakes a tired reviewer makes are scattered and inconsistent, which is far harder to find than a uniform machine error. Where the answer is no — where the rewrite must **decide** something — the scale argument inverts. A script that decides is a script that guesses, and it guesses identically everywhere, producing hundreds of call sites that are wrong in the same plausible way. | Change | Decidable from the call? | Route | | --- | --- | --- | | Rename an action and every call to it | Yes | Sweep | | Reorder two parameters | Yes | Sweep | | Move an action to another module | Yes | Sweep | | Wrap an argument in a new container type | Yes | Sweep | | Add a parameter whose value depends on what the case does | No | Hand edit | | Split one action into two by intent | No | Hand edit | | Replace a loose assertion with a specific one | No | Hand edit | ## How a bad sweep fails, and why that is the worst failure mode A hand edit that goes wrong usually fails loudly: the case does not compile, or it fails on the next run, and someone fixes it. A sweep that goes wrong tends to leave a suite that **compiles, runs and passes** while asserting less than it did before, or exercising a configuration nobody selected. Green is exactly what a bad sweep looks like from the outside. That has a direct consequence for how you verify one. A green run after a large mechanical rewrite is weak evidence, because it mostly proves the cases still execute. The verification has to attack the diff and the coverage instead: - **Review the rule, not the output.** Read the transformation and its match conditions with real care, once. Then read a *sample* of output sites chosen to cover each distinct shape the rule matched — not four hundred near-identical hunks, which nobody reviews honestly past the fortieth. - **Read every refusal.** A rule that cannot decide should record the site and the reason rather than fall through to a default. The refusal list is the most valuable artefact the sweep produces. - **Prove a rewritten case can still fail.** Pick a few and deliberately break the behaviour they cover. If they stay green, the sweep weakened them, and you have found that in an hour instead of at the next incident. - **Land the sweep with nothing else in it.** One change, one explanation. A diff that mixes a mechanical rewrite with three behaviour fixes cannot be reviewed by either standard. ## Splitting the change In practice most migrations are not purely one route or the other; they are a large mechanical majority plus a stubborn minority. The productive shape is to **write the rule so that it refuses**: 1. Express the transformation with explicit conditions under which it applies. 2. At any site where the conditions do not hold, record a refusal with the location and the reason — never a fallback value. 3. Land the sweep of the sites the rule was confident about, as its own change. 4. Work the refusal list by hand, with the case visible, in batches. 5. Count both. If the refusal share is small, the route was correct. If it turns out that a third of the sites need judgement, that is a signal about the change itself: the transformation you described is not really one transformation, and it may be worth redefining it as two narrower rules that each sweep cleanly. That last point is the useful reframing. "Mechanical or by hand" is rarely a whole-migration decision; it is a per-rule decision, and a change that resists sweeping often does so because it bundles two different intents that the old call shape happened to spell the same way. Separating them makes both halves decidable, and the sweep becomes possible for the part that was always mechanical. ## The cost you are actually trading A hand edit costs attention per site and scales linearly with the count. A sweep costs design effort per rule and scales with the number of *shapes*, not sites — plus a fixed review cost that does not shrink just because you trust the rule. Neither is free, and choosing badly in either direction is expensive: hand-editing a decidable change wastes weeks and introduces inconsistency, while sweeping an undecidable one produces a suite that is uniformly, silently, confidently wrong.

  • The sweep compiles and the whole suite is green. Why is that not yet evidence it was correct?
    Green mostly proves the cases still execute. If the rewrite weakened an assertion or supplied a value the case never intended, green is exactly what you would see. The check has to be on the diff and the coverage: sample sites by hand, read every refusal, and deliberately break the behaviour a few rewritten cases cover to prove they still fail.
  • How do you review a diff that touches four hundred files?
    Review the rule, not the output. Read the transformation and its match conditions closely once, then read a sample of sites chosen to cover each distinct shape the rule matched, plus every site it refused. Land the sweep as its own change with no behaviour edits mixed in, so the diff has exactly one explanation.
  • What makes hand-editing hundreds of sites its own risk?
    Attention. Someone making the same edit two hundred times makes different mistakes in different places, and scattered inconsistency is much harder to find than a uniform machine error. That is why splitting the change matters: the machine does the boring uniform part, and human attention is spent only where a decision was genuinely required.

saying these in an interview costs you the question

  • Assumes a green suite proves the rewrite was correct
  • Lets the rule guess a value the old call never carried
  • Mixes the mechanical sweep and behaviour fixes in one change
  • Reviews the output line by line instead of the rule
  • Hand-edits hundreds of identical sites to feel safer
  • Chooses the route by the number of sites rather than decidability