skip to content

Which structural edits do you let a model propose, and which do you insist a deterministic tool perform?

level: principalimportance: should knowfreq 36%

answer

  1. What checks the result, not what produced it
  2. Resolution, build check, or test?
  3. Per kind of edit, not per tool
  4. Make the evidence exist first

basics

~20 s

Sort the edit by what can check it, not by how good the suggestion looks. Name resolution covers mechanical edits inside the scope it indexes, a build check catches shape, and order, defaults and failure need a test.

solid answer

~50 s

I draw the line per kind of edit, because a refactor is a claim about behaviour and the question is what checks the claim. Where a tool resolves names the way the language does, a rename or an extract carries a real guarantee about the references inside the scope it indexes, so I let the model propose it and let that tool perform it, then skim for the places a rename cannot follow. Where the same edit is textual, the guarantee is gone and I read every changed file I did not predict. Anything that merges, inlines or modernises is semantic: it needs a test pinning the edge case, written before the edit, or it does not happen yet. When that evidence does not exist, the honest options are to write the narrow test first, shrink the edit to the mechanical part, or decline it.

go deeper

for a junior

Know the difference between an edit a tool can guarantee because it resolves names the way the language does, and an edit that is a well-informed guess about text.

for a middle

Explain what each kind of evidence proves: name resolution covers references inside the scope it indexes, a build check catches shape, and order, defaults and failure need a test.

for a senior

Draw the line per kind of edit and say what you do when the evidence does not exist yet - write the pinning test first, shrink the edit to its mechanical part, or decline it.

for a principal

Own the tradeoff. A blanket refusal gives up real value on the mechanical majority, which already had a checker, while blanket trust spends the review budget on the hunks that never needed it.

## The axis is not how clever the suggestion looks The instinct is to sort structural edits by how confident the tool seems, or by how experienced the engineer accepting them is. Neither survives contact with a bad week. The axis that holds is **what can check the result** - because a refactor is a claim about behaviour, and a claim needs a checker that is not the thing that made the claim. ## Three grades of evidence, and what each actually covers 1. **Name resolution by a tool that understands the program.** An editor refactoring engine that resolves names the way the language itself does can rename or extract with a real guarantee about **references inside the scope it indexes**. That is the strongest evidence available for a structural edit, and it is also narrow: it says nothing about a name in configuration, in stored data, in message text, or assembled at run time. 2. **A type or build check.** It catches shapes that no longer fit - a call with the wrong arity, a value of the wrong kind. It runs over the whole program, which is its real value after a model edited part of it. It is silent about order of evaluation, defaults, which failure escapes, and what the caller does with a result it now ignores. 3. **Tests that pin the behaviour you claim to preserve.** Where the edit is a suggestion rather than a checked transformation, this is the only evidence that speaks to semantics at all - and only for the cases someone wrote down. A test written *before* the edit is a comparison; a test written after it is a description of whatever the edit did. | kind of edit | evidence usually available | how much review it needs | |---|---|---| | Rename, where a name-resolving tool did it | resolution guarantee, plus the build | skim the diff, then search for the name outside code | | Rename, done textually | the build, and counted occurrences | read every changed file you did not predict | | Extract or inline | tests only, and only where they exist | read the call sites; pin the edge case first | | Modernise one construct into another | tests, plus whatever the build catches | read the edges: empty input, ordering, failure behaviour | | Restructure across module boundaries | build, tests, and a design opinion | this is a design change wearing a refactor's clothes | ## The asymmetry nobody mentions until they change stack Whether a deterministic refactoring tool exists at all is a property of the **ecosystem**, not of the model. Where the editor resolves names the way the compiler does and indexes the project, a rename or an extract is a mechanical operation with a guarantee attached inside that index. Where names are resolved late, assembled dynamically, or spread across files the editor never parses, the same editor command is textual - a search and replace with good manners. Same request, same model, two entirely different risk profiles. An engineer who has only worked where the guarantee exists tends to over-trust the same edit where it does not, and the tell in an interview is someone saying "renames are safe, it is extractions you watch" without ever asking which kind of rename. ## When the evidence does not exist yet This is the interesting half of the question, because most legacy code has no test pinning the edge you are about to move. Three honest answers, in order of preference: - **Make the evidence exist.** Write the narrow characterisation test around the block first - what the rejected input does *not* cause - and accept the edit only if it still passes. This costs less than diagnosing the regression later, and the test survives the refactor as documentation. - **Shrink the edit until it is mechanical.** If the risky part is the merge of two nearly-identical blocks, do the rename part with the deterministic tool, land it, and leave the merge for a change that carries its own test. - **Decline the edit for now.** A refactor nobody can check is a bet, and taking it in code you cannot afford to break is a bad bet regardless of how good the suggestion looks. ## The tradeoff you are actually making A blanket refusal to let a model near a structural edit gives up real, boring value on the mechanical majority, and it does so in exchange for nothing, because the mechanical majority was already the part with a checker. Blanket acceptance spends the review budget on hunks that never needed it and leaves nothing for the extraction that did. The line is drawn **per kind of edit**, not per tool and not per person - which is also why it survives the tools changing. ## How you would know your line is wrong Look at where behaviour regressions actually came from over a few months. If they cluster on edits you had classified as mechanical, the classification is wrong, not the tool - most often because the rename was textual all along, or because the code resolves names at run time in a way no index can see. Move the line and keep watching; it is a calibration, not a principle.

  • No test pins the behaviour and the refactor is wanted now. What do you do?
    Write the narrow characterisation test around the block first - what the rejected input must not cause - and accept the edit only if it still passes. Failing that, shrink the edit to the part a name-resolving tool covers and leave the semantic half for a change that carries its own test.
  • How would you tell, months later, that you drew the line in the wrong place?
    Look at where behaviour regressions actually came from. If they cluster on edits you had classified as mechanical, the classification is wrong rather than the tool - usually because the rename was textual, or because names are resolved at run time where no index can see them.
  • Does a stronger model move this line?
    It moves how often the suggestion is right, not what can check it. The guarantee on a mechanical rename comes from name resolution and the guarantee on a semantic edit comes from a test, and neither is produced by the thing being checked.

saying these in an interview costs you the question

  • Decides per tool rather than per kind of edit
  • Treats a passing build as evidence that behaviour is unchanged
  • Says renames are always safe because they are mechanical
  • Writes the pinning test after accepting the edit
  • Refuses all tool-suggested restructuring rather than classifying it