skip to content

What evidence would convince you that rewriting an automation harness beats continuing to migrate it?

level: principalimportance: should knowfreq 34%

answer

  1. Measure the burndown, not the frustration
  2. Structural constraint or incidental one
  3. The cases are the asset, not the runner
  4. A rewrite restarts the trust curve
  5. Reversible only while the old suite runs

basics

~20 s

A measured burndown that will not finish before the reason for moving expires, a blocking constraint structural in the old design rather than incidental, and a plan that keeps the cases' behaviour. Dislike is not evidence.

solid answer

~50 s

Both routes are defensible, so decide on measurements you already have. Migrating incrementally keeps coverage continuous — you stay protected by whatever has moved — and costs a long overlap. A rewrite buys a design chosen for what you now know, and pays with a coverage gap and a restarted trust curve. Three numbers pick between them: the **burndown rate against the remaining count**, which says whether the migration finishes before the reason for it expires; the **share of cases needing judgement rather than a mechanical edit**, which is where migration cost actually lives; and whether the blocking constraint is **structural** in the old design — its parallelism or lifecycle model — or merely incidental. If it is incidental, keep migrating. And say what the rewrite keeps: the behaviour the cases encode is the asset, and a rewrite that discards it is a far larger bet than the one being priced.

code

yaml · 11 lines
yaml
harness_decision_record:
  remaining_cases: 610
  migrated_per_period: [41, 38, 29, 33, 30, 22, 27, 25]
  projected_finish_periods: 21
  deadline_driver: "per-case isolation needed before the daily-release target in 12 periods"
  judgement_share: 0.34            # cases that cannot move by rule
  blocking_constraint: "one process per run - structural in the old runner"
  constraint_owner: "runner"       # or: system under test
  rewrite_keeps: [business_actions, data_setup, assertions]
  retirement_bar: "6 periods matching the old suite's caught regressions"
  reversible_until: "the old suite is deleted"

go deeper

for a junior

You are not expected to make this call. Be able to say what makes it hard: the behaviour hundreds of cases encode is the expensive asset, and neither route preserves it automatically.

for a middle

Describe both routes fairly — incremental migration keeps coverage continuous and costs a long overlap, a rewrite buys a clean design and costs a gap — rather than assuming the rewrite is the grown-up choice.

for a senior

Bring numbers. A burndown rate against the remaining count, the share of cases needing judgement, and a named blocking constraint you can classify as structural or incidental are what turn an opinion into a decision.

for a principal

Own the asymmetry and write the decision down. Continuing is reversible at every batch; a rewrite is reversible only while the old suite still runs, so budget the overlap into the rewrite and state the bar the new harness must clear before anything is retired.

## First, name the asset The expensive thing a mature suite owns is not its framework. It is the **behaviour hundreds of cases encode** — the business actions, the data setup, the assertions, and the accumulated knowledge about what the system under test actually does when you poke it. A rewrite that preserves that behaviour is a machinery swap and can be argued on cost. One that discards it and starts the cases again is a much larger bet, and the first job is to establish which is being proposed. Teams routinely price the first and then do the second. ## The two defensible routes | Route | What it buys | What it costs | It wins when | | --- | --- | --- | --- | | Keep migrating incrementally | Continuous coverage — every week you are protected by whatever has already moved; reversible at any batch boundary | A long overlap: doubled run cost, two upkeep surfaces, slow visible progress | The blocking problem is incidental and the burndown projects to finish before the reason for moving expires | | Rewrite the harness | A design chosen for what you now know, unconstrained by the old model | A coverage gap, a restarted trust curve, and the exceptions the old harness accumulated arriving again one at a time | The blocking constraint is structural in the old design and cannot be extended at any migration pace | Both are legitimate. What is not legitimate is choosing between them on the strength of how the old harness feels to work in. ## Three measurements the argument needs 1. **Burndown rate against the remaining count.** Not the original estimate, which every migration overruns. Take the last several periods of migrated cases, take the count remaining, and project. Then compare that projection with the *deadline driver* — the reason the migration was started. A suite moving twenty-five cases a period with six hundred left has twenty-four periods to go; whether that is fine or fatal depends entirely on what you needed the new runtime for and when. 2. **The share of cases that need judgement rather than a mechanical edit.** Migration cost lives almost entirely here. A pack where nine cases in ten move by rule is a fundamentally cheap migration however large it is; a pack where a third of the cases need a human to decide what they meant is expensive per case, and adding people speeds it up much less than the plan assumes. 3. **Whether the blocking constraint is structural or incidental.** Write down the specific thing the old harness will not do — a parallelism model that cannot isolate what needs isolating, a lifecycle that fires in an order the suite has outgrown, a result model that cannot express what you need to report. Then ask whether it can be extended. An incidental constraint is a piece of work; a structural one is the only condition that continued migration cannot solve at any pace, and it is the strongest evidence for a rewrite there is. ## The cost of being wrong, in both directions Wrong toward rewriting is the more expensive error and the easier one to make. - **The constraints often survive.** Many of them belong to the system under test, not the runner: a deployment too slow to give each case a fresh target, no readiness signal worth waiting on, data that cannot be created cheaply. A new harness inherits every one of those on day one. Before committing, list each constraint and mark whether the runner or the product owns it. If most are owned by the product, the rewrite will not deliver what was promised. - **The second system re-accumulates the exceptions.** The old harness is complicated largely because of edge cases discovered painfully over years. Those return one bug report at a time, and they were never in the estimate. - **You pay the overlap anyway.** A responsible rewrite keeps the old suite runnable until the new one has demonstrably earned trust, which means the overlap cost you were trying to avoid appears in the rewrite budget too — and if it does not appear, the plan is to fly without a net. Wrong toward continuing has a quieter cost: a year of small edits, individually justifiable, arriving at a suite that still cannot do the thing that started the conversation. Because the spend is amortised nobody ever sees the total. ## Own the asymmetry Continuing is reversible at every batch boundary. A rewrite is reversible only while the old suite still runs — deleting it is the point of no return, and it is the only irreversible act in the whole programme. So the cheap risk control is to state, in advance, the bar the new harness must clear before anything is retired: it runs the same behaviours, and over an agreed window it catches the regressions the old suite caught with no worse a failure rate on unchanged code. Budget the overlap *into* the rewrite rather than against it. Finally, write the decision down with its numbers — remaining count, recent rate, judgement share, the named constraint and who owns it, what the rewrite keeps, and the retirement bar. Not for ceremony: a rewrite outlives the conditions that justified it, and in month seven someone will need to know whether the argument still holds.

  • The team is confident the rewrite takes eight weeks. How much weight should that estimate carry?
    Little on its own. A rewrite estimate covers the parts already understood and systematically omits the accumulated exceptions — the odd wait, the one data setup needing a real record, the reporting detail somebody depends on — which is most of what made the old harness complicated. Weight the estimate by how much of that complexity the team can currently explain, and treat the unexplained remainder as work not yet estimated.
  • What bar should a new harness clear before the old one is retired?
    A measurable one agreed in advance: it runs the same behaviours, and over an agreed window it catches the regressions the old suite caught, with no worse a failure rate on unchanged code. Until that window closes the old suite stays runnable, because that is the only thing keeping the decision reversible.
  • Why do the constraints that motivated a rewrite often survive it?
    Because many belong to the system under test rather than the runner: a deployment too slow to give each case a fresh target, no readiness signal worth waiting on, data that cannot be created cheaply. A new harness inherits all of those on day one. Before committing, list each constraint and say whether the runner or the product owns it.

Replacing the engine while the vehicle is in service versus building a second vehicle. The first is slow and always drivable; the second is faster on paper and leaves you walking during the handover.

saying these in an interview costs you the question

  • Argues for the rewrite from frustration with the old design
  • Prices the rewrite without the overlap it still needs
  • Assumes the new harness will not re-accumulate the same exceptions
  • Treats the behaviour the cases encode as something to write again
  • Never measures the burndown before declaring the migration stuck
  • Deletes the old suite the day the new one first goes green