How do you execute a refactoring far too large for one commit, across a team sharing one mainline, while keeping the build green and the system releasable at all times? Address the Mikado Method and the parallel-change (expand-contract) technique.
answer
- Never a long-lived refactoring branch
- Mikado: try → note breakage → REVERT → do prerequisites → commit leaves
- Graph of prerequisites is a shareable artefact
- Parallel change: expand → migrate → contract
- Contract phase needs a date, an owner and usage telemetry
basics
~20 sBreak it into many small commits that each keep the build green. Use the Mikado Method to map prerequisites by trying the change, noting what breaks, reverting, and doing the prerequisites first. Use parallel change — add the new form, migrate callers, remove the old — so old and new coexist while others keep working.
solid answer
~60 sLarge refactorings must be delivered as a stream of small, green, individually shippable commits on mainline — never as a long-lived branch, whose merge reintroduces the big-bang risk you were avoiding. **Mikado Method**: attempt the change naively; when it breaks, write the breakage down as a prerequisite node, **revert to green**, and recurse into the prerequisites. You end with a dependency graph whose leaves are safe, independent changes. Work leaves inwards, committing each. Reverting rather than pushing through is the counter-intuitive core — it keeps the tree green while you learn. **Parallel change / expand–contract**: *expand* (add the new API, column, or event alongside the old, both working), *migrate* (move callers incrementally, possibly across releases and repos), *contract* (delete the old once nothing uses it). Combined with branch by abstraction and feature toggles, this lets independent teams and deployables migrate at their own pace. Operationally: automate with codemods where possible, keep each PR mechanically reviewable, publish the graph so others can help, and put a dated deletion milestone on the contract phase — half-finished migrations are the real cost.
code
pseudocode · 16 lines// PARALLEL CHANGE (expand -> migrate -> contract) on a public signature
// 1. EXPAND — new form added; old form kept, delegating. Nothing breaks.
function chargeCustomer(id, amountMinorUnits, currency) { /* new, real logic */ }
@Deprecated("use chargeCustomer(id, amountMinorUnits, currency); removal: 2026-11-01")
function chargeCustomer(id, amountInDollars) {
metrics.count("deprecated.chargeCustomer.legacy") // telemetry gates the contract phase
return chargeCustomer(id, round(amountInDollars * 100), "USD")
}
// 2. MIGRATE — move call sites incrementally (codemod where mechanical),
// one green commit at a time; external clients migrate over releases.
// 3. CONTRACT — when the deprecated counter has read zero for a full retention window,
// delete the old overload, the shim and the metric. Deletion is the goal.go deeper
Say you would break the work into many small changes that each keep the tests passing and get committed, rather than doing it all at once on a branch.
Add the expand–migrate–contract shape (add the new thing, move callers, delete the old) and why coexistence lets other people keep working.
Explain the Mikado loop including the revert-to-green step and the prerequisite graph, combine it with branch by abstraction and toggles, and address deprecation telemetry gating the deletion phase.
Treat it as a programme-design and organisational problem: continuous delivery of value in leaves rather than a freeze, legible progress metrics so sponsorship survives, codemod-generated mechanical commits separated from semantic ones for reviewability, cross-team/cross-repo deprecation with timelines and usage telemetry, explicit budget and ownership for the contract phase, and the property that pausing the effort loses nothing because every committed step stands alone.
## The constraint A refactoring that touches hundreds of files, crosses module or repository boundaries, or must be coordinated across teams cannot be one commit. But the discipline that makes refactoring safe — small steps, green suite, cheap revert — must survive the scale-up. Three properties must hold throughout: 1. **Mainline stays green and releasable** — other teams keep shipping. 2. **Every step is independently revertible** — no step is a point of no return. 3. **Work is visible and shareable** — colleagues can contribute, and the effort survives a person leaving. The anti-pattern is the **long-lived refactoring branch**: it diverges as others commit, its merge is a single high-risk event, its work is invisible until then, and it is frequently abandoned. Its cost grows superlinearly with duration. ## The Mikado Method From Ola Ellnestam and Daniel Brolund. Named after the pick-up-sticks game where the goal stick can only be taken once the sticks on top of it are removed. **Loop** 1. State the **goal** ("replace the home-grown DI container with the framework's"). 2. **Try it naively** — make the change directly. 3. **Observe the errors** (compile failures, red tests). Each is a **prerequisite**. 4. **Write the prerequisites into a graph** as child nodes of the goal. 5. **Revert to green** — throw the attempt away. This is the crucial, counter-intuitive step: your attempt was an *experiment to learn dependencies*, not work to preserve. 6. Pick a prerequisite and recurse from step 2 until you find one you can complete without breaking anything — a **leaf**. 7. **Do the leaf, verify green, commit, push.** Mark it done. 8. Re-attempt its parent; the graph shrinks upward until the goal itself becomes a leaf. **Why revert instead of pushing through?** Because pushing through leaves the tree red and the working copy large and uncommittable, which is exactly the state where refactorings die. Reverting costs minutes of typing and buys a known-good baseline plus a map. It also produces an artefact — the graph — that can be shared, estimated, and worked on in parallel by several people. **Trade-offs**: repeated exploratory work feels wasteful; graphs on very large goals can sprawl; it works best when feedback (compile/test) is fast, and poorly in dynamic environments where breakage only surfaces at runtime — there, lean on characterization tests, static analysis and staged rollout instead. ## Parallel change (expand–contract) Danilo Sato's *parallel change*, also called expand–contract or expand–migrate–contract. Structure a breaking change as three separately deployable phases: 1. **Expand** — introduce the new form *alongside* the old. New method next to the old one; new database column next to the old one, dual-written; new event schema published in parallel; new interface added while the old delegates to it. After expand, both forms work and nothing has broken. 2. **Migrate** — move consumers to the new form, incrementally. Within a repo this can be a codemod; across services or client apps it can take releases or months, and you may need per-consumer tracking. Instrument the old form (deprecation logging, metrics) so you can *see* remaining usage rather than guess. 3. **Contract** — once usage is provably zero, delete the old form, the dual-write, and any compatibility shims. This is the only sane way to change a contract that independently-deployed clients depend on, because at no point do producer and consumer have to deploy simultaneously. For **data**, the same shape appears as expand–migrate–contract migrations: add nullable column → backfill → dual-write/dual-read → switch reads → stop writing old → drop column. Each phase is independently deployable and revertible; dropping a column is never in the same release as the code change that stopped using it. ## Complementary techniques - **Branch by abstraction** — introduce an interface over the thing being replaced, migrate callers, build the new implementation behind it, switch by toggle, delete the old. Keeps the replacement on mainline. - **Feature toggles** with per-tenant/percentage rollout and a tested kill switch, so behaviorally risky switches roll back in seconds. Track toggles as debt with removal dates. - **Codemods / automated refactorings** — for mechanical, repository-wide edits (rename, signature change, import rewrite), machine-generated diffs are safer and more reviewable than hand edits. Reviewers verify the *transformation*, and the tests verify the result. - **Preparatory refactoring** — "make the change easy, then make the easy change"; often the first several Mikado leaves. - **Parallel run / shadow traffic + contract tests** — evidence of equivalence when the suite alone is not convincing. - **Strangler fig** — the same philosophy applied at the system boundary. ## Making it survive contact with an organisation - **Don't stop the world.** Big-bang "code freeze while we refactor" programmes lose sponsorship. Ship continuously in leaves. - **Make progress legible.** Publish the Mikado graph or a checklist with counts ("call sites remaining: 412 → 96"). Progress that cannot be seen gets cancelled. - **Budget the contract phase.** The commonest organisational failure is stopping after migrate: both forms live forever, complexity has been *added* not removed. Put dates and owners on deletion, and use usage telemetry as the gate. - **Keep PRs reviewable.** Separate mechanical commits (huge, boring, ideally codemod-generated, with the transformation described) from semantic ones (small, scrutinised). Never mix behavior changes in. - **Coordinate deprecation with consumers** — deprecation warnings, migration guides, timelines, and telemetry showing who still calls the old form. - **Expect and plan for interruption.** Each committed leaf is standalone value; if the effort is paused for a quarter, nothing is lost and no branch rots.
- Why does the Mikado Method insist you revert your exploratory attempt instead of fixing forward?Because the attempt's value is the *information* about prerequisites, not the code. Fixing forward leaves you with a red build and a large uncommittable working copy — the state in which large refactorings stall and get abandoned. Reverting restores a known-good baseline, costs minutes, and yields a dependency graph you can share, estimate and parallelise across people.
- How do you change a method signature that external clients you don't control depend on?Parallel change. Expand: add the new signature and keep the old one delegating to it, marked deprecated with a removal date and usage telemetry. Migrate: publish a guide and let clients move over releases, watching the counter. Contract: delete only after usage has been provably zero for a full retention window. At no point must producer and consumer deploy simultaneously.
- What is the most common way large refactoring programmes fail organisationally?They stop after the migrate phase. Both old and new forms remain, the compatibility shims stay, the toggles stay, and the codebase is now more complex than before the effort started. Prevent it by treating deletion as the definition of done: dated milestones, named owners, usage telemetry as the gate, and visible metrics for remaining call sites or legacy traffic.
- How should reviewers handle the enormous mechanical diffs these efforts produce?Split commits by nature: mechanical (thousands of lines, ideally generated by a codemod or IDE refactoring, with the exact transformation stated in the message) and semantic (small, hand-written, heavily reviewed). Reviewers verify the transformation rule and spot-check, rather than reading every hunk; the test suite verifies the result. Behavior changes must never ride along in the mechanical commit.
Mikado is the pick-up-sticks game: you can only lift the stick you want once everything resting on it is gone. Trying and putting it back — rather than yanking — is how you learn what is on top without collapsing the pile.
saying these in an interview costs you the question
- Doing a large refactoring on a long-lived branch and merging at the end
- Pushing forward through breakage instead of reverting to green and mapping prerequisites
- Skipping the contract phase, so old and new forms coexist permanently
- Deleting a deprecated API based on a code search rather than production usage telemetry
- Mixing behavior changes into mechanical, repository-wide commits so reviewers cannot separate risk
- Demanding a delivery freeze so the team can "refactor properly"
- Progress that is invisible to sponsors, so the effort is cancelled midway leaving the system half-migrated