skip to content

How should the reversibility of an architectural decision change how much effort you spend choosing, and how do you keep a style choice evolvable over time?

level: principalimportance: nice to knowfreq 34%

answer

  1. two-way door = decide fast; one-way door = analyse
  2. reduce irreversibility, don't just predict better
  3. last responsible moment, not procrastination
  4. strangler fig / branch-by-abstraction / expand-contract
  5. fitness functions + revisit triggers in the ADR

basics

~20 s

Cheap-to-reverse decisions should be made fast and tried; expensive ones deserve real analysis. Keep choices evolvable by deferring the costly ones until you know more, keeping clean enforced boundaries, and migrating incrementally rather than rewriting.

solid answer

~50 s

Classify decisions by reversibility — 'two-way doors' versus 'one-way doors'. Two-way doors (a library, a cache, a queue technology) should be decided quickly by whoever is closest to the work, since reversing costs days. One-way doors — splitting data across services, a public API contract, a consistency model, a datastore with no migration path — justify scenario evaluation, prototyping and an ADR, because reversal costs quarters. Apply the last responsible moment: defer a costly decision until deferring further would itself remove options, and meanwhile keep options open with abstraction at the volatile boundary only. Evolvability comes from three habits: enforced modular boundaries (architecture tests, ownership of schemas per module) so seams exist when needed; fitness functions so drift is detected rather than discovered; and incremental migration — strangler fig with a routing facade, branch by abstraction, expand/contract schema changes — instead of big-bang rewrites. Finally, write the revisit trigger into the ADR so the decision reopens on evidence, not on whoever complains loudest.

go deeper

for a junior

Say that easily changed decisions should be made quickly and hard-to-change ones need more thought, and that big rewrites are risky compared with changing a system piece by piece.

for a middle

Use the reversible versus irreversible distinction with examples, and name at least one incremental migration technique such as the strangler fig pattern.

for a senior

Add the last responsible moment, enforced modular seams, fitness functions, and expand/contract plus branch by abstraction; explain why boundaries without enforcement erode.

for a principal

Frame it as option value and risk: deliberately convert one-way doors into two-way doors, size analysis by reversal cost times probability of error, define evidence-based revisit triggers, and guard against sunk-cost and rewrite traps at the portfolio level.

### Reversibility as the effort dial Not all architecture decisions deserve equal ceremony. A widely used framing distinguishes **two-way doors**, which can be walked back cheaply, from **one-way doors**, which cannot. A related point often attributed to Martin Fowler is that the architect's job is largely to *reduce irreversibility* — to convert one-way doors into two-way doors where possible. Examples of two-way doors: a logging library, a cache technology, internal code layout, a CI tool. Examples of one-way doors: splitting one datastore into several owned by separate services; publishing an external API contract that third parties depend on; choosing eventual over strong consistency in a domain handling money; a datastore with no realistic migration path; and the module boundaries themselves, once teams and data have grown around them. The operational rule: spend evaluation effort proportional to (cost of reversal) times (probability of being wrong). Analysing a two-way door for three weeks is waste; prototyping it for two days is not. ### Last responsible moment From lean thinking: delay a decision until the moment beyond which delaying further would eliminate an important option. The point is not procrastination — it is that information accumulates. A month into building you know real load shapes, real domain boundaries and real team capacity, all of which improve the decision. What makes deferral safe is keeping the decision *isolated*: hide the volatile choice behind one boundary so changing it later touches a small area. The failure mode is over-applying this into speculative generality — abstracting everything 'just in case', which adds indirection now for flexibility that is usually never used. Abstract at the boundary you have concrete evidence is volatile; elsewhere, commit. ### Keeping the choice evolvable 1. **Enforced boundaries.** Modularity is what makes later change affordable: one module per capability, an explicit public surface per module, no cross-module table access, and automated architecture tests that fail the build on a violation. Boundaries that are documented but not enforced erode within months. 2. **Fitness functions.** Automated checks that the chosen characteristics still hold — latency budgets asserted in load tests, dependency rules asserted in architecture tests, failover time asserted by a chaos experiment, budgets on bundle size or startup time. This is the core of *evolutionary architecture*: guided change, where the guides are executable. 3. **Incremental migration.** When the style must change, avoid the big-bang rewrite, which competes with a moving target and delivers nothing until the end. Instead use the **strangler fig** — put a facade in front, route slice by slice to the new implementation, delete the old slice when its traffic reaches zero; **branch by abstraction** — introduce an abstraction over the old implementation, build the new one behind it, switch by configuration; and **expand/contract (parallel change)** for schemas and APIs — add the new shape, write to both, migrate readers, then remove the old shape. Add dark launches and shadow traffic to validate the new path under real load before it serves users. 4. **Revisit triggers.** State in the ADR the observable condition that reopens the decision: a scale threshold, a team-count threshold, a cost ceiling, a regulatory change, a latency budget consistently breached. This makes evolution evidence-driven rather than politics-driven. ### Traps at this level The sunk-cost trap: continuing with a style because of what has already been spent rather than the cost from here forward. The rewrite trap: assuming a clean slate is cheaper than incremental change — it rarely is, because the old system encodes years of undocumented edge cases. The premature-distribution trap: paying microservice costs for optionality you never exercise. And the frozen-architecture trap: no fitness functions, so the system quietly drifts until the documented architecture and the real one have nothing in common.

  • Give a concrete example of converting a one-way door into a two-way door.
    Instead of splitting a database across services immediately, keep one datastore but give each module its own schema with no cross-schema queries, and route all access through the module's interface. The expensive irreversible step — physical data separation — is deferred, while the seam that makes it possible is already enforced.
  • Why is a big-bang rewrite usually worse than incremental migration?
    The old system keeps evolving while you rebuild, so you chase a moving target; nothing ships until the end, so feedback and value are deferred and the project is politically fragile; and the legacy code encodes years of undocumented edge cases that surface only at cutover. Strangler fig delivers value continuously and keeps rollback cheap per slice.

Renting versus buying a home in a new city: rent (two-way door) while you learn the neighbourhoods, and only buy (one-way door) once the evidence is in — meanwhile keeping your belongings packable so moving stays cheap.

saying these in an interview costs you the question

  • Applying the same heavyweight analysis to every decision regardless of reversal cost
  • Confusing 'last responsible moment' with deferring decisions until the options are gone
  • Speculative generality — abstracting everywhere 'for flexibility' that is never used
  • Justifying continuation of a failing architecture by how much has already been invested
  • Documenting boundaries without automated enforcement, then being surprised they eroded
  • Assuming a rewrite will be faster and cleaner than incremental migration

context