When is hiding unfinished work behind a runtime switch better than hiding it on a branch?
answer
- Same problem, two hiding places
- Merge debt traded for flag debt
- Integrated in history, off at run time
- Each switch adds states to test
- A switch needs an owner and expiry
basics
~20 sA runtime switch moves the hiding place from version control into the running system: code integrates daily while the behaviour stays off. It trades merge debt for flag debt, meaning more live paths, more states to test, and switches nobody removes.
solid answer
~50 sA long-lived branch and a runtime switch answer the same question: how do you keep work in progress away from users? The branch hides it in history and pays at integration; the switch hides it at run time and pays continuously. Prefer the switch when the work spans weeks, touches code other people are also changing, or has to be exercised in the real environment before anyone sees it. It keeps the change integrated daily and makes exposure reversible in seconds rather than requiring a redeploy. The costs are real: each switch adds states the system can be in, both paths need the tests that matter, and half-built behaviour sits in production behind a false condition. Treat a switch as debt with a name: an owner, an expected removal, and removal as the last step of the work.
code
pseudocode · 7 linesif switch_on(renewals-price-rule-v2) then
price = new_renewal_price(policy) // unfinished, off in production
else
price = current_renewal_price(policy) // the shipped behaviour
end
// both paths are deployed; only the condition decides who sees whichgo deeper
Know that unfinished work can be hidden either on a separate line of development or behind a condition checked at run time, and that the switch lets the code be integrated while the behaviour stays off.
Explain the mechanics: where the condition is evaluated, why both paths need tests, and why a switch left in place after its behaviour ships becomes code every future reader has to reason about.
Show judgement about which work earns a switch, how you would validate the disabled path in production, and how you have actually removed switches as part of the work rather than as a promised follow-up.
Own the policy: how many live switches the organisation tolerates, who owns expiry, how switches interact with release obligations and audit, and when the honest answer is a short-lived branch instead of a switch.
## Two hiding places for the same problem Work in progress must not reach users before it is ready, and there are exactly two places to hide it: in **version control**, on a line that has not been merged, or in the **running system**, behind a condition that evaluates to off. Everything about the choice follows from where the hiding happens, and therefore from when the cost is paid. | | Hidden on a branch | Hidden behind a runtime switch | |---|---|---| | Where the work sits | Unmerged history | Merged, shipped, inert | | When you pay | At integration, all at once | Continuously, in paths and states | | Exposure decision | A merge plus a release | A configuration change, in seconds | | Reversal | Undo the change and redeploy | Turn the switch off | | Debt it creates | Merge debt | Flag debt | ## When the switch is the better trade Reach for the switch when at least one of these holds: - **The work spans weeks.** Divergence compounds; a switch converts weeks of merge debt into daily integration. - **It touches code other people are changing.** Integrating daily means their work and yours meet while both are still small. - **It must be validated where the real data is.** Some behaviour can only be trusted after it has run against production traffic, for a fraction of it, before anyone sees the result. - **Exposure and deployment must be separable.** A car-insurance renewals team on a six-week cadence, with a contractual penalty attached to a delivery date, cannot afford a fix that waits for the next release window. Being able to turn a half-finished renewal-price rule off 40 seconds after noticing it, rather than in five and a half weeks, is what makes that date survivable. ## What the switch actually costs The classic failure is treating the switch as free: 1. **State explosion.** Each independent switch multiplies the configurations the system can be in. Ten switches are a thousand nominal combinations, of which perhaps five matter, and somebody has to decide which five and test them. 2. **Both paths ship.** Half-built code is in production, readable by everyone, and reachable if the condition is ever wrong. That includes its dependencies and whatever surface it exposes. 3. **Configuration becomes a failure surface.** A wrong switch value is now an incident with no code change behind it, which makes it easy to cause and easy to miss when looking for what changed. 4. **Switches outlive their reason.** The reference failure is the switch whose behaviour shipped fully two years ago and which is still evaluated everywhere, because removing it is nobody's task and looks like pure risk. That renewals team measured 23 live switches, nine of them older than a year and two whose owners had left. None was doing anything except making every change in that area harder to reason about. ## Keeping flag debt bounded - **Name an owner and an expectation at creation.** A switch with no owner is permanent by default. - **Remove it as the last step of the work, not as a follow-up.** A cleanup ticket is a ticket that loses to the next feature, every time. - **Distinguish two species.** A release switch exists to decouple deployment from exposure and should live for weeks. An operational control, such as a kill switch on an expensive path, is meant to be permanent and should be reviewed as configuration rather than tracked as debt. - **Test the paths that matter.** At minimum the shipped-on path, the shipped-off path, and any pair of switches known to interact. - **Cap the live count.** A team that will not remove switches faster than it adds them has chosen the debt deliberately, whatever it says. ## Where the switch is the wrong answer A runtime switch selects which code runs. It does not change **data that has already been written**, so a change to the shape of stored records cannot be hidden behind a condition alone. That needs the expand, migrate, contract sequence, where both shapes stay readable while the migration runs. The same limit applies to a contract that other systems already call: you can switch your own path, not theirs. And a switch is not permission to skip anything. Code behind a switch is reviewed, tested and deployed like every other line, because it is deployed. A team that uses switches as a way to land unreviewed work does not have a branching strategy; it has a way of losing the gate. ## The honest summary Neither hiding place is free. They differ in when the bill arrives and whether you can pay it in small pieces. A branch defers a large, unpredictable payment to the worst possible moment, near a release date. A switch charges a small predictable rent, and what an interviewer is really testing is whether you know that the rent has to be cancelled.
- How do you stop runtime switches accumulating until nobody knows which ones are safe to remove?Give every switch an owner and an expected lifetime when it is created, and make removal the final step of the work that introduced it rather than a follow-up ticket. Keep an inventory with ages, review it on a cadence, and cap how many may be live at once. Separate the short-lived release switches, which must expire, from deliberate operational controls, which are configuration and stay.
- What kind of unfinished work cannot be hidden behind a runtime switch?Anything that changes state outside the code path. A switch chooses which code runs; it cannot alter records already written in the old shape, so a data-shape change needs an expand, migrate and contract sequence with both shapes readable during the transition. The same applies to a published contract other systems already call: you can switch your side, not theirs.
A branch is an extension you have not built; a switch is one you have built with the door locked. You can open it in seconds, but you are heating a room nobody uses.
saying these in an interview costs you the question
- Treats a runtime switch as free once it is added
- Never removes switches after the behaviour is fully released
- Thinks a switch can hide a change to stored data shape
- Tests only the enabled path and never the disabled one
- Uses a switch as an excuse to skip review or tests