skip to content

A feature has served 100% of traffic behind a permanently-enabled flag for six months, and your CI suite only ever runs with that flag on. Why is that a reliability problem, and how would you cover both states without letting the test matrix explode?

level: middleimportance: should knowfreq 48%

answer

  1. the off branch is your rollback path
  2. no traffic, no tests, six months
  3. two to the n combinations
  4. pin defaults, parameterize the movers
  5. nightly run with flags inverted

basics

~20 s

The off path is your rollback path, and untested code that has served no traffic for six months has probably rotted. Test both states only for flags that could realistically be flipped, rather than every combination of every flag.

solid answer

~50 s

The flag exists so you can turn the feature off in an incident, so the off branch is your mitigation path — and a branch with no test coverage and no production traffic for six months is not a mitigation path, it is a second untested code path you will discover at 3am. Meanwhile testing every combination is impossible: n independent flags give 2 to the n states, so twenty flags is over a million. The way out is to scope it. Pin long-lived flags to their production value as the default test configuration, and run the two states explicitly only for flags that are mid-rollout or that are declared kill switches. Add a scheduled job that runs the suite with those flags inverted, so rot in the fallback is found by a nightly failure rather than during an outage.

code

javascript · 14 lines
javascript
const { test } = require('node:test');
const assert = require('node:assert');

function total(items, flags) {
  return flags['new-pricing'] ? items.length * 10 : items.length * 12;
}

// Both reachable states of one mid-rollout flag, not every combination.
for (const enabled of [true, false]) {
  test(`total is serviceable with new-pricing=${enabled}`, () => {
    const result = total(['a', 'b'], { 'new-pricing': enabled });
    assert.ok(result > 0, 'both branches must return a usable total');
  });
}

go deeper

for a junior

Know that a flag creates two code paths and that only one of them is getting exercised, so the disabled branch can quietly stop working without anyone noticing.

for a middle

Explain the combinatorial limit — n flags give 2 to the n states — and describe a concrete scheme: production-matching defaults, parameterized tests for movers, a scheduled inverted run.

for a senior

Frame it as rollback readiness: the untested off path is the mitigation you are counting on, so name the rot modes you have actually seen and how you detect them before an incident does.

for a principal

Set the policy that keeps this bounded — flag expiry with an owner, a cap on permanent flags, and generated test defaults — and be able to defend deleting flags over carrying more coverage.

## Why the off path matters more than the on path A feature flag is only a control because of what happens when you turn it off. The on path gets exercised by every production request; it is the best-tested code you own. The off path gets exercised by nothing — no users, and, in the situation described, no tests either. That asymmetry is the whole problem: the branch you will rely on in the worst possible moment is the branch with the least evidence behind it. Six months is long enough for real rot. Typical ways the fallback stops working while nobody notices: - The old path reads a column or config key that was cleaned up after the migration "finished". - A shared helper was refactored for the new path's needs and the old caller was updated only enough to compile. - The old path's downstream dependency was decommissioned, or its client credentials expired. - The old rendering path lost a template, an asset or a translation key. - Serialization moved on: the old reader cannot parse rows the new writer has been producing for six months. Every one of those is silent until the moment you flip the switch during an incident — turning a one-symptom outage into two. ## Why you cannot just test everything The naive fix is "run the whole suite in both states for every flag". That does not scale, and saying so is part of a good answer. With n independent boolean flags there are 2 to the n combinations: 10 flags is 1,024, 20 flags is over a million. Even running each flag's two states independently doubles suite wall-clock time per flag, which turns a 10-minute pipeline into an hour and then gets switched off by the first team that misses a deadline. So the real question is not "how do I test all states" but "which states are reachable in production, and which of those do I actually intend to use?" ## A workable policy **Classify the flag by whether it can still move.** A flag that is at 100 percent, has an owner, and is scheduled for removal is on a one-way trip; it does not need permanent dual coverage, it needs deleting. A flag that is mid-rollout, or one that is declared an operational kill switch and will live forever, genuinely has two reachable states and earns coverage of both. **Make the default test configuration mirror production.** Tests should run with flags set to their current production values, so CI is testing the system users actually get. Divergence between the test defaults and the production values is a common and confusing source of "passes in CI, fails in prod". **Parameterize the small set that matters.** For the flags you kept, run their two states explicitly at the level where the branch lives — usually a handful of integration or contract tests rather than the entire suite. ```javascript for (const enabled of [true, false]) { test(`checkout total with new-pricing=${enabled}`, async () => { const app = buildApp({ flags: { 'new-pricing': enabled } }); const res = await app.post('/checkout', { items: ['a', 'b'] }); assert.equal(res.status, 200); }); } ``` The point is not that both branches produce identical output — often they must not — but that both produce a *correct, serviceable* result. **Run the inverse on a schedule.** For long-lived kill switches, a nightly or weekly job that runs the suite with those flags flipped catches rot within days, at a cost that never touches the pull-request pipeline. A red nightly is a cheap warning; a dead fallback in an incident is not. **Close the loop with production evidence.** Tests prove the code compiles and behaves in isolation; a periodic drill of the switch in production proves the whole path works. The two are complementary, and neither substitutes for the other. ## Deleting is also an answer For a flag that has been at 100 percent for six months with no plan to ever flip it, the correct move is usually removal: delete the losing branch, delete the conditional, delete the flag record. That collapses two code paths into one and removes the test-matrix entry entirely. Keeping a flag "just in case" while never testing its off state is the worst of both worlds — you carry the branch, the complexity and the false sense of safety, and you get none of the control. ## What interviewers listen for The strong answer names the asymmetry (the fallback is untested precisely because it is the fallback), gives the combinatorial reason you cannot brute-force it, and then makes a *choice*: which flags earn dual coverage, which get deleted, and where the inverse run happens. An answer that stops at "we should test both states" has not engaged with the cost.

  • Which flags would you say do not deserve permanent coverage of both states?
    Release flags that have reached 100 percent, have an owner and a removal date, and that nobody intends to flip back. Those should be deleted rather than tested forever. Permanent coverage is for flags with two genuinely reachable states: anything mid-rollout, and any declared kill switch or degradation lever that is expected to live for the life of the service.
  • How do you keep the test configuration from drifting away from production flag values?
    Generate the test defaults from the production flag state rather than hand-maintaining a second list, and fail the build when a flag exists in one place and not the other. If generation is impractical, run a scheduled comparison that reports flags whose test default differs from their production value, so drift is a visible number instead of a surprise.
  • Does covering both states in CI mean you can skip drilling the switch in production?
    No. CI proves the branch compiles and behaves against test doubles; it says nothing about production configuration, real dependencies, propagation time, or whether the people who need to flip it can. Tests and drills fail differently, which is precisely why you want both.

saying these in an interview costs you the question

  • Both branches must produce identical output to pass
  • Test every combination of every flag
  • The off path is safe because it used to work
  • A flag at 100 percent for months is still real insurance
  • Production traffic counts as coverage for the disabled branch

context