skip to content

Feature Flags as Operational Controls

Decoupling deploy from release so risk can be dialed up and down at runtime. Interviewers probe flags as incident levers — kill switches and gradual ramps — plus the flag-debt hygiene that keeps them trustworthy.

on this pageshow

questions

5

Your service can only disable a bad new code path by redeploying the previous build, which takes about 25 minutes end to end. What does putting that code path behind a runtime feature flag change about your recovery, and what has to be true of the flag for that improvement to be real?

level: juniorimportance: must knowfreq 70%

answer

  1. deploy is not release
  2. the branch already shipped
  3. config change, not a rebuild
  4. evaluated per request, not at boot
  5. the off path must still work

basics

~20 s

A runtime feature flag turns recovery from a redeploy into a configuration change, cutting mitigation from tens of minutes to seconds. That only holds if the flag is evaluated per request, the old path still works, and the change reaches every instance quickly.

solid answer

~50 s

Behind a flag the new code is already in production, so recovery stops being a rebuild-promote-redeploy cycle and becomes a config change that propagates in seconds. That is the deploy/release split: deploying puts bytes on servers, releasing turns behaviour on for users, and only the second one needs to be reversible in seconds. For the 25 minutes to actually become seconds, three things must hold. The flag has to be evaluated at request time, not read once into a constant at boot, or you are back to restarting processes. The old code path has to still be present and still work — a flag does not test the path it falls back to. And the change has to reach every instance quickly and visibly, so you can confirm the flip took effect instead of assuming it did. If any of those is false, you still own the redeploy.

code

javascript · 20 lines
javascript
// Build-time constant: changing it requires a new deploy.
const NEW_PRICING = true;

function priceCartCompiled(cart) {
  return NEW_PRICING ? newPricing(cart) : legacyPricing(cart);
}

// Runtime flag: changing it is a config write, effective on the next request.
function priceCartFlagged(cart, flags, userId) {
  if (flags.isEnabled('new-pricing', { userId })) {
    return newPricing(cart);
  }
  return legacyPricing(cart);
}

function newPricing(cart) { return cart.items.length * 10; }
function legacyPricing(cart) { return cart.items.length * 12; }

console.log(priceCartCompiled({ items: [1, 2] }));
console.log(priceCartFlagged({ items: [1, 2] }, { isEnabled: () => false }, 'u1'));

go deeper

for a junior

Be able to say plainly that the flagged code is already deployed and the flag only chooses which branch runs, so turning a feature off is a config change rather than another deploy.

for a middle

Explain the mechanics: request-time evaluation versus a boot-time constant, how the value reaches instances by push or poll, and roughly how long that propagation takes.

for a senior

Show you know the limits — an untested off path, side effects already written, and flag changes bypassing every pipeline gate — and say what you put in place to cover each.

for a principal

Own the tradeoff of which changes earn a flag at all. Each flag is a permanent branch and a test-matrix entry, so argue the criteria by expected incidents avoided rather than flagging everything.

## The problem the flag is solving When the only way to withdraw a change is to redeploy the previous build, your mitigation time is bounded by your delivery pipeline: rebuild or re-promote the artifact, roll it across instances, wait for health checks and connection draining. Twenty-five minutes is an ordinary figure for that, and during an incident every one of those minutes is user impact and error-budget spend. A runtime feature flag attacks the problem from a different angle. Instead of changing *which code is running*, it changes *which branch that code takes*. The new path and the old path both ship in the same artifact; a value read at request time decides which one executes. ```javascript // build-time constant: changing it means a redeploy const NEW_PRICING = true; // runtime flag: changing it means a config write if (flags.isEnabled('new-pricing', { userId })) { return newPricing(cart); } return legacyPricing(cart); ``` The first form is not an operational control at all — it is a compile-time decision wearing a flag's clothes. ## Deploy is not release The vocabulary matters in interviews. **Deploy** is the act of getting an artifact onto production machines. **Release** is the act of exposing a behaviour to users. Without flags the two are the same event, so the blast radius of a deploy is the blast radius of every change inside it, and the undo for both is the same slow mechanism. With flags they separate: you can deploy at any cadence, keep the new behaviour dark, and choose the moment — and the audience — for the release independently. Crucially, the *undo* for a release becomes a different, much cheaper operation than the undo for a deploy. This is why flags are treated as a release-safety practice and not merely a coding convenience. They give the on-call engineer a lever that can be pulled during an incident by someone who does not need to know how the build system works. ## The three conditions **Evaluated at request time.** If the flag value is read once at process start into a module-level constant or a cached singleton that never refreshes, flipping it changes nothing until the processes restart — and restarting every instance is a redeploy by another name. Real control means the value is consulted on the code path each time, or at least refreshed on a short interval that you know and can quote. **The off path still works.** A flag's promise is that turning it off returns you to the previous behaviour. That is only true if the previous behaviour is still exercised and still correct. Once a feature has run at 100 percent for months, the off branch has had no traffic, and it can rot silently — it may reference a column that was dropped, or produce a page that no longer renders. The flag looks like insurance and turns out to be a second, untested failure mode. **Propagation is fast and observable.** Flag changes usually reach instances by push (a streaming connection from the flag platform) or by poll (each instance re-fetching every few seconds). Either way there is a real number — a few seconds, sometimes tens of seconds — and you should know it, because during an incident the gap between "I flipped it" and "it took effect" is the window where people start making a second change on top of the first. You also want a signal that confirms the flip landed: the error rate turning over, or a metric that reports the effective flag value per instance. ## What the flag does not give you A flag does not make the change safe to ship untested; it makes it cheap to withdraw. It does not undo side effects that already happened — rows written in the new format, messages published to a queue, emails sent — so a flag guarding a write path is a weaker control than one guarding a read or render path, and you should say so when asked. It does not replace rollback either: a bad change that is not behind a flag, or a bad change in shared code that both branches run, still needs the deploy-level undo. And every flag has a cost. It is a branch in the code, a state in the test matrix, and an entry in an inventory somebody has to prune. The reason experienced teams still take that cost for risky changes is arithmetic: minutes of user impact avoided per incident, multiplied by the number of incidents where the lever gets pulled. ## How to answer it in an interview Give the number first — redeploy in tens of minutes versus a config change in seconds — then name the deploy/release split, then immediately volunteer the conditions. Candidates who stop after "flags let you turn things off quickly" sound like they have read about flags; candidates who say "and the off path has to still work, which is why we test both states" sound like they have owned one.

  • Does the same reasoning hold if the flagged code path writes to the database?
    Weakly. Turning the flag off stops new writes, but the rows already written in the new shape are still there and the old path has to cope with them. Flags guarding read or render paths are clean levers; flags guarding writes need a forward-compatible reader, or the flip leaves you with data your fallback cannot parse.
  • A flag change is a production change that skips the pipeline entirely. What does that cost you?
    It skips code review, CI and any staged rollout the pipeline would have given you, so a mistyped flag value can hit 100 percent of traffic instantly. Teams compensate by treating flag changes as audited events with an actor and a timestamp, restricting who can raise a flag's exposure, and alerting on flag changes so responders correlate them with the graphs.
  • If flags make withdrawal cheap, why keep rollback capability at all?
    Because most production defects are not behind a flag. Shared code, dependency upgrades, config baked into the image and infrastructure changes all fail outside any branch you guarded, and a flag cannot reach them. Flags shrink the set of incidents that need a rollback; they never empty it.

A redeploy is rebuilding the wiring to remove a lamp from a room; a flag is the light switch that was installed while the wall was open.

saying these in an interview costs you the question

  • Deploying the code is the same event as releasing the feature
  • A flag read once at startup still gives runtime control
  • Flags make rollback capability unnecessary
  • Flipping a flag undoes side effects that already happened
  • Any code behind a flag is safe to ship untested

context

open as a page

A feature has served 100% of traffic behind a permanently-enabled flag for six months, and your CI suite only ever runs with that flag on. Why is that a reliability problem, and how would you cover both states without letting the test matrix explode?

level: middleimportance: should knowfreq 48%

basics

~20 s

The off path is your rollback path, and untested code that has served no traffic for six months has probably rotted. Test both states only for flags that could realistically be flipped, rather than every combination of every flag.

open as a page

At 3am the on-call engineer sees error rate climbing on a feature that shipped behind a flag at 20% of traffic. What should decide whether they are allowed to ramp that flag to zero without waking the feature owner, and what should the ramp itself look like?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pre-authorization decided when the flag was created, not at 3am: a registered owner, a known safe off state, and no irreversible side effects. Given that, ramp straight to zero — ramping down fast is not the same decision as ramping up slowly.

open as a page

A team's incident plan says the mitigation for their recommendation service failing is a kill switch that disables recommendations, but that switch has never been flipped in production. Why is that not yet a control, and what would you do to make it one?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An unexercised switch has unverified wiring, an unverified fallback experience and unverified access. Make it a control by drilling it on a schedule in production, measuring time from flip to effect, and confirming the degraded state against a service level indicator.

open as a page

Your production flag service holds around 300 flags and nobody can say which are still needed. As the engineer setting policy, how would you run flag hygiene as an ongoing operational review, and which flags would you deliberately never expire?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Treat stale flags as untested code paths with a tracked count: require an owner and an expiry at creation, review overdue flags on a fixed cadence, and delete both the flag and the losing branch. Permanent operational levers are exempt but must be drilled.

open as a page