skip to content

Release Safety

Treating deploys as the leading cause of incidents: progressive rollouts, feature flags as operational controls, and rollback readiness. Interviewers ask how you would ship a risky change safely — this is the SRE-side answer that complements CI/CD tooling.

on this pageshow

questions

17

Your service can only disable a bad new code path by redeploying the previous build, which takes about 25 minutes end to end. What does putting that code path behind a runtime feature flag change about your recovery, and what has to be true of the flag for that improvement to be real?

level: juniorimportance: must knowfreq 70%

answer

  1. deploy is not release
  2. the branch already shipped
  3. config change, not a rebuild
  4. evaluated per request, not at boot
  5. the off path must still work

basics

~20 s

A runtime feature flag turns recovery from a redeploy into a configuration change, cutting mitigation from tens of minutes to seconds. That only holds if the flag is evaluated per request, the old path still works, and the change reaches every instance quickly.

solid answer

~50 s

Behind a flag the new code is already in production, so recovery stops being a rebuild-promote-redeploy cycle and becomes a config change that propagates in seconds. That is the deploy/release split: deploying puts bytes on servers, releasing turns behaviour on for users, and only the second one needs to be reversible in seconds. For the 25 minutes to actually become seconds, three things must hold. The flag has to be evaluated at request time, not read once into a constant at boot, or you are back to restarting processes. The old code path has to still be present and still work — a flag does not test the path it falls back to. And the change has to reach every instance quickly and visibly, so you can confirm the flip took effect instead of assuming it did. If any of those is false, you still own the redeploy.

code

javascript · 20 lines
javascript
// Build-time constant: changing it requires a new deploy.
const NEW_PRICING = true;

function priceCartCompiled(cart) {
  return NEW_PRICING ? newPricing(cart) : legacyPricing(cart);
}

// Runtime flag: changing it is a config write, effective on the next request.
function priceCartFlagged(cart, flags, userId) {
  if (flags.isEnabled('new-pricing', { userId })) {
    return newPricing(cart);
  }
  return legacyPricing(cart);
}

function newPricing(cart) { return cart.items.length * 10; }
function legacyPricing(cart) { return cart.items.length * 12; }

console.log(priceCartCompiled({ items: [1, 2] }));
console.log(priceCartFlagged({ items: [1, 2] }, { isEnabled: () => false }, 'u1'));

go deeper

for a junior

Be able to say plainly that the flagged code is already deployed and the flag only chooses which branch runs, so turning a feature off is a config change rather than another deploy.

for a middle

Explain the mechanics: request-time evaluation versus a boot-time constant, how the value reaches instances by push or poll, and roughly how long that propagation takes.

for a senior

Show you know the limits — an untested off path, side effects already written, and flag changes bypassing every pipeline gate — and say what you put in place to cover each.

for a principal

Own the tradeoff of which changes earn a flag at all. Each flag is a permanent branch and a test-matrix entry, so argue the criteria by expected incidents avoided rather than flagging everything.

## The problem the flag is solving When the only way to withdraw a change is to redeploy the previous build, your mitigation time is bounded by your delivery pipeline: rebuild or re-promote the artifact, roll it across instances, wait for health checks and connection draining. Twenty-five minutes is an ordinary figure for that, and during an incident every one of those minutes is user impact and error-budget spend. A runtime feature flag attacks the problem from a different angle. Instead of changing *which code is running*, it changes *which branch that code takes*. The new path and the old path both ship in the same artifact; a value read at request time decides which one executes. ```javascript // build-time constant: changing it means a redeploy const NEW_PRICING = true; // runtime flag: changing it means a config write if (flags.isEnabled('new-pricing', { userId })) { return newPricing(cart); } return legacyPricing(cart); ``` The first form is not an operational control at all — it is a compile-time decision wearing a flag's clothes. ## Deploy is not release The vocabulary matters in interviews. **Deploy** is the act of getting an artifact onto production machines. **Release** is the act of exposing a behaviour to users. Without flags the two are the same event, so the blast radius of a deploy is the blast radius of every change inside it, and the undo for both is the same slow mechanism. With flags they separate: you can deploy at any cadence, keep the new behaviour dark, and choose the moment — and the audience — for the release independently. Crucially, the *undo* for a release becomes a different, much cheaper operation than the undo for a deploy. This is why flags are treated as a release-safety practice and not merely a coding convenience. They give the on-call engineer a lever that can be pulled during an incident by someone who does not need to know how the build system works. ## The three conditions **Evaluated at request time.** If the flag value is read once at process start into a module-level constant or a cached singleton that never refreshes, flipping it changes nothing until the processes restart — and restarting every instance is a redeploy by another name. Real control means the value is consulted on the code path each time, or at least refreshed on a short interval that you know and can quote. **The off path still works.** A flag's promise is that turning it off returns you to the previous behaviour. That is only true if the previous behaviour is still exercised and still correct. Once a feature has run at 100 percent for months, the off branch has had no traffic, and it can rot silently — it may reference a column that was dropped, or produce a page that no longer renders. The flag looks like insurance and turns out to be a second, untested failure mode. **Propagation is fast and observable.** Flag changes usually reach instances by push (a streaming connection from the flag platform) or by poll (each instance re-fetching every few seconds). Either way there is a real number — a few seconds, sometimes tens of seconds — and you should know it, because during an incident the gap between "I flipped it" and "it took effect" is the window where people start making a second change on top of the first. You also want a signal that confirms the flip landed: the error rate turning over, or a metric that reports the effective flag value per instance. ## What the flag does not give you A flag does not make the change safe to ship untested; it makes it cheap to withdraw. It does not undo side effects that already happened — rows written in the new format, messages published to a queue, emails sent — so a flag guarding a write path is a weaker control than one guarding a read or render path, and you should say so when asked. It does not replace rollback either: a bad change that is not behind a flag, or a bad change in shared code that both branches run, still needs the deploy-level undo. And every flag has a cost. It is a branch in the code, a state in the test matrix, and an entry in an inventory somebody has to prune. The reason experienced teams still take that cost for risky changes is arithmetic: minutes of user impact avoided per incident, multiplied by the number of incidents where the lever gets pulled. ## How to answer it in an interview Give the number first — redeploy in tens of minutes versus a config change in seconds — then name the deploy/release split, then immediately volunteer the conditions. Candidates who stop after "flags let you turn things off quickly" sound like they have read about flags; candidates who say "and the off path has to still work, which is why we test both states" sound like they have owned one.

  • Does the same reasoning hold if the flagged code path writes to the database?
    Weakly. Turning the flag off stops new writes, but the rows already written in the new shape are still there and the old path has to cope with them. Flags guarding read or render paths are clean levers; flags guarding writes need a forward-compatible reader, or the flip leaves you with data your fallback cannot parse.
  • A flag change is a production change that skips the pipeline entirely. What does that cost you?
    It skips code review, CI and any staged rollout the pipeline would have given you, so a mistyped flag value can hit 100 percent of traffic instantly. Teams compensate by treating flag changes as audited events with an actor and a timestamp, restricting who can raise a flag's exposure, and alerting on flag changes so responders correlate them with the graphs.
  • If flags make withdrawal cheap, why keep rollback capability at all?
    Because most production defects are not behind a flag. Shared code, dependency upgrades, config baked into the image and infrastructure changes all fail outside any branch you guarded, and a flag cannot reach them. Flags shrink the set of incidents that need a rollback; they never empty it.

A redeploy is rebuilding the wiring to remove a lamp from a room; a flag is the light switch that was installed while the wall was open.

saying these in an interview costs you the question

  • Deploying the code is the same event as releasing the feature
  • A flag read once at startup still gives runtime control
  • Flags make rollback capability unnecessary
  • Flipping a flag undoes side effects that already happened
  • Any code behind a flag is safe to ship untested

context

open as a page

You roll a change out to 1% of production traffic, the canary's error rate and latency stay clean for an hour, and then the change breaks one enterprise customer's integration the moment it goes to 100%. Why can a 1% traffic canary contain 0% of the users a change affects, and how would you choose the canary population instead?

level: middleimportance: must knowfreq 60%

basics

~20 s

A traffic percentage is a random sample, not a representative one. A rare code path, a single tenant, or one old client version can send so few requests that the canary receives none of them. Choose the canary population deliberately, not just its size.

open as a page

A service runs from an immutable image identified by digest, but its environment configuration is applied separately and always from the config repository's HEAD. Redeploying the previous image digest did not restore the previous behaviour. Why, and what would you change so that one action reverts the whole release?

level: middleimportance: must knowfreq 60%

basics

~20 s

Only the code was versioned. Configuration applied from a moving HEAD stays at its new value, so half the release is still live after the rollback. Bind the artifact digest and the config revision into one versioned release record you revert as a unit.

open as a page

A change passed every canary stage, reached 100% of the fleet, and the failure surfaced two days later. Which classes of failure does a progressive rollout structurally fail to catch, and what controls would you add for them?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Progressive rollouts miss anything that needs time, scale, a specific event, or a specific population: slow leaks and disk fill, dependencies that only saturate at full traffic, monthly batches, and one tenant or old client. Each needs its own control.

open as a page

A release ships application code together with a database migration that renames a column, and the release turns out to be bad. Why can you no longer simply redeploy the previous version, and how should that migration have been sequenced so the code stays rollback-safe?

level: seniorimportance: must knowfreq 62%

basics

~20 s

A renamed column leaves rolled-back code reading a column that no longer exists, and the reverse migration destroys data. Sequence it expand-contract: add the new column, dual-write, backfill, switch reads, and drop the old one only after the rollback window closes.

open as a page

A production release is bad and you decide to roll back. Your delivery system can either redeploy the previously built artifact or revert the commit and let CI build and deploy a fresh one. Which one is the rollback, and why does the other still have to happen?

level: juniorimportance: should knowfreq 55%

basics

~20 s

Redeploy the previously built artifact - it already exists, already ran in production, and returns in minutes instead of a full build. Revert the source commit as well, or the next deploy reships the bad code.

open as a page

A feature has served 100% of traffic behind a permanently-enabled flag for six months, and your CI suite only ever runs with that flag on. Why is that a reliability problem, and how would you cover both states without letting the test matrix explode?

level: middleimportance: should knowfreq 48%

basics

~20 s

The off path is your rollback path, and untested code that has served no traffic for six months has probably rotted. Test both states only for flags that could realistically be flipped, rather than every combination of every flag.

open as a page

An automated canary check scores the new version by comparing its error rate and latency against the metrics the rest of the production fleet reported over the same hour. Why is that comparison misleading, and what should the canary be compared against instead?

level: middleimportance: should knowfreq 45%

basics

~20 s

The fleet has been running for days; the canary started minutes ago. Cold caches, unfilled connection pools and runtime warm-up make a healthy canary look slow. Compare it against a control running the old version, started at the same time on the same class of host.

open as a page

A rollback of a bad release finished two minutes ago and the server-side error-rate graph is falling. What do you check before you call the rollback successful?

level: middleimportance: should knowfreq 48%

basics

~20 s

Confirm every instance actually runs the old version, compare the user-facing SLI against the pre-release baseline rather than the incident peak, watch request rate alongside errors, and clear residual bad state such as poisoned caches and parked messages.

open as a page

At 3am the on-call engineer sees error rate climbing on a feature that shipped behind a flag at 20% of traffic. What should decide whether they are allowed to ramp that flag to zero without waking the feature owner, and what should the ramp itself look like?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Pre-authorization decided when the flag was created, not at 3am: a registered owner, a known safe off state, and no irreversible side effects. Given that, ramp straight to zero — ramping down fast is not the same decision as ramping up slowly.

open as a page

A team's incident plan says the mitigation for their recommendation service failing is a kill switch that disables recommendations, but that switch has never been flipped in production. Why is that not yet a control, and what would you do to make it one?

level: seniorimportance: should knowfreq 44%

basics

~20 s

An unexercised switch has unverified wiring, an unverified fallback experience and unverified access. Make it a control by drilling it on a schedule in production, measuring time from flip to effect, and confirming the degraded state against a service level indicator.

open as a page

You are setting the bake time for each stage of a canary rollout — how long the new version soaks at 1% before it advances to 10%. How do you decide that number, and why is a 10-minute bake worthless for some changes?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Two clocks set the bake: how long it takes to collect enough events at that traffic share to detect the regression you care about, and how long the failure mode you fear needs to appear. Bake for the longer of the two.

open as a page

You roll a service back to its previous version an hour after a bad release. During that hour the new version wrote records containing a field the older version knows nothing about. What problems does the rollback now create, and what limits how far back you can safely go at all?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Data written by the newer version can be unreadable, silently dropped, or overwritten by the restored older code. Your real rollback horizon is the oldest version still compatible with today's schema, message formats and cached data - usually one release back.

open as a page

You set release policy for a service that runs in three regions with three availability zones each. After the initial canary passes, how would you order the rollout across those zones and regions, and how do you decide the total wall-clock time a rollout is allowed to take?

level: principalimportance: should knowfreq 40%

basics

~20 s

Expand along a blast-radius ladder: one zone in the least critical region, then that region, then the remaining regions one at a time, keeping a known-good region until last. Total duration is bounded by mixed-version tolerance, emergency-fix latency and how often you deploy.

open as a page

Your automated canary analysis aborts roughly one rollout in three, nobody can reproduce a defect in the aborted changes afterwards, and engineers are now asking for a flag to skip the check. What is going wrong, and how do you fix it without going blind?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

The gate is scoring noise: too many metrics, too few samples, and thresholds tuned tighter than normal variance. Fix precision — require a minimum sample count, score a small set of high-signal SLIs relative to a control, and return inconclusive instead of aborting.

open as a page

Your production flag service holds around 300 flags and nobody can say which are still needed. As the engineer setting policy, how would you run flag hygiene as an ongoing operational review, and which flags would you deliberately never expire?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Treat stale flags as untested code paths with a tracked count: require an owner and an expiry at creation, review overdue flags on a fixed cadence, and delete both the flag and the losing branch. Permanent operational levers are exempt but must be drilled.

open as a page

You set release-engineering standards for an organisation running several dozen services. Which services would you require to maintain a tested, single-action rollback path, and where would you accept that fixing forward is the only realistic option?

level: principalimportance: nice to knowfreq 35%

basics

~20 s

Mandate a drilled, measured rollback for stateless services on the critical user path, where it is cheap. Accept fix-forward where effects are already irreversible outside the system - then make the pipeline fast, because its duration becomes the recovery floor.

open as a page