skip to content

At 3am the on-call engineer sees error rate climbing on a feature that shipped behind a flag at 20% of traffic. What should decide whether they are allowed to ramp that flag to zero without waking the feature owner, and what should the ramp itself look like?

level: seniorimportance: should knowfreq 40%

answer

  1. decide authority at creation time
  2. down is safe, up is a release
  3. zero is last week's configuration
  4. watch for side effects that strand data
  5. record it; the next shift inherits it

basics

~20 s

Pre-authorization decided when the flag was created, not at 3am: a registered owner, a known safe off state, and no irreversible side effects. Given that, ramp straight to zero — ramping down fast is not the same decision as ramping up slowly.

solid answer

~50 s

The authority question should already be answered before the pager fires. When a flag is created it gets registered with an owner, an expected effect, and a verdict on whether reducing its exposure is safe to do unilaterally. Ramping *down* a release flag is normally pre-authorized because it returns the system to the pre-release state, which is the best-known-good configuration you have. The exceptions are flags where off is not a safe state: anything that writes data the old path cannot read, one-way migration flags, or flags with billing or compliance consequences — those escalate. On the ramp shape, go straight to zero. Stepping 20 to 10 to 5 preserves diagnostic signal but prolongs user impact and burns budget for minutes per step, and diagnosis is not the job at 3am. Then announce it, record who flipped what and when, and hand the non-default state to the next responder.

go deeper

for a junior

Know that reducing a flag's exposure is the safety direction and that the flip should be announced, because others reading the graphs need to know the system is no longer in its default configuration.

for a middle

Explain the asymmetry between ramping down and ramping up, and name the case where off is not a safe state — a flagged path writing data the disabled path cannot read.

for a senior

Demonstrate that authority was settled at flag-creation time, justify going straight to zero in cost terms, and describe verifying propagation and handing the non-default state to the next shift.

for a principal

Own the policy: what on-call may do unilaterally, how flags are classified when created, and how you keep the pre-authorized set small enough to be trustworthy and large enough to be useful.

## The decision is made in advance, not at 3am The worst version of this scenario is an on-call engineer who can see the lever, believes it will help, and spends fifteen minutes deciding whether they are permitted to pull it. That hesitation is an organizational failure, not an individual one, and the fix belongs at flag-creation time. A flag that is going to be an operational control should be registered with, at minimum: an **owner** (a team, not a person), the **expected effect** of turning it off in one sentence, and an explicit **safe-to-disable verdict**. That verdict is the pre-authorization. With it recorded, the 3am decision collapses to a lookup. ## What makes a flag safe to ramp down unilaterally **The off state is a state you have run before.** For a release flag at 20 percent, zero is where the system was last week, and it is the configuration with the most production evidence behind it. That is a strong reason to prefer it under uncertainty. **The change is reversible.** Ramping back up costs another flip. Nothing is destroyed by going to zero. **No side effects strand data or obligations.** This is the real disqualifier. If the enabled path writes records in a format the disabled path cannot read, or completes half of a migration, or changes what customers are billed, then "off" is not a return to a known state — it is a third, novel state. Those flags need the owner, and they should be marked that way when they are created so the on-call engineer does not have to work it out from first principles at 3am. **Down, not up.** Pre-authorization is asymmetric. Reducing exposure is a safety action; increasing it is a release decision. On-call should always be able to reduce, and should essentially never raise a flag's exposure during an incident. ## The shape of the ramp Once the decision is to reduce, the next question is whether to step down or go to zero. There is a real tension: - **Stepping down (20 to 10 to 5 to 0)** keeps some traffic on the new path, which preserves signal. If the error rate falls proportionally you have confirmed the flag is the cause, which is useful for the postmortem. - **Going straight to zero** ends the impact immediately, at the cost of that signal. During an active incident with user-visible errors, go to zero. The reasoning is about cost: each intermediate step costs the propagation delay plus an observation window — call it a couple of minutes — while users keep hitting errors and the budget keeps burning. You can recover the causal evidence afterwards by re-enabling in a controlled way, in daylight, with the owner present. You cannot recover the minutes. Stepping down is defensible when the symptom is degradation rather than failure, when the feature is load-related and you are trying to find the level the system tolerates, or when turning it fully off has its own cost. Being able to articulate *when* each applies is what separates a real answer from a rule recited. ## After the flip The flip is not the end of the responder's obligation: **Verify the effect landed.** Know the propagation time and watch the indicator turn over. If the error rate does not fall within that window, the flag was not the cause — a genuinely valuable finding that stops the team from anchoring on a wrong theory. **Announce it in the incident channel** with the flag name, the old and new value, and the timestamp. Anyone joining later, and anyone reading graphs afterwards, needs to know the system is not in its default configuration. **Leave the state discoverable.** The system is now running with a feature disabled. That belongs in the handoff, because the next shift must not be surprised by it, and because a flag left at zero with nobody accountable becomes a stale flag: an untested path that will eventually be re-enabled by accident. **Do not re-enable during the incident** to test a theory unless the incident lead explicitly agrees. Toggling back and forth while people are debugging destroys the correlation everyone is reading the graphs for. ## What interviewers are listening for They want to hear that you had the authority conversation before the incident, that you distinguish reducing exposure from increasing it, that you can name the class of flags where off is not safe, and that you treat the flip as a recorded production change rather than a private action. Candidates who say "I would call the owner to be safe" without noticing that the call costs fifteen minutes of user impact are describing an organization that has not thought about this in advance.

  • Which flags would you explicitly mark as requiring the owner before anyone touches them?
    Flags whose off state is novel rather than previous: anything writing data the disabled path cannot read, one-way migration switches, flags that change billing or a contractual behaviour, and anything mid-flight in a schema change. Mark them at creation with the reason, so the on-call engineer reads a one-line answer instead of reconstructing the risk under pressure.
  • The error rate does not fall after you ramp to zero. What does that tell you?
    Once you are past the known propagation window, the flagged feature was probably not the cause — or not the only one. That is a useful negative result: say it out loud in the channel so the team stops anchoring on that theory. Leave the flag at zero for now, since re-enabling adds a variable while you are still hunting.
  • Should the flag go back on before or after the postmortem?
    After there is an understood cause and a fix, and never quietly. Re-enable in daylight with the owner present, at low exposure, watching the same indicator that fired. A feature silently switched back on because someone assumed it had been fixed is a repeat incident waiting to happen, and the flag's state should be an explicit action item with a name against it.

saying these in an interview costs you the question

  • Always wake the owner before touching any flag
  • Ramp down in small steps to preserve diagnostic signal
  • Re-enable the flag to confirm the theory mid-incident
  • Turning a flag off is safe by definition
  • A flag flip needs no announcement because nothing was deployed

context