skip to content

After shipping a feature, is a permanent 1% holdback kept for a quarter worth its cost?

level: principalimportance: should knowfreq 30%

answer

  1. measurement continues after the experiment ends
  2. one percent is a very small arm
  3. every holdback is a live code fork
  4. one company-level holdback beats many
  5. named owner and a sunset date

basics

~20 s

Sometimes, and rarely per feature. A holdback buys continued measurement after the experiment ends, but 1% is insensitive and every holdback forks the product. A single company-level holdback covering many launches usually beats one per feature.

solid answer

~50 s

The case for a post-ship holdback is that experiments stop measuring the moment they end, while the feature keeps running: a holdback is the only way to say later what shipping actually bought. The costs are real and often understated — 1% of users on a divergent experience, two code paths to build, test and support indefinitely, and a comparison so small that it can only ever resolve large differences, so a quarter buys less certainty than people assume. My default is a single company-level holdback that withholds a whole slate of launches from a small population, since the combined effect is much larger than any one feature's and there is one fork rather than dozens. Per-feature holdbacks earn their keep only when the feature plausibly carries slow-building harm or a large ongoing cost. Every holdback needs a named owner and a sunset date agreed before launch, and features that fix accessibility, security or correctness get no holdback at all.

go deeper

for a junior

Know what a post-ship holdback is: a small slice of users kept on the old experience after launch so the comparison still exists later.

for a middle

Be able to name both sides — continued measurement and regression insurance against a maintained second code path and a comparison too small to resolve modest effects.

for a senior

Show you would state the sensitivity up front, so an inconclusive result from a tiny holdback is never presented as evidence that the feature did nothing.

for a principal

Own the policy: default to one company-level holdback with a fixed refresh and sunset cadence, restrict per-feature holdbacks to a defined short list, and make ownership and end dates a launch requirement.

## What a post-ship holdback is When a feature wins its experiment and ships, the usual move is to turn it on for everyone. A **holdback** keeps a small slice of users — typically 1% to 5% — on the old experience after launch, so that the comparison continues to exist. A quarter later, you can still ask what the feature is worth. ## The case for keeping one **Experiments have short horizons; products do not.** A two-week test tells you about two weeks. A holdback is the only mechanism that keeps a randomized comparison alive long enough to say anything about a longer horizon. **Aggregate accountability.** Individually measured launches usually add up to more than the company metric actually moved. A holdback that withholds many launches at once from a small group measures the *sum* directly, and the gap between that sum and the claimed total is the single most informative number a measurement organisation can produce about its own calibration. **Regression insurance.** A comparison that still exists lets you detect that something went wrong after launch — a change interacting badly with a later release, or a cost that only becomes visible over time — without reconstructing a counterfactual from before-and-after data. ## The costs, stated honestly **Sensitivity.** One percent of traffic is a very small arm. Against the other 99%, the precision of the comparison is limited by the small side, so a 1% holdback resolves only comparatively large differences even over a quarter. People who ask for a 1% holdback "to keep measuring" often want a level of certainty it cannot deliver, and the resulting flat, wide interval gets misread as evidence of no effect. **Engineering drag.** Every holdback is a live fork: two code paths, two sets of tests, two support experiences, and a constraint on every subsequent refactor that touches the area. The cost is not the day it is created; it is every change afterwards that has to keep the old path alive. **Product and ethical cost.** A slice of real users is deliberately kept on a worse experience. That is acceptable for a discretionary improvement and unacceptable when the change fixes an accessibility barrier, a security hole, a correctness bug, or a legal obligation. Those ship to everyone. **Organisational drift.** Holdbacks created without an owner and an end date survive by inertia. Years later nobody knows why 1% of users see an old checkout, and nobody dares delete the branch. ## The decision I would actually make Prefer **one company-level holdback** over many per-feature ones. Withhold a whole quarter's slate of launches from a small randomized population, refresh the population on a fixed cadence, and report the combined effect. The advantages compound: the combined effect is large enough for a small holdback to detect, there is a single owner and a single sunset review, and the fork count stays at one instead of scaling with launches. Reserve **per-feature holdbacks** for a short list: features with a plausible slow-building harm, features with a large ongoing operational cost you may want to justify later, and features whose reversal would be very expensive if the initial read turns out to be wrong. For those, write the sunset date and the owner into the launch document, along with what result would cause the holdback to end early. ## Governance that makes it work Three rules keep holdbacks from becoming debt. First, **no holdback without a named owner and an end date** agreed before launch. Second, **a standing review** — a recurring pass over every live holdback that either renews it with a stated reason or removes it. Third, **an honest sensitivity statement** at creation time: write down roughly what size of effect this holdback can resolve in the intended window, so nobody later interprets an inconclusive result as proof of no effect. ## How to frame it in an interview The answer an interviewer is listening for is not yes or no. It is: what does the holdback buy, what does it cost, at what size of effect does it stop being informative, who owns it, when does it end, and is there a cheaper design that buys most of the value — which, at the level of a company, there usually is.

  • Why does a company-level holdback usually beat a holdback per feature?
    Because the quantity you most want to know — did the quarter's launches move the company metric as much as the individual reads promised — is a combined effect, and a combined effect is much larger than any single feature's, so a small holdback can actually resolve it. It also collapses dozens of forks into one, with a single owner, a single sunset review and a single set of maintenance constraints.
  • When would you refuse a holdback outright?
    When the shipped change fixes something users should not be denied: an accessibility barrier, a security or privacy defect, a correctness bug, or anything with a legal or regulatory obligation behind it. Also when the old path cannot be kept alive safely — an unsupported dependency or a data model that has already migrated — since a half-maintained fork is a source of incidents, not measurement.
  • How do you keep a holdback from turning into permanent technical debt?
    Require a named owner and an end date at creation, run a standing review that either renews each live holdback with a stated reason or deletes it, and record at creation roughly what size of effect the holdback can resolve. The third point matters most: without it, an inconclusive small comparison gets read as proof of no effect and the fork is kept for a reason that was never valid.

saying these in an interview costs you the question

  • Creates a holdback with no owner or sunset date
  • Expects a 1% holdback to detect small effects
  • Holds back an accessibility or security fix
  • Adds a per-feature holdback to every launch
  • Reads an inconclusive holdback as proof of no effect

context