skip to content

A 5% long-term holdback six months after ship shows no gain — how do you interpret it?

level: seniorimportance: should knowfreq 38%

answer

  1. null result or underpowered readout?
  2. precision is set by the small arm
  3. does the interval exclude the launch effect?
  4. leakage attenuates toward zero
  5. frozen arm versus a moving product

basics

~20 s

Do not read it as proof the launch gain was fake. Check the interval first: a 5-versus-95 split is precision-limited by the small arm, so it may still contain the original launch effect. Then rule out leakage.

solid answer

~40 s

The first question is whether this is a null result or an underpowered one. With 5% held back and 95% treated, the variance of the difference is dominated by the small arm — the standard error is roughly 2.3 times what a balanced split of the same total traffic would give — so six months of accrual can still leave an interval wide enough to contain the launch effect. If the interval excludes it, the gain genuinely faded and the launch number was largely a transient. If it contains it, the readout is uninformative. Second, check contamination: if held-back users reached the feature anyway, the true contrast is diluted toward zero. Third, the holdback has been frozen against a moving product, so the contrast is the feature plus everything built on top of it.

go deeper

for a junior

Know what a long-term holdback is: a small slice of users kept on the old experience after launch, re-measured later to see whether the gain is still there.

for a middle

Be ready to explain why a small holdback arm limits precision, and why a non-significant result is not the same as a demonstrated absence of effect.

for a senior

Show the full checklist before you accept a null: interval versus the launch effect, leakage into the held-back arm, differential churn, and the fact that the frozen arm is being compared with a product that kept moving.

for a principal

Own what the organisation does with the answer. The lasting value is the calibration it gives every future short-horizon estimate on that surface, and the credibility cost if teams are allowed to explain the readout away.

A long-term holdback keeps a small slice of users on the pre-launch experience for months after ship, and re-reads the metric later. Its purpose is precisely to answer the question a short experiment cannot: does the gain still exist once the reaction to the change has worn off? Reading a null from one correctly is a senior skill because there are at least four ways to get that null. ## 1. It may be a power problem, not an effect problem With a fraction f of users held back and 1 minus f treated, the variance of the difference in means is proportional to 1/f plus 1/(1 minus f), per unit of total traffic. At f = 0.05 that is 20 + 1.05, roughly 21, against 2 + 2 = 4 for a balanced split. The standard error therefore runs about sqrt(21/4), roughly 2.3 times wider than a 50/50 test on the same total traffic; to match a balanced split you would need about five times the traffic. The practical reading: the precision of the whole exercise is set almost entirely by the size of the small arm, and adding more treated users barely helps. So the very first move is to look at the interval, not the point estimate or the significance verdict. Ask whether the interval excludes the original launch effect. If it does not, the readout has failed to distinguish a fully persistent gain from no gain at all, and the correct report is that it is uninformative — not that the effect is gone. Treating a non-significant result as evidence of no effect is the single most common error in this analysis. ## 2. It may be real decay If the interval does exclude the launch effect, that is a genuine finding and an important one: the number the launch was justified on was substantially transient. This is exactly what a novelty-driven readout looks like six months later, and it is the reason long-term holdbacks exist at all. A team that responds by re-litigating the analysis rather than updating its belief about the launch is wasting the instrument. ## 3. It may be contamination Anything that lets a held-back user experience the feature pulls the measured contrast toward zero. Common routes: the same person on a second device or account that lands in the treated group; a shared surface such as a public page, an export, or a notification that reflects the treated experience; a partial gate that covers the main entry point but not every path to the feature; word of mouth or support articles describing the new behaviour. All of these attenuate, and they attenuate toward exactly the null being observed. Ruling contamination out before accepting the null is mandatory, and the cheapest check is whether held-back users show any exposure to the feature at all in logs. ## 4. It may be a comparability problem A six-month holdback compares a frozen experience against a product that kept moving. Everything shipped since launch has been built against the treated experience, and some of it may depend on the launched feature or may have been tuned in its presence. The contrast is therefore the launched feature *plus its downstream consequences*, not the feature in isolation. That is often the more decision-relevant quantity, but it must be described as what it is. In the other direction, a long-lived small arm can drift: users who churn are not replaced identically in both arms, and if the frozen experience is worse the churn is differential, which biases the surviving comparison. ## Putting it together A defensible interpretation reads roughly like this. Report the estimate with its interval and say explicitly whether the interval excludes the launch effect. If it does not, declare the readout underpowered and say what would change that. If it does, check the exposure logs for held-back users to rule out dilution, note the frozen-versus-moving-product caveat, and only then conclude that the launch gain was substantially transient. The value of the conclusion is not that one feature underperformed; it is the correction it supplies to every future estimate the team makes from short experiments, because it calibrates how much of a first-week lift on this surface tends to survive.

  • Why does a 5-versus-95 split cost so much precision compared with a balanced one?
    The variance of the difference adds a term for each arm, proportional to 1 over that arm's share. At 5% the small arm contributes about 20 units against roughly 1 from the large arm, versus 2 and 2 when balanced. That is around 2.3 times the standard error for the same total traffic, and enlarging the treated arm barely moves it.
  • What single check would you run before accepting that the gain decayed?
    Look for feature exposure among held-back users in the logs. Any leakage — a second device, a shared surface, an incomplete gate — dilutes the contrast toward zero and produces exactly the null being observed. If held-back users show real exposure, the readout is contaminated and the decay conclusion is unsupported.
  • How does the moving product complicate the six-month contrast?
    The held-back arm has been frozen while everything else shipped against the treated experience, some of it depending on the launched feature. The measured contrast is the feature plus its downstream consequences rather than the feature alone. That is often the more useful quantity for a keep-or-remove decision, but it must be reported as such.

saying these in an interview costs you the question

  • Reads a non-significant holdback readout as proof of no effect
  • Ignores that precision is dominated by the small arm
  • Never checks whether held-back users saw the feature anyway
  • Assumes the frozen arm is comparable after six months
  • Concludes the original launch result was a false positive

context