skip to content

Your capped revenue metric is flat but uncapped revenue is up 6% — how do you decide whether to launch?

level: principalimportance: should knowfreq 38%

answer

  1. get the interval before the argument
  2. drop the top few and recompute
  3. the pre-registered metric governs
  4. switching metrics after readout inflates errors
  5. design a tail-targeted follow-up

basics

~20 s

Treat the uncapped 6% as unresolved until you see its confidence interval and how much survives dropping the largest few spenders. The pre-registered capped metric governs the launch; a whale-driven swing calls for a tail-targeted follow-up, not an override.

solid answer

~40 s

First, quantify rather than choose. Put a confidence interval on the uncapped 6%: on a heavy-tailed metric that interval usually contains zero and a great deal more. Then run a drop-the-top-k check — recompute the arm difference without the single largest spender, then the top five — and see how much of the 6% survives. If it collapses, the number is a few users landing in one arm, not an effect. Second, honour the pre-registration: the capped metric was named as the decision metric precisely so a tail swing could not be read as a win after the fact. Third, separate the substantive question: if the feature genuinely targets high spenders, the capped metric was the wrong instrument, and the answer is a tail-targeted follow-up plus a fixed metric plan next time.

go deeper

for a junior

Know that a large uncapped revenue swing is often a few users landing in one arm, and that a point estimate without a confidence interval is not a result.

for a middle

Be able to run and read the checks: the interval on the uncapped difference, the effect after removing the top few spenders, and the count of extreme users in each arm.

for a senior

Show that you can hold the pre-registered decision metric under pressure while still taking the possibility of a tail effect seriously, and specify the follow-up that would actually measure it.

for a principal

Own the policy, not just this call: who may change a decision metric and how that exception is recorded, when a feature gets a tail-aware metric plan up front, and how post-launch holdbacks answer questions a short experiment cannot.

## Why this scenario is a trap A flat capped metric with a large positive uncapped reading is the single most common way a heavy-tailed experiment gets mis-shipped. It offers a story everyone in the room wants to believe — 'the cap is hiding our win' — and that story is sometimes true. The job of a lead is to make the organisation resolve it with evidence rather than with enthusiasm, and to have set the rules before the numbers arrived. ## Step one: quantify the 6% before arguing about it A point estimate on an uncapped revenue metric means very little without its interval, and on a heavy-tailed metric that interval is typically enormous. Ask for three things: 1. **The confidence interval on the uncapped difference.** If it spans, say, minus 4% to plus 16%, the honest reading is that the experiment did not measure uncapped revenue at all. 2. **A drop-the-top-k sensitivity.** Recompute the arm difference with the largest single spender removed, then the top five, then the top twenty. If most of the 6% evaporates by k = 5, the difference is a story about where a few individuals were randomised. 3. **The per-arm composition.** How many users above the cap in each arm, and how much total spend does each arm's tail contribute? An imbalance of two or three extreme users is entirely ordinary under valid randomisation and produces exactly this pattern. Often those three numbers end the conversation without anyone having to invoke policy. ## Step two: the pre-registration governs If the capped metric was declared the decision metric, it decides. This is not bureaucracy — it is the only thing that keeps the false-positive rate near its nominal level. The moment a team may choose, after readout, between capped and uncapped, the effective test is 'either metric moved', and the error rate reflects that broader test. Allowing the override once teaches the organisation that the rule is negotiable, and every future ambiguous result will be resolved by whichever metric flatters it. The defensible position is: the capped metric is flat, so this experiment does not authorise a launch on revenue grounds; here is what we will do to find out whether the tail effect is real. ## Step three: ask whether the metric plan was wrong The genuinely interesting case is when the feature's mechanism is aimed at high spenders — a concierge tier, a high-value bundle, a whale-facing retention play. Then a p99 cap deletes the hypothesis by construction, and the flat capped reading is not evidence of no effect; it is evidence that the wrong instrument was chosen. That is a real failure, and the correct handling is: - acknowledge that the metric plan did not match the hypothesis; - do **not** rescue this experiment by switching metrics after the fact; - design the follow-up around the tail: pre-register a high-spender stratum defined by pre-period behaviour (a pre-treatment quantity, so conditioning on it is legitimate), power it explicitly for the subpopulation, run a longer horizon since large purchases are infrequent, and choose a looser cap or an uncapped analysis knowing what it costs in precision; - consider a post-launch holdback that keeps a fraction of users in control for months, which is often the only realistic way to measure a tail effect at all. ## Step four: what a lead actually decides There are three defensible outcomes, and they should be distinguished out loud: - **Do not launch on revenue grounds.** The capped metric is flat, the uncapped signal does not survive sensitivity. Ship only if some other pre-registered metric justified it. - **Launch on other grounds, monitor revenue.** If the feature was justified by a bounded proxy or a product goal that did move, and the guardrails are clean, launch and keep a holdback to answer the money question over a longer horizon. - **Do not decide yet, run the tail-targeted follow-up.** When the mechanism plausibly acts on the tail and the sensitivity check leaves a real possibility of an effect, the right answer is another experiment, not a coin flip dressed as analysis. ## The organisational fix The lasting output of this situation is not the launch call, it is the policy: which metric decides is written down before the experiment; the cap value and its provenance are platform-level, not per-team; a change of decision metric after readout requires an explicit, recorded exception; and features aimed at high spenders get a tail-aware metric plan from the start. Litigating each case on its merits after the fact guarantees the organisation will eventually ship noise. ## The sentence that scores 'Show me the interval and the drop-the-top-five result. If the 6% does not survive, it is a few users. If it might be real, the answer is a tail-targeted follow-up and a fix to the metric plan — not overriding the metric we pre-registered precisely to prevent this argument.'

  • What single diagnostic settles most of these arguments fastest?
    Recompute the arm difference with the top spender removed, then the top five. On a heavy-tailed metric a swing of this size usually collapses by k = 5, which shows the reading is about where a few individuals were randomised rather than about the feature. Report it alongside the interval on the uncapped difference.
  • Is it ever right to override the pre-registered decision metric?
    Rarely, and never silently. If the metric plan plainly did not match the hypothesis — a whale-facing feature judged on a p99-capped metric — the honest move is to record the mismatch, decline to call this experiment, and rerun with a tail-aware plan. An override without a recorded exception turns the decision rule into whichever metric looks best.
  • How would you design the follow-up so the tail effect is measurable?
    Pre-register a high-spender stratum defined by pre-period behaviour, which is a pre-treatment quantity and so safe to condition on. Power that stratum explicitly, run a longer horizon because large purchases are infrequent, choose a looser cap knowing the precision cost, and plan a post-launch holdback to answer the magnitude question over months.
  • What do you tell a leader who says the cap is obviously hiding the win?
    Agree that it might be, then make it checkable: here is the interval on uncapped revenue, here is the effect after removing the top five users, here is how many extreme spenders each arm happened to get. If the signal survives, we design a test that can see it; if it does not, we have prevented an expensive launch based on three people.

saying these in an interview costs you the question

  • Reads the uncapped 6% as a result without an interval
  • Switches the decision metric after seeing the readout
  • Never checks how much survives dropping the top spenders
  • Blames the cap without asking if the feature targets whales
  • Ships on the wider, noisier metric because it is bigger

context