skip to content

When would you replace revenue per user with a bounded proxy such as purchased-or-not in an A/B test?

level: seniorimportance: nice to knowfreq 33%

answer

  1. bounded variance beats a fat tail
  2. Bernoulli variance is p(1-p)
  3. blind to how much anyone spends
  4. match the proxy to the mechanism
  5. keep the money metric in the readout

basics

~20 s

Use a bounded proxy when the hypothesis is about whether users buy at all, not how much they spend, and revenue per user cannot be powered at available traffic. A yes/no purchase indicator has variance at most 0.25.

solid answer

~50 s

The case for the swap is purely about variance. A binary purchased-or-not indicator has variance `p(1-p)`, which cannot exceed 0.25, whereas revenue per user has a standard deviation many times its own mean, so the binary version detects effects an order of magnitude smaller at the same traffic. The case against is that it answers a different question: it is blind to how much anyone spends, so a treatment that raises basket size among existing buyers reads as exactly zero, and one that wins many small purchases while losing a few large ones can look positive while revenue falls. So swap when the mechanism plausibly acts on the decision to purchase at all, keep capped revenue in the readout as the money-shaped check, and say plainly that the proxy is the powered metric and the revenue reading is directional.

go deeper

for a junior

Know that a yes/no purchase indicator has variance p times one minus p, at most 0.25, which is why it needs far less traffic than revenue per user to detect a change.

for a middle

Explain the tradeoff concretely: the proxy buys power by discarding magnitude, so any effect on how much existing buyers spend registers as exactly zero.

for a senior

Show the operating package — proxy chosen to match the feature's mechanism, declared before results, capped revenue reported alongside, and a plan to verify magnitude after launch.

for a principal

Own when the organisation is allowed to decide on a proxy at all, what evidence links the proxy to money, and how a proxy win that coincides with a revenue decline gets escalated rather than shipped.

## Why a bounded metric is so much easier to test A yes/no indicator — did this user purchase at all during the window — is a Bernoulli variable. Its variance is `p(1-p)`, maximised at `0.25` when `p = 0.5` and smaller for the rarer conversion rates most products actually see. Compare that with revenue per user, where the standard deviation is routinely several times the mean and is set by a small number of very large spenders. Because the detectable effect scales with the standard deviation over the square root of the sample size, the binary metric is dramatically cheaper to power. A conversion-rate experiment that needs tens of thousands of users per arm can correspond to a revenue-per-user experiment needing millions — and the revenue one still comes back with an interval spanning zero. This is the whole appeal. A bounded metric has no tail to dominate it because the metric cannot exceed one. ## What you give up The binary metric is deliberately blind to magnitude, and that blindness cuts several ways: - **Basket-size effects vanish.** A feature that persuades existing buyers to spend more, without converting anyone new, moves the proxy by exactly zero. If that is your hypothesis, the proxy is the wrong metric, not a conservative one. - **The proxy and revenue can disagree in sign.** A cheaper entry offer can lift the purchase rate while lowering revenue per user. Reading only the proxy would ship a revenue loss as a win. - **The proxy is not the business outcome.** Nobody's plan is denominated in purchase indicators. The proxy is an instrument for detecting a change; the money question is answered, if at all, by the capped revenue reading and by post-launch monitoring. ## The middle ground: bounded versions of the metric itself Between 'unbounded revenue' and 'pure yes/no' sits a family of bounded metrics that keep some magnitude information: - revenue per user capped at a high quantile; - number of purchases per user, capped at a small integer; - watch-time per user capped at three hours, so a handful of always-on sessions stop deciding the average. These retain more of the signal than a binary indicator while still having bounded influence per user. In practice the strongest readout is a small set: a bounded metric powered enough to decide, a capped money metric to keep the decision honest, and the raw unbounded metric reported with its wide interval and no decision rights. ## How to make the swap defensibly 1. **Match the proxy to the mechanism.** Ask what the feature is supposed to change. If it changes whether someone buys, the binary proxy is genuinely measuring the hypothesis. If it changes how much, it is not. 2. **Declare it before the experiment.** Which metric decides the launch is a pre-registration question. Switching to the proxy after revenue comes back flat is result-driven metric selection. 3. **State the mapping you are assuming.** Moving the proxy is only useful because you believe purchase rate maps to revenue in a known direction. Write that belief down; it is the assumption a reviewer should attack. 4. **Keep the money metric visible.** Report capped revenue with its interval next to the proxy. If the proxy is up and capped revenue is meaningfully down, that is a stop signal, not a rounding error. 5. **Verify after launch.** A bounded proxy plus a post-launch holdback on the real metric is a much stronger package than either alone, because it gives the magnitude question a longer horizon and a bigger sample. ## When not to swap If the whole point of the feature is monetisation depth — upsell, bundle size, subscription tier — a purchase indicator will report nothing and you will conclude the feature is inert when it may be valuable. In that case the honest answers are a capped money metric, a longer run, or accepting that this particular decision cannot be made from a single short experiment. ## The sentence that scores 'I swap to a bounded proxy when the mechanism acts on whether users convert rather than on how much they spend — the binary variance ceiling of 0.25 buys me the power — and I keep capped revenue in the readout so a proxy win that costs money cannot be shipped by accident.'

  • How much power does the binary version actually buy?
    A Bernoulli indicator's variance is at most 0.25, while revenue per user often has a standard deviation several times its mean. Since the detectable effect scales with the standard deviation, the binary metric can resolve relative changes that are an order of magnitude smaller at the same traffic — which is the difference between a decidable and an undecidable experiment.
  • What kind of feature is the purchase indicator wrong for?
    Anything whose mechanism is monetisation depth — upsell, bundling, tier changes, larger baskets. Those move spend among users who were already going to buy, so the indicator reports zero and the feature looks inert. There the honest choices are a capped money metric, a longer run, or accepting that one short experiment cannot decide it.
  • The proxy is up and capped revenue is down. What now?
    Treat it as a stop signal and investigate the composition: more buyers at lower value is the usual story, for instance a discount that converts people who would have spent more. Do not ship on the proxy alone; the proxy earned its place as a sensitive detector, not as the definition of success.

saying these in an interview costs you the question

  • Switches to the proxy only after revenue reads flat
  • Treats a proxy win as a proven revenue gain
  • Ignores that basket-size effects are invisible to a binary metric
  • Never reports the capped money metric alongside
  • Assumes any bounded metric is automatically well powered

context