skip to content

When choosing an A/B test OEC, how do you trade sensitivity against long-term alignment?

level: middleimportance: must knowfreq 63%

answer

  1. two axes, not one
  2. detectable versus meaningful
  3. spread relative to the mean
  4. noisy revenue, quiet sessions
  5. validate a proxy against past outcomes

basics

~20 s

Pick the most aligned metric that can still detect a realistic effect with the traffic you have. Sensitive short-term metrics decide fast but drift from what the business wants; truer long-run metrics are often too noisy or too slow to see.

solid answer

~50 s

Sensitivity is whether a real effect of plausible size will actually show up in the metric at your sample size, driven by the metric's spread relative to its mean rather than by how much the metric matters. Alignment is whether moving the metric really means the product got better over the long run. The two conflict: a per-user revenue criterion is well aligned with the business but noisy enough that a few percent change may be undetectable in two weeks, while a per-user sessions criterion is quieter and moves fast but can rise for reasons nobody wants. My rule is to take the most aligned metric still detectable at realistic traffic, and to justify any step down toward a shallower proxy with evidence — past experiments where the proxy and the long-run outcome agreed. If nothing is both aligned and detectable, run longer or reduce variance rather than quietly redefining success.

go deeper

for a junior

Know that some metrics are too noisy to show a small change in a short test, and that a flat result with a wide interval is not the same as no effect.

for a middle

Be able to explain what makes a metric insensitive — spread relative to its mean at the sample size available — and to name the tradeoff against how well the metric tracks what the business wants.

for a senior

Show judgment about when to run longer or reduce variance instead of downgrading the criterion, and describe how you have checked a proxy against the outcome it stands in for.

for a principal

Own the policy. Expect to argue when an org should accept a short-horizon proxy at all, what evidence licenses it, and how you prevent a convenient proxy from quietly becoming the definition of success everywhere.

## Two different properties, often confused **Sensitivity** is a statistical property. It asks: if the treatment really changes this metric by an amount worth caring about, will the experiment detect it? For a per-user average metric estimated from `n` users per arm, the standard error of the difference shrinks roughly like `1/sqrt(n)`, and the scale of the noise is the metric's standard deviation across users, `s`. What matters for detecting a *relative* change is the ratio `s / mean` — the coefficient of variation. A metric with a large coefficient of variation needs far more users to resolve the same percentage effect than a metric with a small one. **Alignment** is a substantive property. It asks: if this metric goes up, is the product actually better in the way the business means over months and years? Nothing statistical tells you this. It is a claim about the world that has to be argued, and ideally checked, not assumed. A metric can score high on one and low on the other, and the pair almost always trade off. ## Why they conflict The outcomes the business really wants — a customer still active a year from now, lifetime value, brand trust — are either unobservable inside an experiment window or enormously variable across users. The outcomes that are easy to move and easy to detect — clicks, page views, sessions — are shallow, and a variant can move them by doing something the user does not benefit from. Take the concrete choice between **revenue per user** and **sessions per user** as the criterion for the same experiment. *Revenue per user* is close to the thing the business ultimately cares about. But spending is extremely uneven: most users contribute nothing in a short window and a small minority contribute a lot, which makes the spread across users large relative to the mean. That large coefficient of variation is exactly what suppresses sensitivity, so a genuine 1% improvement may sit comfortably inside the confidence interval at the traffic a single team can get in two weeks. Reading that experiment as *no effect* would be wrong; it is *no detectable effect*, which is a statement about the design, not the product. *Sessions per user* is much quieter — most users have a small number of sessions, the distribution is tighter relative to its mean — so the same experiment resolves a small relative change. The price is alignment: sessions can rise because users find what they need faster and come back, or because the product became confusing and they had to keep returning. The metric does not distinguish those. That is the trade in its cleanest form: the aligned metric cannot see the effect, and the sensitive metric cannot tell you what the effect means. ## The short-horizon problem A related version of the same tension is time. Suppose the outcome the team genuinely cares about is whether a user is still retained twelve months from now. No two-week experiment can observe that. So the criterion becomes a short-term proxy chosen precisely because the true outcome is out of reach: an early-lifecycle behaviour believed to predict long-run retention. This is a defensible move, but it converts an empirical question into an assumption. The proxy is standing in for the real outcome, and the strength of the whole decision now rests on how well it does that. Two things follow: 1. **The substitution should be argued explicitly**, in the design doc, not smuggled in. Write down which long-run outcome the proxy stands for and why you believe the link. 2. **The link should be validated with data where possible.** The strongest form of evidence is historical: take past experiments where both the proxy and the eventual long-run outcome are now observed, and check whether treatment effects on the proxy predicted treatment effects on the outcome — in direction and roughly in magnitude. Correlation between the two metrics *across users* is much weaker evidence, because a metric can be a fine descriptive correlate of good users while being useless at tracking what an intervention does. A proxy that has never been checked against the outcome it replaces is a hypothesis being used as a decision rule. ## Working the trade in practice A usable procedure: - **Start from the most aligned metric you can name**, whether or not it is measurable. That is the target the criterion is approximating. - **Estimate detectability** for each candidate: given historical variance and available traffic, what relative effect could this metric resolve in the planned window? Anything that cannot resolve an effect of the size the change plausibly produces is disqualified as a criterion. - **Choose the most aligned surviving candidate**, and record what you gave up. - **Try to buy back sensitivity before compromising alignment.** Running longer, pooling traffic across surfaces, restricting to an eligible population where the change can actually bite, and variance-reduction techniques all improve detectability without changing what the metric means. Weakening the criterion should be the last move, not the first. - **Re-examine the choice periodically.** A proxy that was validated two years ago on a different product surface is not automatically still a good stand-in. ## What interviewers are listening for The weak answer treats sensitivity as a property of how important a metric is, or asserts that the business metric must be the criterion because the business cares about it, with no account of whether the experiment can see it. The strong answer separates the two axes cleanly, names the mechanism that makes a metric insensitive (spread relative to mean, at the sample size available), and treats any move toward a shallower proxy as a claim that has to be defended with evidence.

  • An experiment on revenue per user comes back flat with a wide interval. Does that mean no effect?
    No — it means no detectable effect. A wide interval that comfortably contains both zero and the effect size you were hoping for is uninformative, not negative. Report the interval rather than a binary verdict, state what effect the design could have resolved, and treat the result as a power problem: run longer, pool traffic, or reduce variance before concluding the change does nothing.
  • How would you actually validate that a short-term proxy predicts the long-run outcome?
    Use the back catalogue of experiments. For past tests where the long-run outcome is now observable, compare the treatment effect on the proxy with the treatment effect on the outcome across experiments, and check whether the proxy got the sign right and tracked magnitude. Agreement across many experiments is the evidence you want; user-level correlation between the two metrics is much weaker and routinely misleads.
  • What can you do to improve sensitivity without changing which metric is the criterion?
    Buy sample or reduce noise. Run longer, pool traffic across eligible surfaces, and restrict the analysis population to users who could actually be affected by the change, since users who never see it only add noise. Variance reduction using pre-experiment covariates lowers the standard error without redefining the metric. All of these preserve alignment, which is why they come before swapping in a shallower proxy.

saying these in an interview costs you the question

  • Treating sensitivity as how important a metric is
  • Reading a wide flat interval as proof of no effect
  • Assuming a proxy predicts the long-run outcome without checking
  • Swapping in a shallow metric before trying variance reduction
  • Claiming user-level correlation proves a metric is a good proxy

context