skip to content

Your experiment platform reports per-event standard errors for every ratio metric: how do you prioritise fixing it?

level: principalimportance: nice to knowfreq 32%

answer

  1. measure the damage before proposing a fix
  2. count decisions that would have flipped
  3. A/A significance rate versus nominal
  4. make it the platform default, not opt-in
  5. publish new sample sizes before cutover

basics

~20 s

Quantify the damage first by re-analysing past experiments with user-level variance and counting how many decisions flip. Then make correct variance the platform default for ratio metrics, and prepare teams for wider intervals and larger sample-size requirements.

solid answer

~50 s

Start with evidence, not theory. Re-run a few hundred historical experiments with user-level variance alongside the shipped per-event numbers, and report two things: the typical ratio between the two standard errors, and how many launch decisions would have flipped. That converts an abstract statistical objection into a count of bad decisions. Then fix it at the platform layer rather than metric by metric — a closed-form delta-method variance over per-user totals is cheap enough to be the default for the whole catalogue, with resampling reserved for metrics whose assumptions fail. Sequence the rollout: correct the primary metrics that gate launches first, publish the new sample-size requirements so teams are not surprised mid-experiment, and be explicit that past "wins" computed the old way are not being retroactively revoked but should not be cited as precedent. Expect pushback about lost power and answer it with the A/A false-positive rate.

go deeper

for a junior

Recall that fixing a platform-wide variance bug is not just a formula change: results get less significant, and teams need warning. Knowing that the correction widens intervals is the level-appropriate takeaway.

for a middle

Be ready to describe how you would demonstrate the problem empirically — reanalysing past experiments, or repeated random splits of control — rather than arguing from theory alone in a room that does not share your notation.

for a senior

Show you can scope the work: which metrics gate decisions and go first, which estimator becomes the default, and what the compute and migration costs look like. Expect to explain how you avoid changing variance mid-experiment.

for a principal

Own the tradeoff between short-term experiment velocity and the credibility of every future result, and defend a sequencing that spends the least trust. Be ready to say what you would not fix this quarter and why.

## Frame it as a decision-quality problem Nobody outside the analytics team will fund "our standard errors are wrong". They will fund "we shipped changes that did nothing". The first move is therefore measurement, not migration. **Quantify with history.** Take a large sample of completed experiments and recompute each ratio metric's uncertainty at the randomization unit. Report the distribution of the ratio between the honest and the shipped standard error, and the share of past launch decisions that would have gone the other way. A finding like "standard errors were typically two to four times too small and roughly a quarter of significant results were not significant" is the entire business case. **Corroborate with A/A.** Repeatedly split control users at random and record how often each ratio metric comes back significant. Under correct variance that should sit near the nominal rate; a sharply higher rate is a direct, assumption-free demonstration that intervals are too narrow. It is slower than re-analysis but far harder to argue with, so it is worth running on the handful of metrics that gate the most decisions. ## Choose the default, not a menu The engineering decision is what the platform computes automatically. - **Closed-form variance at the randomization unit as the default.** It runs over per-user aggregates, costs about one pass, and therefore scales to a whole metric catalogue on a schedule. Making it the default matters more than which of the valid estimators you pick: an option that analysts must remember to enable will be forgotten precisely when the result is exciting. - **Resampling as the escape hatch.** Reserve it for the metrics whose assumptions genuinely fail — near-zero or frequently-zero denominators, small user counts, extreme skew. It costs orders of magnitude more compute, so scope it deliberately rather than turning it on everywhere. - **Prefer unit-aligned metrics where they answer the question.** Some ratio metrics exist only because someone summed events. If the decision is genuinely about users, a per-user metric sidesteps the whole problem and is easier to explain. ## Sequence the rollout 1. **Metrics that gate launches first.** The primary and guardrail metrics that decide ship-or-not carry all the decision risk; long-tail diagnostic metrics can follow. 2. **Do not change variance mid-experiment.** Switching the calculation on a running test invites the accusation that you moved the goalposts. Cut over at experiment boundaries and stamp each result with which convention produced it. 3. **Publish revised sample-size requirements before the cutover.** Honest variance means larger required samples for the same detectable effect. Teams that discover this when their experiment fails to conclude will blame the change rather than the arithmetic. 4. **Do not retroactively revoke past launches.** Recompute them for learning, tell teams which past results no longer clear the bar, but relitigating shipped decisions burns the credibility you need for the rollout. ## Handle the organisational reaction The predictable objection is that experimentation just got slower and fewer things reach significance. Three answers work: - **The A/A evidence.** The old regime was not faster, it was wrong: a large share of those quick wins were noise, and the team paid for them in unexplained metric drift and in follow-up work built on false premises. - **Power is a design lever.** If the honest sample sizes are infeasible, the response is to change the design — longer runs, better-chosen metrics, variance-reduction techniques — not to restore an understated standard error. - **Trust compounds.** A platform whose intervals mean what they say lets leadership act on results without a private discount factor. That is the durable argument, and it is why this is a platform default rather than an analyst's discretionary choice. ## What you own as a lead The judgment call is not whether the mathematics is right — it is. It is how much of the organisation's short-term velocity you are willing to spend, and in what order, to make experiment results mean what they claim. Deciding to fix the launch-gating metrics this quarter and let the diagnostic long tail wait two more is a defensible allocation; leaving the default wrong because the correction is unpopular is not.

  • How do you answer a product lead who says experimentation just got slower?
    Show that the old speed was borrowed, not earned: the A/A significance rate under the previous calculation was far above nominal, so a large share of past quick wins were noise. Then treat power as a design problem — longer runs, better metric choice, variance reduction — rather than restoring an understated standard error, which buys velocity by making the results meaningless.
  • Would you recompute and republish results for experiments that already shipped?
    Recompute them for learning, yes; revoke the decisions, no. Publish the reanalysis so teams see which conclusions no longer clear the bar and stop citing them as precedent, but relitigating shipped launches spends credibility you need for the rollout and rarely changes anything that is still reversible.
  • How do you decide which metrics get the expensive resampling estimator?
    Scope it by where the closed-form approximation actually fails: denominators that are often zero or near zero, small user counts, and per-user totals so skewed that the linearisation is unreliable. Those are a minority of the catalogue. Everything else gets the cheap default, because a correct estimator that runs everywhere beats a better one that runs on a few metrics.

saying these in an interview costs you the question

  • Proposes the fix with no estimate of the damage
  • Leaves correct variance as an opt-in analyst setting
  • Switches the calculation mid-experiment
  • Promises the change will not reduce statistical power
  • Retroactively revokes previously shipped launch decisions

context