skip to content

What is dilution in an A/B test where only 3% of assigned users ever see the change?

level: middleimportance: must knowfreq 72%

answer

  1. most assigned users cannot move
  2. effect shrinks, noise does not
  3. multiply the absolute effect by the trigger rate
  4. about one over the trigger rate more traffic
  5. power problem, not a bias problem

basics

~20 s

Dilution is the shrinking of a measured effect when most assigned users never encounter the change. At a 3% trigger rate an effect among exposed users appears about thirty times smaller in the all-up read, while the noise stays.

solid answer

~40 s

If untriggered users are genuinely unaffected, the all-up absolute effect equals the trigger rate times the effect among triggered users. A feature that only fires on the settings page for 3% of assigned users therefore turns a real 1 percentage-point gain among those users into about 0.03 points across everyone. The estimate is still honest; what collapses is sensitivity. The effect shrinks by the trigger rate while per-user variance stays roughly where it was, so detecting the same underlying effect through the all-up read needs on the order of one over the trigger rate more assigned users, which at 3% is roughly thirty times the traffic. That is why a genuinely good narrow feature routinely reads flat all-up: the experiment was never powered for a thirty-fold dilution, not because nothing happened.

go deeper

for a junior

Be ready to say why a change that only 3% of users encounter shows up tiny when averaged over everyone, and that the other 97% still add noise.

for a middle

Explain the arithmetic: all-up absolute effect equals trigger rate times triggered effect, and required traffic grows roughly with one over the trigger rate.

for a senior

Show that you catch dilution at design time by powering on the exposed population, rather than discovering it when a real win reads flat.

for a principal

Frame the diluted number as the input to an investment call: is the next gain in improving the feature or in raising how many people ever reach it?

## What dilution is Dilution is what happens to an effect size when the population you average over is much larger than the population the change can touch. It is arithmetic, not a bug, and it is the dominant reason narrowly scoped experiments read flat. ## The arithmetic on the absolute scale Let `p` be the trigger rate, the share of assigned users who reach the change. Let `d` be the average absolute effect on triggered users, and assume untriggered users are unaffected, so their effect is exactly zero. Averaging over everyone assigned: `d_allup = p * d + (1 - p) * 0 = p * d` With `p = 0.03` and `d = 1 percentage point of conversion`, the all-up effect is `0.03` percentage points. Nothing was lost or hidden; the population average of a change that only touches 3% of the population simply is small. ## The relative scale needs one more term Relative lift does not scale by `p` alone. If `m_t` is the control-arm metric level among users who trigger and `m_all` is the level across all assigned users, then `relative_lift_allup = relative_lift_triggered * p * (m_t / m_all)` When triggered users are typical, `m_t / m_all` is about 1 and the shortcut of multiplying the relative lift by the trigger rate works: a `+9%` triggered lift at `p = 0.03` becomes roughly `+0.3%` all-up. But users who reach a settings page are frequently heavier, more engaged users whose baseline is well above average, in which case the shortcut understates the impact. The robust habit is to convert to absolute units, scale there, and convert back at the end. ## Why dilution costs power, not validity The diluted estimate is unbiased for the thing it estimates: the effect of shipping this change on the whole assigned population. Untriggered users contribute the same distribution of outcomes to both arms, so they do not push the difference in either direction in expectation. What they do contribute is variance. Sample size for a fixed power scales like `variance / effect^2`. Dilution divides the effect by roughly `1/p` while leaving per-user variance broadly unchanged, so the squared term dominates. Comparing the two designs in units of assigned users, the all-up read needs on the order of `1/p` times more traffic than a correctly logged triggered read to detect the same underlying effect, holding the outcome variance comparable. At a 3% trigger rate that is a factor of about thirty. An experiment sized for a one-week read becomes a seven-month read. ## How this shows up in practice - A team ships a feature that clearly works when you watch people use it, and the readout is a wide interval straddling zero. - The point estimate moves erratically week to week, because the signal is a rounding error relative to the noise of the untriggered majority. - Someone proposes rescuing the result by filtering to exposed users after the fact, which is the right analysis done at the wrong moment and with no control-side record to support it. ## The healthy response Decide the exposed population before launch, instrument the trigger on both arms, and power the experiment on the triggered read while still reporting the all-up number as the business impact. If the trigger rate is genuinely tiny and cannot be instrumented, be honest at design time that the experiment cannot answer the question in the available traffic, rather than running it and reading tea leaves afterwards. ## The number that dilution actually teaches you Dilution is also information, not just an obstacle. A change with a large triggered effect and a 3% trigger rate has a small business impact today. That points at a different investment question: is the next win in making the feature better, or in making more people reach it? The diluted number is the honest input to that decision, and it should never be quietly replaced by the triggered number when someone asks how much the launch was worth.

  • Does dilution bias the all-up estimate of the effect?
    No. Untriggered users draw outcomes from the same distribution in both arms, so they cancel in expectation and the all-up difference remains an unbiased estimate of the effect of shipping on the whole population. What they add is variance. Dilution is a sensitivity problem: the interval gets wide relative to the effect, so real wins fail to reach significance even though nothing about the estimate is systematically wrong.
  • Triggered users convert at twice the site-wide rate. Can you still multiply the relative lift by the trigger rate?
    Not safely. Relative lift scales by the trigger rate times the ratio of the triggered baseline to the overall baseline, so a triggered group converting at twice the average roughly doubles the scaled figure compared with the naive shortcut. Convert the triggered lift into absolute units, multiply by the trigger rate, then express the result against the all-user baseline.
  • If dilution is this severe, why report the all-up number at all?
    Because it is the number that answers what shipping does to the business, and it is the only one that is comparable across launches competing for the same headroom. The triggered number tells you whether the feature works; the diluted number tells you what it is worth today. A memo that quotes only the triggered figure will overstate the launch by whatever the reciprocal of the trigger rate happens to be.

Stirring a spoonful of dye into a bathtub. The dye is still all there, but no camera pointed at the tub will register the colour change.

saying these in an interview costs you the question

  • Concludes from a flat all-up read that the feature does nothing
  • Says dilution biases the estimate instead of costing power
  • Multiplies a relative lift by the trigger rate with no baseline check
  • Filters to exposed users only after seeing a disappointing result
  • Sizes the experiment on the triggered effect but reads it all-up

context