How do you check whether an A/B test effect is decaying over exposure time?
answer
- the pooled number hides the shape
- calendar day mixes exposure ages
- re-index on each user's own clock
- compare arms within the same bucket
- the shape should repeat per cohort
basics
~20 sPut every user on their own clock: days since that user first saw the change. Plot the treatment-minus-control gap against exposure age, grouped by first-exposure date. A gap that shrinks with exposure age, repeating per group, indicates decay.
solid answer
~50 sThe single pooled number cannot show decay, and a plot against calendar date is misleading because each calendar day mixes users who are on their first day of exposure with users who are on their thirtieth. So I re-index on exposure age: for each user, days since their own first exposure, then compute the treatment-minus-control gap within each exposure-age bucket. Comparing arms inside the bucket matters, because activity naturally falls after anyone's first visit in both arms, so treatment alone would show a decline that has nothing to do with the effect fading. Then I split by first-exposure date, so that each group of users who entered together gets its own curve. Genuine decay shows the same rise-and-fade shape for every group, offset in calendar time. A dip that appears in all groups on the same calendar day is an external event, not decay.
go deeper
Know that a single overall lift number cannot tell you whether an effect faded, and that the fix is to look at the effect over time rather than as one summary figure.
You are expected to explain the exposure clock precisely: days since each user's own first exposure, gaps computed between arms inside a bucket, and why the calendar axis confounds the trend with the mix of exposure ages.
Demonstrate the discipline of separating decay from noise, external events and survivorship in the deep buckets. Interviewers want to hear that you require the shape to reproduce across cohorts before you act on it.
Frame this as measurement infrastructure rather than a one-off analysis. Whether exposure-age readouts exist by default determines whether every team can detect decay or only the teams that thought to look.
Decay is a claim about how the effect behaves as an individual user accumulates exposure to the change. Diagnosing it is mostly a matter of choosing the right time axis and then not fooling yourself with the composition of each bucket. ## Why the pooled number is blind to it A single treatment-minus-control estimate over the whole run averages every user's first day with their thirtieth. A +8% first-day reaction and a +1% steady state pool to something in between, and that in-between number is reported with a comfortable interval around it. Nothing in the summary reveals that the two components exist. ## Why the calendar axis is the wrong axis The obvious next step is to plot the daily effect against the calendar. This is better than nothing but is confounded by composition. On any given calendar day the treatment arm contains a mixture of users at every stage of exposure: people who joined at launch and have adapted, and people arriving today who are seeing the change for the first time. As the test runs, the mixture keeps shifting, so a downward calendar trend can be produced by changing composition rather than by any individual's effect changing at all. The converse also holds: an effect that decays sharply per user can look almost flat on the calendar if fresh users keep arriving and keep contributing fresh first-day reactions. ## The exposure clock Re-index the data. For every user record the date on which they were first exposed to their assigned experience, then compute, for each integer exposure age d, the metric among treatment users at age d and among control users at age d, and take the difference. Two details make this work: - **The control arm needs the same clock.** Engagement falls naturally in the days after anyone's first visit, in both arms. Plotting only the treatment metric against exposure age will always slope downward, and reading that as decay is a common error. The quantity of interest is the *gap between arms* at equal exposure age. - **Later buckets are thinner and more selected.** Only users who keep coming back reach exposure age 30, and they are systematically the more engaged ones. That makes deep buckets both noisier and drawn from a different population, so a difference between age 1 and age 30 is partly a difference between two groups of people. Restricting to users who have had at least the full window available to them makes the buckets comparable, at the price of dropping recent arrivals. ## Cohorting by first-exposure date The strongest version of the diagnosis splits users into groups by the date each user was first exposed, and draws one exposure-age curve per group. This separates two very different stories that the pooled calendar view confuses: - If every group shows the same shape on its own clock — a large gap early, shrinking to a smaller one — and each group's shape is simply shifted in calendar time relative to the last, that is decay of the per-user effect. It is a property of exposure, and it will repeat for every user who ever gets the feature. - If all groups move together on the same calendar days, regardless of how long each has been exposed, the driver is external to the experiment: a marketing push, an outage, a holiday, a change shipped elsewhere in the product. That is not decay, and it will not repeat per user. - If the earliest group shows a big effect and later groups never do, the change did something one-off at launch, or the population arriving later differs from the population that was already there. ## Distinguishing decay from noise A sequence of point estimates that happens to descend is not decay. Each bucket has its own uncertainty, and the later ones are wide because they hold fewer users. Before calling decay, check that the drop is large relative to those intervals, that it is monotone rather than a single low bucket, and above all that it reproduces across independent groups of users. Reproducing across cohorts is the most convincing evidence available, because each cohort is a separate set of people generating the same shape. ## Reading the direction The same machinery diagnoses the opposite transient. A gap that begins negative and climbs toward zero as exposure accumulates is change aversion wearing off; if it climbs past zero and settles positive, the change carried a real benefit that only materialised once users habituated to it. The decision the readout supports is different in each case, but the diagnostic — exposure clock, arm-to-arm gap within bucket, repeat across cohorts — is identical.
- Why not simply plot the treatment metric against days since first exposure?Because activity falls naturally after anyone's first visit, in both arms. The treatment curve alone will slope down whether or not the effect is fading. Only the treatment-minus-control gap at equal exposure age isolates the effect, since the same natural decline is present in the control curve and cancels out of the difference.
- What does it mean if all cohorts dip on the same calendar day rather than at the same exposure age?That is an external event, not decay. Something outside the experiment moved the metric for everyone at once: an incident, a campaign, a seasonal break, or another release. Per-user decay is anchored to each user's own exposure clock, so it appears at the same exposure age in every cohort and is offset in calendar time.
- Why are the deepest exposure-age buckets the least trustworthy?They contain only users who kept returning, so they are both small and selected toward the most engaged people. The wide intervals invite over-reading a single low or high point, and the population difference means a gap at age 30 is not measured on the same users as the gap at age 1. Restricting to users with a full window available keeps the comparison honest.
saying these in an interview costs you the question
- Plots the effect against calendar date and calls it exposure
- Reads a declining treatment-only curve as effect decay
- Declares decay from three descending point estimates
- Ignores that late exposure buckets hold only returning users
- Never checks whether the shape repeats for later cohorts