skip to content

Validity Threats

The failure modes that make a clean-looking result false: stopping early, a traffic split that never was 50/50, effects that fade, and users who influence each other across arms.

on this pageshow

explore

questions

20

What are novelty and primacy effects in an A/B test?

level: juniorimportance: must knowfreq 74%

answer

  1. reaction to change, not to the feature
  2. one inflates early, one depresses early
  3. both fade as users habituate
  4. control arm has nothing to react to
  5. week-one number is not the steady state

basics

~20 s

Both are temporary reactions to the fact that something changed, not to what it changed into. Novelty inflates an early result as users explore something new; primacy, or change aversion, depresses it as users unlearn a habit.

solid answer

~50 s

A novelty effect is a short-lived boost in the treatment arm: users notice the change, poke at it, and generate extra clicks or sessions out of curiosity rather than sustained value, so the early lift overstates the steady state. A primacy effect, normally called change aversion, is the mirror image: users with an established habit are slowed down while they relearn where things are, so the early result understates the steady state. Both live only in the treatment arm, because the control arm carries no change to react to, and both fade as each user accumulates exposure. That makes the measured difference a mixture of a transient reaction and the real effect. The practical consequence: a lift or a drop seen in the first days of exposure is not a safe estimate of what the feature will be worth once users have settled.

go deeper

for a junior

Be ready to define both terms in one sentence each and to say which direction each pushes the measured effect. Knowing that they are temporary and that they sit in the treatment arm is enough at this level.

for a middle

Expect to explain the mechanism: why the transient does not cancel in a treatment-minus-control difference, and why the effect must be tracked on each user's exposure clock rather than the calendar to be seen at all.

for a senior

Show that you separate what is broken from what is not. Randomization is intact; the horizon is wrong. Demonstrate that you would check the pattern against its uncertainty before declaring decay rather than eyeballing a shrinking number.

for a principal

Own the framing problem. Novelty and change aversion are the two explanations most easily reached for to explain away an inconvenient result in either direction, so define what evidence your organisation accepts before either label is allowed to change a decision.

Novelty and primacy are transient reactions to *the fact that something changed*, not to what it changed into. They matter because an experiment measures the difference between arms during the window in which it runs, and if that window is dominated by a transient reaction, the number a team ships on is not the number it will live with. ## Novelty A novelty effect is a temporary lift in the treatment arm caused by newness itself. Users notice that something is different, explore the new surface, click the unfamiliar control to see what it does, and briefly increase whatever engagement metric is being tracked. None of that behaviour reflects durable value; it reflects curiosity. The characteristic signature is an effect that starts large and shrinks as each individual user's exposure accumulates: big on a user's first day of exposure, smaller on their fifth, smaller again on their twentieth. ## Primacy, or change aversion A primacy effect is the mirror image: a temporary depression in the treatment arm caused by disruption of an existing habit. The term in product work is change aversion. Users of a long-established tool have built muscle memory around where things sit. A redesign forces them to relearn the layout, and during that relearning period they are slower and make more mistakes even when the new arrangement is better once learned. The signature is an effect that starts negative and improves with exposure. The special case where the curve keeps climbing past the old baseline, because the new design is genuinely better once mastered, is called a learning effect: a keyboard shortcut that saves time only after users habituate to reaching for it behaves exactly this way. ## Why they are a validity threat It helps to be precise about what is broken. Randomization is intact; the two arms are still comparable; the comparison is an unbiased estimate of the difference *during the observed window*. What fails is the extrapolation from that window to the steady state — a threat to the temporal generalizability of the result, not to the internal comparison. A weak candidate treats novelty as though it were a bucketing or logging failure; it is not. The arms are fine. The horizon is wrong. ## Why the transient sits in one arm Control users see the product they already knew. There is nothing new to explore and no habit to unlearn, so no transient. Treatment users get both the feature and the disruption of receiving it. Because the estimate is treatment minus control, the transient does not cancel — it is added to, or subtracted from, the durable effect. ## Who shows which The two effects usually fall on different populations. Someone who has never used the product before has no prior layout to unlearn, so change aversion has nothing to bite on; they evaluate the new design on its own merits. Heavy, long-tenured users are the opposite: they have the most invested habit and pay the largest relearning cost. That is why a redesign can show a positive effect for people arriving for the first time and a negative one for daily users of the same product, and why the pooled number depends on how those two groups are mixed in the traffic. ## What it looks like in data Both effects are defined on each user's own clock — days since that user first saw the change — not on the calendar. A single-number readout hides them completely. A curve of the treatment-minus-control gap against exposure age reveals them: monotone decay toward a smaller value for novelty, monotone recovery from a deficit for change aversion. ## What it is not A shrinking point estimate is not automatically decay. Estimates wobble, and later exposure buckets carry fewer users and wider intervals, so a lift that falls from +8% to +5% may be nothing but noise. Decay is a claim that needs the pattern to be consistent, to appear for successive groups of users as they enter, and to be large relative to the uncertainty. Equally, a stable effect is not proof of no transient — it can be a novelty boost and a real gain of similar size overlapping in a short window. ## What can be done about it after the fact Three readouts do most of the work. Read the effect restricted to users who have had the change for a while, rather than pooling everyone. Read the segment with no prior habit separately from the tenured segment, since they are answering different questions. And read a long-term holdback, where a slice of users has still never received the feature, to see whether the launch gain is still there months later. Each of these targets the same underlying question: what remains once the reaction to the change itself has worn off.

  • Which of the two would you expect from a redesigned navigation bar in a long-established tool?
    Predominantly change aversion. Daily users have muscle memory for the old bar and pay a relearning cost, so the early effect for them is likely negative and should recover as they adapt. Users arriving for the first time have no habit to unlearn and may show the opposite sign, so the pooled early number depends on the traffic mix between the two groups.
  • Do novelty effects threaten the internal validity of the comparison?
    No. Assignment is still random and the arms are still comparable, so the estimate is a valid measurement of the difference during the observed window. What fails is generalising that window to the steady state, which is a temporal or external validity problem. Saying the test is invalid overstates it; saying the horizon is wrong is accurate.
  • Can a transient reaction show up in the control arm?
    Only if control users also experienced a change. In a standard test the control is the unchanged status quo, so there is nothing novel to explore and no habit to unlearn, and the transient sits entirely in treatment. If both arms were altered, each carries its own transient and they partially cancel in the difference.

Moving the furniture in a room you know well: for a week you bump into things and hate it, and a visitor who has never been in the room thinks it looks great. Neither reaction tells you whether the new layout is better to live in.

saying these in an interview costs you the question

  • Labels every early lift a novelty effect without evidence
  • Assumes novelty always inflates and never depresses a result
  • Treats change aversion as proof the new design is fine
  • Calls novelty a randomization or logging failure
  • Reads the first-week lift as the long-run effect
  • Confuses a noisy point estimate with genuine decay

context

open as a page

On day 2 of a planned 14-day A/B test the dashboard shows p = 0.04 — do you ship?

level: juniorimportance: must knowfreq 66%

basics

~20 s

No. The 0.05 threshold is calibrated for a single analysis at the planned sample size, so a day-2 reading is not the test you designed. Run to the agreed horizon unless the plan already allowed an early call.

open as a page

What is a sample ratio mismatch in an A/B test?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A sample ratio mismatch is when the observed split of users across experiment arms differs from the planned split by more than chance can explain. It signals a broken assignment or logging pipeline, so the comparison itself is untrustworthy.

open as a page

In a marketplace test where treated sellers win demand from control sellers, which way is the effect biased?

level: middleimportance: must knowfreq 66%

basics

~20 s

Upward. Treated sellers capture bookings that control sellers would otherwise have received, so treatment is pushed up while control is pushed down. The measured lift counts one transfer twice and mixes real growth with redistribution that vanishes at full rollout.

open as a page

How do you check whether an A/B test effect is decaying over exposure time?

level: middleimportance: must knowfreq 56%

basics

~20 s

Put every user on their own clock: days since that user first saw the change. Plot the treatment-minus-control gap against exposure age, grouped by first-exposure date. A gap that shrinks with exposure age, repeating per group, indicates decay.

open as a page

Why does stopping an A/B test as soon as its p-value drops below 0.05 inflate Type I error?

level: middleimportance: must knowfreq 74%

basics

~20 s

A fixed-horizon p-value is calibrated for one analysis at one pre-set sample size. Each extra look gives noise another chance to cross the threshold, so stopping at the first p < 0.05 rejects far more than 5% of null tests.

open as a page

A 50/50 A/B test logged 100,000 control and 98,500 treatment users — is that a sample ratio mismatch?

level: middleimportance: must knowfreq 62%

basics

~10 s

Yes. Against a 50/50 plan those counts give a chi-square goodness-of-fit statistic near 11.3 on one degree of freedom, a p-value around 0.0008. Far too extreme for chance, so treat the experiment as invalid.

open as a page

Why does randomizing communities in a social graph, rather than users, reduce spillover bias?

level: seniorimportance: must knowfreq 52%

basics

~20 s

Clustering puts most of a user's connections in the same arm, so treatment rarely crosses the arm boundary. Residual contamination is confined to the edges cut between clusters, and a good partition keeps that fraction small.

open as a page

In a social feed A/B test, why can users in the control arm be affected by the treatment?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Users influence each other. If treated users post or share more, their control-arm friends see that extra content, so the control arm is partly treated and the measured gap between arms understates the feature's true effect.

open as a page

When is a switchback design, alternating treatment on time slices, better than a user-level test?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Use it when interference is system-wide — a shared pool of drivers or inventory that no user-level split can separate. Alternating the whole market between conditions on short slices makes the time slice the randomized unit.

open as a page

A 5% long-term holdback six months after ship shows no gain — how do you interpret it?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Do not read it as proof the launch gain was fake. Check the interval first: a 5-versus-95 split is precision-limited by the small arm, so it may still contain the original launch effect. Then rule out leakage.

open as a page

Your nav redesign is up for new users and down for returning power users — what do you conclude?

level: seniorimportance: should knowfreq 50%

basics

~20 s

That split is the signature of change aversion: users with no prior habit judge the design on its merits, while users with muscle memory pay a relearning cost. Confirm it by checking whether the returning-user deficit shrinks with exposure.

open as a page

How would you design an experiment results dashboard so it does not invite peeking?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Hide the verdict until the planned end date. Show progress toward the target sample size, the decision date and health diagnostics, but withhold p-values and significance badges until the test is due, then present the decision as read.

open as a page

How do you find the root cause of a sample ratio mismatch in a running A/B test?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Localise the loss. Recount users at each pipeline stage — assignment, exposure, then any post-hoc filters — and slice by platform, browser, device, country and day. The first stage and segment where the ratio breaks names the bug.

open as a page

An executive asks 'is it significant yet?' on day 2 of a planned 14-day test — how do you handle it?

level: principalimportance: should knowfreq 36%

basics

~20 s

Answer the decision behind the question, not the number. Say what is honestly known now, give the date the result arrives, and treat recurring demands for early reads as a planning problem to fix before the next launch.

open as a page

How would you set sample ratio mismatch policy for a platform running hundreds of experiments?

level: principalimportance: should knowfreq 34%

basics

~20 s

Make the check automatic and blocking: every experiment is tested on arm counts, a flagged test hides its metric readout, and clearing it requires a named, documented cause. Set the threshold strict enough that alerts stay believable at portfolio scale.

open as a page

If you stop an A/B test at the first significant look, what happens to the measured effect size?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

It is biased upward. Stopping when the statistic crosses the threshold selects the moments when noise happened to help, so the effect you report is systematically larger than the truth — and the earlier you stop, the worse the exaggeration.

open as a page

A sample ratio mismatch alert fired mid-test but arm counts matched by the end — do you trust the result?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Not automatically. Late-arriving logs from one arm can create a mismatch that resolves on its own. But a defect that started and stopped leaves corrupted data behind. Establish which you saw before reading any metric.

open as a page

In an ad auction where treatment and control bid against each other, when is a budget-split test worth its cost?

level: principalimportance: nice to knowfreq 24%

basics

~20 s

Worth it when the change moves bidding enough that the arms distort each other's win rates and prices. Splitting traffic and budget into two isolated auctions removes that interference, but each runs at reduced market scale.

open as a page

How would you set a team policy for change-aversion claims that override a negative A/B result?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Require the claim to be falsifiable before it can change a decision: a written mechanism, a predicted signature in the data, and a pre-agreed evidence bar with a scheduled long-term re-read. Otherwise it becomes an argument that is always available to whoever lost.

open as a page