What are novelty and primacy effects in an A/B test?
answer
- reaction to change, not to the feature
- one inflates early, one depresses early
- both fade as users habituate
- control arm has nothing to react to
- week-one number is not the steady state
basics
~20 sBoth are temporary reactions to the fact that something changed, not to what it changed into. Novelty inflates an early result as users explore something new; primacy, or change aversion, depresses it as users unlearn a habit.
solid answer
~50 sA novelty effect is a short-lived boost in the treatment arm: users notice the change, poke at it, and generate extra clicks or sessions out of curiosity rather than sustained value, so the early lift overstates the steady state. A primacy effect, normally called change aversion, is the mirror image: users with an established habit are slowed down while they relearn where things are, so the early result understates the steady state. Both live only in the treatment arm, because the control arm carries no change to react to, and both fade as each user accumulates exposure. That makes the measured difference a mixture of a transient reaction and the real effect. The practical consequence: a lift or a drop seen in the first days of exposure is not a safe estimate of what the feature will be worth once users have settled.
go deeper
Be ready to define both terms in one sentence each and to say which direction each pushes the measured effect. Knowing that they are temporary and that they sit in the treatment arm is enough at this level.
Expect to explain the mechanism: why the transient does not cancel in a treatment-minus-control difference, and why the effect must be tracked on each user's exposure clock rather than the calendar to be seen at all.
Show that you separate what is broken from what is not. Randomization is intact; the horizon is wrong. Demonstrate that you would check the pattern against its uncertainty before declaring decay rather than eyeballing a shrinking number.
Own the framing problem. Novelty and change aversion are the two explanations most easily reached for to explain away an inconvenient result in either direction, so define what evidence your organisation accepts before either label is allowed to change a decision.
Novelty and primacy are transient reactions to *the fact that something changed*, not to what it changed into. They matter because an experiment measures the difference between arms during the window in which it runs, and if that window is dominated by a transient reaction, the number a team ships on is not the number it will live with. ## Novelty A novelty effect is a temporary lift in the treatment arm caused by newness itself. Users notice that something is different, explore the new surface, click the unfamiliar control to see what it does, and briefly increase whatever engagement metric is being tracked. None of that behaviour reflects durable value; it reflects curiosity. The characteristic signature is an effect that starts large and shrinks as each individual user's exposure accumulates: big on a user's first day of exposure, smaller on their fifth, smaller again on their twentieth. ## Primacy, or change aversion A primacy effect is the mirror image: a temporary depression in the treatment arm caused by disruption of an existing habit. The term in product work is change aversion. Users of a long-established tool have built muscle memory around where things sit. A redesign forces them to relearn the layout, and during that relearning period they are slower and make more mistakes even when the new arrangement is better once learned. The signature is an effect that starts negative and improves with exposure. The special case where the curve keeps climbing past the old baseline, because the new design is genuinely better once mastered, is called a learning effect: a keyboard shortcut that saves time only after users habituate to reaching for it behaves exactly this way. ## Why they are a validity threat It helps to be precise about what is broken. Randomization is intact; the two arms are still comparable; the comparison is an unbiased estimate of the difference *during the observed window*. What fails is the extrapolation from that window to the steady state — a threat to the temporal generalizability of the result, not to the internal comparison. A weak candidate treats novelty as though it were a bucketing or logging failure; it is not. The arms are fine. The horizon is wrong. ## Why the transient sits in one arm Control users see the product they already knew. There is nothing new to explore and no habit to unlearn, so no transient. Treatment users get both the feature and the disruption of receiving it. Because the estimate is treatment minus control, the transient does not cancel — it is added to, or subtracted from, the durable effect. ## Who shows which The two effects usually fall on different populations. Someone who has never used the product before has no prior layout to unlearn, so change aversion has nothing to bite on; they evaluate the new design on its own merits. Heavy, long-tenured users are the opposite: they have the most invested habit and pay the largest relearning cost. That is why a redesign can show a positive effect for people arriving for the first time and a negative one for daily users of the same product, and why the pooled number depends on how those two groups are mixed in the traffic. ## What it looks like in data Both effects are defined on each user's own clock — days since that user first saw the change — not on the calendar. A single-number readout hides them completely. A curve of the treatment-minus-control gap against exposure age reveals them: monotone decay toward a smaller value for novelty, monotone recovery from a deficit for change aversion. ## What it is not A shrinking point estimate is not automatically decay. Estimates wobble, and later exposure buckets carry fewer users and wider intervals, so a lift that falls from +8% to +5% may be nothing but noise. Decay is a claim that needs the pattern to be consistent, to appear for successive groups of users as they enter, and to be large relative to the uncertainty. Equally, a stable effect is not proof of no transient — it can be a novelty boost and a real gain of similar size overlapping in a short window. ## What can be done about it after the fact Three readouts do most of the work. Read the effect restricted to users who have had the change for a while, rather than pooling everyone. Read the segment with no prior habit separately from the tenured segment, since they are answering different questions. And read a long-term holdback, where a slice of users has still never received the feature, to see whether the launch gain is still there months later. Each of these targets the same underlying question: what remains once the reaction to the change itself has worn off.
- Which of the two would you expect from a redesigned navigation bar in a long-established tool?Predominantly change aversion. Daily users have muscle memory for the old bar and pay a relearning cost, so the early effect for them is likely negative and should recover as they adapt. Users arriving for the first time have no habit to unlearn and may show the opposite sign, so the pooled early number depends on the traffic mix between the two groups.
- Do novelty effects threaten the internal validity of the comparison?No. Assignment is still random and the arms are still comparable, so the estimate is a valid measurement of the difference during the observed window. What fails is generalising that window to the steady state, which is a temporal or external validity problem. Saying the test is invalid overstates it; saying the horizon is wrong is accurate.
- Can a transient reaction show up in the control arm?Only if control users also experienced a change. In a standard test the control is the unchanged status quo, so there is nothing novel to explore and no habit to unlearn, and the transient sits entirely in treatment. If both arms were altered, each carries its own transient and they partially cancel in the difference.
Moving the furniture in a room you know well: for a week you bump into things and hate it, and a visitor who has never been in the room thinks it looks great. Neither reaction tells you whether the new layout is better to live in.
saying these in an interview costs you the question
- Labels every early lift a novelty effect without evidence
- Assumes novelty always inflates and never depresses a result
- Treats change aversion as proof the new design is fine
- Calls novelty a randomization or logging failure
- Reads the first-week lift as the long-run effect
- Confuses a noisy point estimate with genuine decay