Your permanent sponsored-ranking holdout shows +3.1% while the year's shipped launches summed to +5.6% - why the gap?
answer
- short readouts are not additive
- novelty fades, the launch stays
- you only shipped the positive draws
- six launches, the same three slots
- the holdout gap is a floor
basics
~20 sShort launch readouts are not additive. Each was measured while the change was still novel, only the launches that read positive were kept, and launches touching the same slots cannot both claim their full effect. The holdout measures what actually survived together.
solid answer
~50 sThe sum is the wrong arithmetic in three ways at once. Each launch was read in its first weeks, when part of the lift is a reaction to change that decays; you shipped the ones that read positive, and a selected set of positive readings is biased upward; and six changes competing for the same three slots overlap, so their effects cannot be laid end to end - even compounding them gives about +5.7%, barely different. The holdout is the only number measured against a population that received none of them, all at once, after the novelty is gone. Push the other way too: the holdout itself understates, because sellers and inventory are shared between the arms and any improvement the frozen policy cannot re-decide leaks in. The practical output is a discount factor you apply to future launch readouts, not an accusation of fraud.
code
pseudocode · 18 lineslaunch_deltas = [0.012, 0.008, 0.005, 0.020, 0.004, 0.007]
naive_sum = 0.0
compounded = 1.0
for d in launch_deltas:
naive_sum = naive_sum + d
compounded = compounded * (1.0 + d)
// naive_sum = 0.056 -> +5.6%
// compounded - 1 = 0.0572 -> +5.7%
holdout_delta = 0.031 // launched population vs permanent holdout, this week
unexplained = (compounded - 1.0) - holdout_delta
// 0.0572 - 0.031 = 0.0262 -> about 2.6 points claimed but not visible
retention = holdout_delta / (compounded - 1.0)
// 0.031 / 0.0572 = 0.54 -> roughly half of the short-run readouts survivedgo deeper
Two weeks of measurement right after a change is not the same as its effect a year later, and six such measurements cannot simply be added together. The holdout is the arm that never got any of them.
Be able to name the three main inflators - novelty in the early window, selecting the launches that read positive, and overlap between changes touching the same slots - and say which way each pushes.
Show that you know the holdout has its own bias in the other direction through shared sellers and inventory, and turn the finding into a discount factor plus a reverse test for any specific launch you want to interrogate.
This is a planning and credibility question. Decide what the organisation commits to when summed claims and the holdout disagree, who owns the discount factor, and how launch reviews are worded so the gap is expected rather than scandalous.
## The two numbers and what each was measured against Six ranking launches shipped to the sponsored strip over a year, each read in a two-week window against the traffic not yet receiving it. Their measured deltas were +1.2%, +0.8%, +0.5%, +2.0%, +0.4% and +0.7% on the launch metric - a sum of **+5.6%**, or **+5.7%** compounded. Meanwhile 1% of users were permanently excluded from every one of them, and this week that holdout reads **+3.1%** for the launched population. About **2.6 points** of claimed progress do not appear. This is one of the most common findings in any organisation that keeps a long-term holdout, and the answer is not that anybody lied. ## Why the sum is the wrong arithmetic | cause | direction | why | |---|---|---| | novelty decay | inflates each launch readout | early response to a changed strip fades as it becomes the norm | | selection of what shipped | inflates the set of readouts | launches that read flat or negative were dropped, so the kept set skews high | | overlap between launches | inflates the sum | two changes improving the same three slots cannot both capture the same headroom | | a moving baseline | either direction | each launch was read against the world of its own quarter, not against a year-ago world | | holdout contamination | deflates the holdout | shared sellers, inventory and page-level changes reach the held-out users too | The first three are the bulk of it in most systems, and the first one is usually the largest. ## Where the missing points went - **Novelty and habit.** A user who has seen the same strip for months has a settled response to it. A change interrupts that, and the interruption itself produces clicks that are not sustained interest. Two weeks is inside the interruption. - **Headroom is shared.** Three slots can only be so well filled. The fourth launch of the year is competing with the improvements the first three already captured, so the honest question per launch is not what it added in isolation but what it added on top of what was already live. - **Cannibalised gains.** A launch that moves clicks from organic results into the strip reads as a strip win and a smaller whole-page win, and the holdout only ever sees the whole page. - **The kept set.** If launches are shipped when the readout looks good, the shipped population is a set of high draws. Nothing dishonest happens; the selection is doing the inflating. - **Slow-moving costs.** Fatigue with a denser or more aggressive strip accumulates over months and is invisible to any two-week window. It shows up only as the holdout gap widening over time. ## What the holdout is actually telling you The holdout answers one question that nothing else can: **what is the marketplace's current outcome versus a world with none of this year's launches in it, right now, after the novelty is gone.** It is a single number covering the whole programme. Crucially, it cannot attribute: it never says which launch was overstated, so it is a programme-level instrument, not an audit of an individual change. It also has its own bias, and in the opposite direction. Sellers adapt their listings and bids to the ranking the other 99% see, the retrieval index is shared, and page-level changes ship to everyone - so some of what the launches did reaches the holdout too, and the gap the holdout reports is a lower bound on the true programme effect. A mature answer states both biases and does not pretend the holdout is truth. ## What to do with the number 1. **Derive a discount factor,** not a verdict. If the programme repeatedly retains roughly half of its summed short-run readouts, plan future launches with that haircut applied and say so out loud in review. 2. **Re-read individual launches at a longer horizon.** A reverse test - turning one shipped change back off for a slice for a few weeks - gives a decayed reading for that specific change, which the holdout cannot give. 3. **Keep the accounting honest going forward.** Record for each launch what it was measured against, how long the window was, and which slots it touched, so overlap is visible before it is summed. 4. **Report the holdout as a range,** acknowledging that contamination makes it a floor. The instinct to defend in a design round: the gap is expected, it is measurable, and a team that cannot produce it is not measuring its own progress.
- Does the holdout tell you which of the six launches was overstated?No. It is one number for the whole programme against a population that missed all of it, so it cannot attribute. To interrogate a single launch you need a reverse test: turn that change back off for a slice for a few weeks and read the difference after the novelty has already gone.
- Which direction does holdout contamination push the gap?It narrows the measured gap, so the holdout understates the programme's true effect. Sellers, listings, bids and the retrieval index are shared between the arms, and page-level changes ship to everyone, so part of what the launches did reaches held-out users. Report the holdout reading as a floor.
- Two launches both claimed a lift on the same three slots. How should they have been accounted for?Each should have been read against the world as it stood when it launched, with the earlier change already live, and the overlap recorded. Sequential measurement is the honest version; what you cannot do afterwards is lay two independently measured deltas end to end and call the total progress.
A household adds up six separate deals it switched to during the year and claims a large saving, but the actual annual bill fell by much less. Each deal was priced against a different month, some of them overlapped on the same expense, and the first month of each was unusually cheap.
saying these in an interview costs you the question
- Adds two-week launch readouts to get annual progress
- Calls the gap evidence that a team faked its results
- Thinks the holdout can attribute the shortfall to one launch
- Believes novelty inflates the long-run effect, not the early one
- Treats the holdout reading as unbiased truth
- Assumes overlapping launches capture separate headroom