In an ad auction where treatment and control bid against each other, when is a budget-split test worth its cost?
answer
- bidders compete for the same impressions
- one bidder's win is another's loss
- isolate traffic and budget together
- each universe has thinner competition
- half-scale answer, not launch scale
basics
~20 sWorth it when the change moves bidding enough that the arms distort each other's win rates and prices. Splitting traffic and budget into two isolated auctions removes that interference, but each runs at reduced market scale.
solid answer
~50 sIn an auction the arms are not independent. If treatment makes a bidder bid higher, it wins impressions control bidders would have won and pushes their clearing prices up, so the control arm is actively harmed by the treatment and the contrast is inflated. A budget-split design partitions both the traffic and each advertiser's budget into two isolated auction universes so treated and control bidders never meet. That buys a clean within-universe comparison, but each universe runs at a fraction of real liquidity: fewer bidders per impression means softer competition, different clearing prices and different pacing, so the estimate describes the effect at half scale rather than at launch scale. I would spend it on changes that plausibly move equilibrium prices — bidding, pacing, reserve or ranking logic — keep ordinary splits for changes local to one advertiser's own experience, and confirm any budget-split result with a staged ramp.
go deeper
Know that ad auctions are competitive: advertisers in different experiment arms fight over the same impressions, so the two arms cannot be treated as separate independent worlds.
Explain the mechanism of the bias — treated bidders win impressions and raise prices for control bidders — and what partitioning both traffic and budget is meant to fix.
Be ready to design it: isolate traffic and budget together, verify no cross-universe bidding or pacing remains, and state precisely which quantity the resulting estimate identifies.
Own the policy: which classes of change require auction isolation, what the platform investment and lost liquidity are worth, and how a half-scale estimate becomes a defensible launch forecast.
## Why auctions are the hardest interference case An ad auction is a zero-sum allocation over a fixed inventory: for each impression, exactly one bidder wins, and the price the winner pays is determined by the other bids. That makes every advertiser's outcome a direct function of every other advertiser's behaviour. Randomly splitting advertisers into treatment and control does not create two independent experiments — it creates one auction in which some participants were handed a change and the rest have to live with the consequences. The bias is structural. If the treatment raises effective bids or improves ranking for treated advertisers, then in the same auction: - treated bidders **win more** impressions, - control bidders **win fewer** and, on the impressions they still win, often **pay more**, because the treated bids sit above them in the price ladder, - so the treated-minus-control contrast picks up both the gain and the induced loss. The measured effect is the effect of *having the change while your rivals do not*, which is not the effect of launching it to everyone. At full rollout, every bidder has it, the relative advantage disappears, and what remains is only whatever the change does to total efficiency. ## What a budget split actually isolates A budget-split design creates two parallel auction universes. Traffic is partitioned — each impression is routed to exactly one universe — and each advertiser's budget is partitioned alongside it, so the same advertiser participates in both universes with separate, non-fungible spend. The change is applied in one universe only. Because a bidder in universe A never competes for an impression with a bidder in universe B, the cross-arm distortion is gone and the within-universe comparison is a legitimate contrast between two complete auction environments. Splitting the traffic alone is not sufficient and is a classic implementation failure: if the budgets remain shared, spend exhausted in one universe throttles the same advertiser in the other, and the arms are coupled through the pacing system even though they never bid against each other. ## What it costs **Liquidity.** Each universe runs with a fraction of the bidders and a fraction of the impressions. Auction outcomes are highly non-linear in participation: with fewer competing bids per impression, competition is softer, clearing prices are lower, and the marginal bidder is different. The design therefore removes interference between the arms while changing the market being measured. The estimate is unbiased *for the half-scale world* and only approximately right for the full one — a subtler failure than cannibalization, and one that no amount of running time repairs. **Pacing behaviour.** Budget pacing algorithms behave differently against smaller budgets over the same time window. Effects that appear or vanish because pacing is operating in a different regime will not reproduce at launch. **Engineering.** Routing traffic deterministically, splitting budgets in the accounting system, and keeping the universes leak-proof is substantial work, usually a platform investment rather than an experiment configuration. That is why the decision is a policy decision, not a per-test one. ## The decision rule The question to ask about any candidate change is: **does it reallocate impressions between advertisers?** - Changes to bidding, pacing, reserve prices, ranking or eligibility do reallocate. They need isolation, or the read is inflated. - Changes to reporting, creative rendering, or an advertiser-facing tool that does not alter what gets bid do not reallocate. An ordinary split answers them well, with far more power and no platform investment. When the classification is unclear, a cheap probe helps: run the change at a small treated share and watch the **control side**. If control-arm win rates fall or control clearing prices rise as treatment exposure ramps, the arms are coupled and the ordinary split is compromised. ## The cheaper alternative when isolation is not feasible Ramp the treatment across several exposure shares and study how the estimated per-advertiser effect moves with the treated share. If the effect decays as more advertisers are treated, most of what the small-share test measured was redistribution among bidders. A stable estimate across shares is evidence of genuine incrementality. This is weaker than isolation — the ramp stages differ in more than exposure, and it takes longer — but it costs one experiment configuration rather than a partitioned marketplace, and it often settles the question well enough to make the launch call. ## Translating a half-scale answer into a forecast Even a well-run budget split leaves a translation problem: the business wants a launch number, and the design produced a number from a thinner market. The defensible practice is to treat the budget-split result as evidence about **direction and mechanism**, confirm the magnitude with a staged ramp on live traffic, and state the residual uncertainty explicitly rather than presenting the half-scale point estimate as a forecast. Owning that distinction — unbiased for what, approximate for what — is what separates a principal answer from a competent one.
- Why can a budget-split result still fail to predict the full launch?Because each universe is a smaller auction. With a fraction of the bidders chasing a fraction of the impressions, competition density, clearing prices and pacing all differ from the live market, and auction outcomes are non-linear in participation. The design removes interference between arms but changes the market being measured, so the estimate is exact for the split world and approximate for the full one.
- How would you decide whether a change needs auction isolation at all?Ask whether it reallocates impressions between advertisers. Anything touching bids, budgets, reserves, ranking or eligibility does and needs isolation. Reporting, creative rendering and advertiser tools that do not alter bidding can use an ordinary split. When unsure, ramp exposure and watch whether control-side win rates and clearing prices move.
- A team splits traffic into two universes but leaves advertiser budgets shared. What breaks?The arms stay coupled through pacing. Spend consumed in one universe drains the same budget available in the other, so a treatment that increases spend starves the control universe and the contrast is contaminated again — this time through the budget system rather than through the auction. Isolation must cover budget as well as traffic.
saying these in an interview costs you the question
- Assumes advertisers in different arms are independent
- Reports the split-universe estimate as a launch forecast
- Ignores that thinner auctions change clearing prices
- Splits traffic but leaves budgets shared
- Runs bidding changes on an ordinary split by default