skip to content

In a marketplace test where treated sellers win demand from control sellers, which way is the effect biased?

level: middleimportance: must knowfreq 66%

answer

  1. buyers are a roughly fixed pool
  2. one seller's gain is another's loss
  3. control arm is actively depressed
  4. the same transfer counted twice
  5. per-seller lift, flat totals

basics

~20 s

Upward. Treated sellers capture bookings that control sellers would otherwise have received, so treatment is pushed up while control is pushed down. The measured lift counts one transfer twice and mixes real growth with redistribution that vanishes at full rollout.

solid answer

~50 s

In a two-sided marketplace, buyer demand is close to fixed in the short run, so a feature that makes treated listings more attractive largely moves bookings from control listings to treated ones. Treatment rises, control falls, and the difference between the arms counts the same transfer twice, so the estimate overstates what launching to every seller would deliver — sometimes by a large multiple. At 100% rollout nobody has a relative advantage any more, and the true global effect is only whatever incremental demand the feature genuinely creates, which can be close to zero. The diagnostic tell is a big per-seller lift alongside flat marketplace totals: bookings per treated seller up, total bookings unchanged. Containment means putting competing sellers in the same condition — randomizing whole markets or time slices instead of individual sellers — rather than correcting the number afterwards.

go deeper

for a junior

Know that in a marketplace both arms compete for the same buyers, so a treated seller's extra bookings may be a control seller's lost bookings rather than new business for the platform.

for a middle

Explain why the treated-minus-control difference double-counts a transfer, and why the same feature can look strong at a 1% treated share and flat at full rollout.

for a senior

Be ready to diagnose it live: reconcile per-seller lift against marketplace totals, read the estimate across ramp stages, and choose a containment design before any forecast is committed.

for a principal

Own which launches earn an expensive market-level or time-sliced readout and which can ship on a cheap biased test, and defend the number that goes to the business as a forecast rather than a lift.

## Why a marketplace test is not a clean comparison A marketplace has a shared pool that both arms draw from: buyers. When you randomize sellers into treatment and control and give the treated group something that makes their listings more attractive — better placement, a badge, a faster response tool — you have not created two independent worlds. You have created **one auction for a roughly fixed amount of demand, in which one group was handed an advantage**. The treated sellers' gain is, at least in part, the control sellers' loss. This is interference in its most financially consequential form, because the bias points in the direction everyone wants to believe. ## The arithmetic of double counting Write the observed contrast as the treated arm's average outcome minus the control arm's. Suppose the feature genuinely creates a small amount of new demand, and additionally shifts an amount D of existing demand per treated seller away from control sellers. Then: - the treated average is inflated by D, - the control average is deflated by roughly the same D scaled by the arms' relative sizes, - and the **difference** picks up both, so the transfer enters the estimate twice. With a 50/50 split and pure redistribution — no new demand at all — the true global effect is zero while the measured contrast is comfortably positive and statistically significant. The experiment is internally valid: within this split, treated sellers really did do better. It just does not answer the launch question, which is what happens when *every* seller has the feature and no one is relatively advantaged. ## The scale dependence that gives it away Cannibalized estimates depend on the **treated share of the market**, which an honest effect does not. Ramp the same feature at 1%, 10% and 50% of sellers: pure redistribution shows a large per-seller lift at 1% (a few advantaged sellers raiding a large untreated pool) that shrinks as the treated share grows and there are fewer control sellers left to take demand from. A genuinely incremental feature shows a roughly stable per-seller effect across shares. Watching the estimate collapse as the ramp proceeds is one of the most reliable field diagnostics available, and it is also why a 1% test can be spectacularly misleading. ## Reading totals against arms The second diagnostic is a reconciliation. Look at three numbers over the same window: bookings per treated seller, bookings per control seller, and total marketplace bookings. - Treated up, control **down**, total **flat** — redistribution. The lift is a transfer. - Treated up, control flat, total **up** by about the treated arm's gain — incremental. The lift is real. - Treated up, control **up**, total up more than the arms' gap suggests — positive spillover, and the arm contrast now *understates* the effect. This happens when the feature draws additional buyers onto the platform who then transact with everyone, including control sellers. The third case matters because it shows the bias is not always upward. Substitution between units inflates the contrast; complementarity between units deflates it. "Marketplace test, therefore overstated" is a heuristic, not a law — the mechanism decides. ## Fixing it by design No analysis of a seller-level split recovers the global effect, because the control arm was damaged by the treatment and there is nothing left in the data that shows what an undamaged control looks like. You need a design in which competing sellers share a condition, so that the comparison is between whole competitive environments rather than between rivals inside one: - **Market-level assignment.** Whole markets are treated or control, so every seller competing for the same buyers has the same condition. Costly in precision: the independent units are markets, and there are rarely many of them. - **Time-slicing.** The entire marketplace alternates between conditions over successive windows, so each window contains a complete competitive equilibrium. Suited to changes whose effect appears within a short window, and it inherits carryover between adjacent windows. - **Fixed-supply-side holdouts read on totals.** Rather than reading a per-seller contrast, read marketplace-level totals against a permanently held-out portion of the market — but only if the held-out portion does not itself compete with the treated one. Each costs power, and that is the honest trade: a biased estimate with tight intervals versus an unbiased estimate with wide ones. The senior move is to notice that the tight interval is worthless if it is centred on the wrong number. ## What weak answers look like Weak candidates report the seller-level lift as the launch forecast, cite randomization as proof of validity, or propose adjusting for seller size — all of which miss that the mechanism is competition for a shared resource and happens strictly after assignment. Strong candidates ask, unprompted, what happened to the marketplace total.

  • What single diagnostic would tell you the lift is redistribution rather than growth?
    Reconcile arm-level metrics against the marketplace total over the same window. If bookings per treated seller are up, bookings per control seller are down by a comparable amount, and total bookings are flat, the feature is moving demand rather than creating it. Genuine growth lifts the total.
  • Does cannibalization always bias the estimate upward?
    No. Substitution between units inflates the contrast, but complementarity deflates it. If the feature draws extra buyers onto the platform who also transact with control sellers, the control arm rises and the contrast understates the true effect. The mechanism decides the direction, not the setting.
  • How would a ramp across traffic shares expose the problem?
    Redistribution shrinks as the treated share grows, because there are progressively fewer untreated sellers left to take demand from. Estimating the same effect at 1%, 10% and 50% treated and watching it decay is strong evidence of transfer; a stable estimate across shares points to genuine incrementality.

saying these in an interview costs you the question

  • Reports the per-seller lift as the launch forecast
  • Says randomization guarantees a valid marketplace estimate
  • Ignores that short-run buyer demand is roughly fixed
  • Confuses relative advantage with value created
  • Proposes covariate adjustment to remove the transfer

context