skip to content

Why does a bandit allocator updating hourly lock onto a stale winner when conversions land three days later?

level: seniorimportance: nice to knowfreq 30%

answer

  1. exposures count now, conversions later
  2. measured rate depends on exposure age
  3. the arm promoted most recently looks worst
  4. update cadence versus conversion window
  5. promotion generates its own demotion evidence

basics

~20 s

Because exposures count immediately while conversions arrive days later, any recently promoted arm looks artificially bad at update time. The early leader, whose conversions have matured, keeps winning the comparison and reinforces its own head start.

solid answer

~50 s

The reward the allocator reads is censored: the denominator (exposures) fills instantly, the numerator (conversions) fills over three days. So an arm's measured rate depends on the *age* of its exposures, not just its quality. An arm whose share just rose is dominated by fresh, not-yet-converted exposures and scores low; the arm that led early has mature exposures and scores high. The allocator promotes the high scorer, which ages its data further, while the demoted arms never get old enough traffic to recover — a self-reinforcing loop that locks in whichever arm happened to lead first. Fixes, in order of simplicity: score only exposure cohorts older than the conversion window; update no faster than that window; keep a minimum allocation floor so starved arms can recover; use a fast, correlated proxy reward; or drop adaptivity for this metric entirely.

go deeper

for a junior

Be ready to spot the timing problem: a conversion counted three days after the exposure cannot appear in the number the allocator reads this hour.

for a middle

Explain the numerator-denominator mismatch — exposures count immediately, conversions arrive later — and why that makes any recently promoted arm score worse than it truly is.

for a senior

Diagnose it from data using rate by exposure age and allocation share against conversion arrival, then fix it with matured cohorts, a slower cadence and an exploration floor.

for a principal

Decide which metrics are allowed to drive adaptive allocation at all given their maturation time, and what proxy signals may stand in for a lagging outcome.

## The mismatch that causes it An allocator scores each arm with something like `conversions observed / exposures served`. Those two quantities fill on different clocks. An exposure is recorded the instant the user is served. A conversion may land minutes, hours or days later. At any moment, the conversions you can see for an arm are only those from exposures old enough to have converted — the recent exposures are **censored**, contributing to the denominator with a numerator still in flight. That makes the measured rate a function of the **age distribution of an arm's exposures**, which the allocator never intended to be part of the comparison. An arm whose exposures are mostly a week old shows nearly its full rate. An arm whose exposures are mostly from the last hour shows nearly zero. ## Why it becomes a lock-in rather than just noise Censoring alone would be an annoyance. The feedback loop is what makes it pathological: 1. Some arm leads early, for real reasons or by luck. 2. Traffic shifts toward it. Its share rises, but crucially, its *existing* exposures keep ageing and converting, so its measured rate stays healthy. 3. Any arm the allocator boosts next receives a burst of fresh exposures. Its measured rate is immediately dragged down by that burst, because the conversions from it have not landed. 4. The allocator reads the drop as evidence the arm is worse and pulls its traffic back — before the burst had time to convert. 5. Those conversions eventually land, but the arm now has so little traffic that they barely move a rate computed over its full history. The system therefore *punishes any arm it has just promoted*, which is precisely the wrong incentive, and it keeps the incumbent whether or not the incumbent is best. Notice the direction: the pathology is not that the allocator explores too little in general; it is that promotion itself generates the evidence used to demote. ## How to detect it in a live run - **Plot measured rate by exposure age.** Bucket each arm's exposures by how long ago they occurred and compute the conversion rate per bucket. If the rate rises steadily with age and only flattens after roughly three days, the metric is censored and any score computed over all exposures is age-dependent. - **Overlay allocation share against conversion arrival.** Lock-in shows as an arm's traffic share moving sharply *before* the conversions from the preceding period have arrived. - **Check whether demotions follow promotions.** A signature pattern is an arm's share spiking and collapsing within a window shorter than the conversion lag. - **Compare to a uniform holdout.** If a uniformly allocated slice ranks the arms differently from the adaptive traffic, the allocator is being steered by something other than arm quality. ## Fixes, cheapest first **Score matured cohorts only.** Include an exposure in the score only once it is older than the conversion window — for a three-day window, score exposures from at least three days ago and ignore everything newer. Every arm is then judged on fully observed data and the comparison is like-for-like. The cost is reaction speed: the allocator is always acting on three-day-old evidence. **Match the update cadence to the lag.** Updating hourly on a three-day metric is the core error. If the evidence takes three days to mature, allocation cannot meaningfully change more often than that without steering on censorship. **Model the lag curve.** If you know what fraction of an arm's eventual conversions have typically arrived by age `t`, you can inflate the observed count by the reciprocal of that fraction to estimate the eventual rate. This preserves fast reaction but adds a modelling assumption — and if lag genuinely differs across arms (a variant that attracts slower-deciding users), the correction itself introduces bias. **Use a fast proxy reward.** Optimise a same-session signal that correlates well with the delayed outcome, and validate the correlation separately. The risk is optimising the proxy at the expense of the real outcome, so this needs a periodic check against the delayed metric. **Keep an exploration floor.** Guarantee every arm a minimum traffic share regardless of score. This does not fix the censoring, but it prevents permanent lock-in: a wrongly demoted arm keeps receiving enough traffic that its matured evidence can eventually overturn the decision. **Or do not run adaptively at all.** When the conversion window is a large fraction of the decision horizon, the allocator can only ever act on censored data, and the regret it saves is small compared with the risk of converging on the wrong arm. A fixed allocation with a single readout after the metric matures is both simpler and more likely to be right. ## The general lesson Adaptive allocation assumes the reward it reads is an unbiased current picture of arm quality. Any process that makes the observed reward depend on how long an arm has been serving — delayed conversions, seasonality, novelty effects, gradual metric maturation — violates that assumption, and the violation is amplified rather than averaged away, because the allocator acts on it and thereby changes the data it next sees.

  • What is the simplest fix that requires no model of the lag?
    Score only exposure cohorts older than the conversion window, and update allocation no faster than that window. Every arm is then compared on fully matured data. You give up reaction speed — the allocator always acts on evidence that is at least a conversion-window old — but you add no assumptions and nothing can be biased by censoring.
  • How would you confirm this diagnosis from the logged data?
    Bucket each arm's exposures by age and plot conversion rate per bucket. A rate that climbs with age and only flattens near the known lag proves the metric is censored. Then overlay allocation share against conversion arrival times: lock-in appears as share moving sharply before the corresponding conversions have landed.
  • When should you abandon adaptive allocation for this metric entirely?
    When the conversion window is a large fraction of the decision horizon. The allocator can then only act on censored evidence, so the regret it saves is small against the risk of converging on the wrong arm. A fixed allocation with one readout after the metric matures is simpler and safer.
  • Why can an exploration floor rescue the run even without fixing censoring?
    A floor guarantees every arm a minimum traffic share regardless of its current score, so a wrongly demoted arm keeps accumulating exposures. Those exposures eventually mature and can overturn the decision. It converts a permanent lock-in into a slow correction, and it also keeps assignment probabilities bounded for later reweighting.

It is like judging new hires on revenue closed this week: whoever started earliest always looks best, and every new starter looks like a mistake.

saying these in an interview costs you the question

  • Blames random noise instead of censored conversions
  • Speeds up the update cadence to react faster
  • Assumes conversion lag is identical across all arms
  • Treats the observed rate at exposure time as final
  • Adds arms instead of investigating the reward's maturity

context