skip to content

How do you keep Thompson sampling responsive when an arm's true rate changes over time?

level: seniorimportance: should knowfreq 45%

answer

  1. Bayesian does not mean adaptive
  2. a hundred thousand observations will not budge
  3. the changed arm barely gets served
  4. forget on purpose
  5. effective memory of one over one-minus-gamma

basics

~20 s

A long-running Thompson sampler locks in because its posteriors become extremely tight and the changed arm barely gets served. Make it forget on purpose: discount old observations each round, or keep only a sliding window of recent ones.

solid answer

~50 s

The failure mode first: after months of traffic, each arm's posterior is built from tens of thousands of observations and is correspondingly narrow. If a promotion's true click rate genuinely jumps in December, that arm's posterior no longer covers the new value, its draws still lose, and it barely gets served — so the handful of new observations that would correct it never arrive. The sampler is locked in. The fix is deliberate forgetting. Either discount the accumulated counts every round by a factor `gamma` slightly below one before adding new data, which caps an arm's effective memory at roughly `1 / (1 - gamma)` observations, or keep a sliding window of the last W observations per arm. Both keep posteriors permanently wide enough that a changed arm can win draws again and re-explore. The cost is real: in a genuinely stationary world you are discarding information, so you explore forever and pay extra regret.

go deeper

for a junior

Know that a sampler which has seen a lot of data becomes very confident, and that confidence is what stops it noticing a change. The idea of deliberately forgetting old data is the takeaway.

for a middle

Explain the two-part lock-in — tight posteriors plus almost no traffic to the changed arm — and describe discounting and sliding windows as mechanisms, including what the discount factor means in terms of effective sample size.

for a senior

Show operational judgment: how you size the forgetting to the real timescale of change, how you validate it on replayed logs, how you tell a genuine rate shift from a tracking bug, and how you cap how fast allocation is allowed to move.

for a principal

Own the staleness-versus-variance trade as a policy. Decide how much permanent regret the organisation should pay for adaptability, whether seasonality is handled by discounting or by scheduled resets, and what monitoring must exist before an allocator is allowed to swing traffic unattended.

## Why a mature sampler stops adapting Thompson sampling's convergence is also its trap. The posterior width for an arm shrinks roughly like one over the square root of its observation count. An arm with 100,000 observations has a posterior so tight that its plausible range is a fraction of a percentage point wide. That is exactly what you wanted while the world was stationary — the sampler stopped wasting traffic and settled on the leader. Now suppose the world moves. A seasonal promotion whose true click rate genuinely rises in December is a canonical case: yesterday the arm was worth 2%, today it is worth 6%. Two things now conspire: 1. The arm's posterior is concentrated near 2% and puts essentially no mass near 6%, so its draws still lose to the leader. 2. Because its draws lose, it is served almost never — so the new, higher-converting observations that would move the posterior arrive at a trickle. Even when they do arrive, each new observation is one voice against tens of thousands of old ones. The update is swamped. This is lock-in, and it is not a bug in the implementation; it is what an estimator that treats all history as equally relevant is supposed to do. ## Deliberate forgetting The cure is to stop treating a year-old observation as evidence about today. **Discounting.** Each round, before folding in new data, multiply each arm's accumulated counts by a discount factor `gamma` slightly below one, pulling them back toward the prior. Old observations then decay geometrically. In steady state the effective number of observations an arm retains is about `1 / (1 - gamma)`: with `gamma = 0.99` an arm remembers roughly the last 100 observations, with `gamma = 0.999` roughly the last 1,000. The posterior width stops shrinking past the width implied by that effective count, so the arm can always be re-explored. Applying the discount to *all* arms each round, not only the played one, is what lets an unserved arm drift back toward the prior and start winning draws again — which is precisely the re-exploration you need. **Sliding window.** Simpler and easier to explain: keep only the last W observations per arm and rebuild the posterior from those. Sharper forgetting than geometric discounting, at the cost of a hard boundary and more storage. **Periodic reset.** Blunt but sometimes right: at a known boundary — a seasonal changeover, a catalogue refresh — reset every arm's posterior toward the prior. Use it when you *know* the world changed rather than trying to detect it. **Change detection plus reset.** Monitor recent reward rates per arm against what the posterior predicts, and reset only the arm that breaks. This adapts fastest to abrupt shifts but adds a detector with its own false-alarm rate. ## Choosing how much to forget The discount factor is a bet on how fast the world changes, expressed as an effective memory in observations. Pick it from the timescale over which reward rates are actually stable in your system, converted into per-arm observation counts: if rates hold roughly constant for a week and a typical arm sees 5,000 impressions a week, an effective memory of a few thousand observations is defensible. The honest way to set it is to replay historical logs at several values and compare accumulated reward, rather than picking a round number. ## The cost when nothing changes Forgetting is not free. In a genuinely stationary environment, discounting throws away information you paid for. Posteriors stay permanently wider than the data justifies, the sampler keeps sending exploratory traffic to arms it has already resolved, and regret accumulates without bound instead of flattening. Aggressive discounting also makes the allocation noisy: served shares swing around on short-run luck, which looks alarming on a dashboard and can trigger unnecessary intervention. So the choice is a variance-versus-staleness trade: forget quickly and you track change fast but chase noise; forget slowly and you are stable but late. There is no setting that avoids both. ## Operational notes worth raising - **A wide prior at launch does not help.** It only affects the first few observations; months later the posterior is tight regardless. Non-stationarity has to be handled continuously, not at initialisation. - **Distinguish a real shift from a measurement break.** A click-tracking regression looks exactly like a rate change to the sampler, and discounting will happily adapt to the broken numbers. Instrument health checks belong alongside the sampler. - **Bound the allocation velocity.** If the sampler can swing 80% of traffic overnight in response to a discounted posterior, a transient glitch becomes an outage. Rate-limiting how fast served shares can move is a cheap guardrail. ## The one-line answer A stationary Thompson sampler converges and then cannot un-converge; add geometric discounting or a sliding window so posteriors keep a bounded effective sample size, sized to the timescale on which your rates genuinely move.

  • How would you pick the discount factor?
    Translate it into an effective memory: a per-round discount of `gamma` retains about `1 / (1 - gamma)` observations per arm. Choose that number from the timescale over which your reward rates are genuinely stable, expressed in impressions per arm, then validate by replaying historical logs at several candidate values and comparing accumulated reward. Do not pick a round number by feel.
  • What is the cost of discounting if the arms are actually stationary?
    You are throwing away information you paid for. Posteriors stay wider than the data justifies, so the sampler keeps sending exploratory traffic to arms it has already resolved and regret keeps accumulating instead of flattening. Allocation also gets noisier, with served shares swinging on short-run luck.
  • Would a wide prior at launch protect against a later change?
    No. The prior only matters while observations are scarce; after months of traffic the posterior is dominated by data and is narrow no matter what it started as. Non-stationarity is a continuous property of the environment and needs a continuous mechanism — discounting, a window or resets — not a one-time initialisation choice.

A posterior built from a hundred thousand observations is a witness so confident that a few new statements cannot cross-examine it. Discounting makes the witness forget last year on purpose, so today's evidence can still change the verdict.

saying these in an interview costs you the question

  • Assumes posteriors adapt automatically because the method is Bayesian
  • Thinks a wide prior at launch handles later change
  • Restarts the entire experiment whenever any metric moves
  • Discounts so aggressively the sampler never settles on anything
  • Cannot say what the discount factor means in observations

context