Why does Thompson sampling keep serving an arm with only 1 success in 2 trials?
answer
- how sure, not how good
- two trials measure almost nothing
- posterior width, not posterior mean
- a wide draw can beat a tight one
- spread of 0.22 against 0.016
basics
~20 sTwo trials leave a very wide posterior, so a draw from that arm lands above the leader's draw often enough to win some requests. Uncertainty, not a tuned exploration parameter, is what buys the arm its traffic.
solid answer
~50 sBecause the arm's posterior is still enormously wide. With a flat prior, one success in two trials leaves a posterior with a standard deviation of roughly 0.22 — it plausibly covers rates from about 0.1 to about 0.9 — so a single draw from it frequently comes out above the leader's draw, and the arm wins the request. Contrast an arm sitting at 500 successes in 1,000 trials: its posterior standard deviation is about 0.016, so its draws stay in a narrow band and it essentially never loses a draw to another well-measured but clearly worse arm. That is the whole design. Exploration is spent where the uncertainty is, not spread evenly across arms, and it anneals itself: as the barely-tested arm accumulates trials its posterior narrows and its share of winning draws shrinks without anyone decaying a parameter.
go deeper
Be ready to say that few trials means a wide range of plausible rates, and that a random draw from a wide range can easily land above a well-measured competitor. That is why an untested arm still shows up in the traffic.
Explain the mechanics with numbers: roughly 0.22 of spread after two trials against roughly 0.016 after a thousand, and why the rule compares draws rather than centres. Then explain why exploration winds down without any decay parameter.
Demonstrate the operational edges — over-confident priors starving an arm, delayed rewards keeping a posterior artificially wide so an arm is over-served, and how you would sanity-check served shares against posterior overlap in a live system.
Own the prior as the real policy lever. Decide what a new arm should be assumed to be worth before it has data, how much traffic the organisation is willing to spend on unproven arms, and who signs off on that number.
## The observation Run a Thompson sampler over a few arms and watch the served shares. An arm that has been tried twice keeps picking up requests, seemingly out of proportion to anything it has demonstrated. Meanwhile an arm with a thousand trials and a mediocre rate gets almost nothing. Both behaviours come from the same source: the width of the posterior. ## Width, quantitatively The posterior is a distribution over the arm's unknown true rate, and its width shrinks roughly with the square root of the number of trials. Two concrete cases, both with a flat starting prior: - **1 success in 2 trials.** The posterior is centred at 0.5 with a standard deviation of about 0.22. Draws from it routinely land anywhere from 0.15 to 0.85. This arm knows essentially nothing about itself. - **500 successes in 1,000 trials.** The posterior is centred at 0.5 with a standard deviation of about 0.016 — roughly fourteen times narrower. Draws stay inside about 0.45 to 0.55. Both arms have the same point estimate. They behave completely differently under sampling, because the rule compares draws, not centres. ## What that does to allocation Suppose the leader's posterior is concentrated near 0.55. The 500/1,000 arm will practically never draw above it — its whole distribution sits below. The 1-of-2 arm draws above 0.55 a large fraction of the time, so it wins requests regularly. From the outside this looks like the sampler favouring an unproven arm; in fact it is the sampler correctly saying "I cannot yet rule out that this arm is the best one, and the only way to find out is to serve it". The symmetric case is instructive. If two arms have posteriors centred at the same value but very different widths, neither is favoured — by symmetry each wins about half the draws. Width does not by itself attract traffic; width attracts traffic *when the arm is behind*, because width is what puts posterior mass above the leader. ## Self-annealing The important consequence is that exploration decays on its own. Every request the uncertain arm wins produces an observation, which narrows the posterior, which reduces the chance it wins the next draw. If the arm really is mediocre, a few dozen trials shrink it out of contention. If the arm really is excellent, the same trials pull its posterior up and it takes over. Either way, no schedule is written by hand and no parameter needs decaying — the feedback loop between traffic and certainty does the annealing. This is what people mean when they say Thompson sampling has no exploration knob. The knob exists, but it is the prior: a strong, confident prior on an arm makes its posterior narrow before any data arrives, and it will be under-explored. A weak prior leaves it wide and it will be explored generously. Prior specification is the real lever. ## Where the intuition breaks **Confident priors starve arms.** If you initialise a new arm with a prior equivalent to hundreds of pseudo-observations at a low rate, its posterior is narrow from the start, it never wins draws, and it never gets the data to correct itself. The self-correcting loop requires that the arm's posterior overlap the leader's. **Delayed rewards distort width.** The posterior only narrows when rewards land. If an arm has served thousands of impressions whose outcomes have not yet arrived, its posterior is still wide and it keeps winning draws, so it gets served far more than intended before the evidence catches up. **Width is not goodness.** A weak candidate answer is "the sampler prefers uncertain arms". It does not prefer them; it gives them a hearing proportional to how plausible it is that they are best. An uncertain arm whose posterior lies entirely below the leader's — possible if the prior was informative — gets nothing. ## The interview-ready summary Thompson sampling allocates traffic to an arm in proportion to the posterior probability that the arm is the best one. A barely-tested arm has a wide posterior, so that probability is meaningfully above zero and it earns traffic. A heavily-tested but clearly worse arm has a narrow posterior sitting below the leader, so that probability is near zero and it gets squeezed out. Exploration is bought with uncertainty and paid off with data.
- What happens to that arm's exploration share as it accumulates trials?It shrinks on its own. Every request it wins produces an observation, the posterior narrows, and the chance its draw beats the leader's falls. If the arm is genuinely mediocre it is squeezed out within a few dozen trials; if it is genuinely good its posterior climbs and it takes over. The annealing is a feedback loop, not a schedule.
- Two arms have the same posterior mean but very different widths. Which gets more traffic?Neither — with symmetric posteriors centred at the same value, each wins about half the draws. Width does not attract traffic by itself; it attracts traffic when the arm is behind, because width is what puts posterior mass above the current leader.
- Can an arm be starved of exploration under Thompson sampling?Yes, if its posterior is narrow and low from the outset. That happens when the prior is over-confident — equivalent to hundreds of pseudo-observations at a poor rate — so its draws never cross the leader's and it never earns the data that would correct it. A weak prior on a new arm is the safeguard.
A heavily measured mediocre arm has nothing left to claim about itself. A barely tested arm can still credibly claim it is brilliant, and Thompson sampling gives that claim a hearing in proportion to how plausible it is.
saying these in an interview costs you the question
- Says the sampler prefers uncertain arms for their own sake
- Thinks exploration must be forced with a fixed random traffic share
- Believes a narrow posterior means the arm is good
- Assumes exploration needs a decay schedule to wind down
- Ignores that an over-confident prior can starve an arm entirely