skip to content

Bandit Feedback

Feedback replaces labels — a bandit only sees the reward of the arm it pulled, a contextual bandit adds features, full RL adds state and delayed reward. Interviewers test the line between them.

on this pageshow

questions

3

How does bandit feedback differ from the labels a supervised learner sees?

level: middleimportance: must knowfreq 55%

answer

  1. one action per round, one outcome
  2. the options not taken stay unknown
  3. no loss defined over all actions
  4. your own policy writes the dataset

basics

~20 s

Bandit feedback reveals the outcome only for the action actually taken; a supervised label states the correct answer regardless of what the model did. Showing one of six headlines teaches you nothing about the other five.

solid answer

~50 s

Under full supervision every example arrives with the answer attached, so you can score every option you might have chosen and compute a loss over all of them. Under bandit feedback you observe the outcome of the one action you took and nothing else — the reward is partial, and the counterfactual is gone for good. A news homepage that shows one of six headlines per visit learns whether that headline was clicked, and never learns what the other five would have earned from that visitor. Two consequences follow. First, the log is written by the system's own choices, so it is not a random sample of action-outcome pairs and is skewed toward whatever the current policy already favours. Second, the system has to sometimes take actions it does not currently believe are best, or it will never collect the evidence that could change its mind.

go deeper

for a junior

Be ready to state the core fact plainly: you see the outcome of the action you took and nothing about the alternatives. Recognise the products that have this shape — feeds, ads, recommendations, any slot where one option is served per request.

for a middle

Explain why a standard classification loss cannot be formed: there is no target over all actions, only one observed reward. Say clearly that the log is produced by the running policy and is therefore not a random sample of action-outcome pairs.

for a senior

Show what this forces in the pipeline: logging context, served action and reward as one record, keeping a slice of traffic whose action is independent of the model, and treating offline accuracy on logged actions as weak evidence for a policy that would act differently.

for a principal

Own the call of whether a surface should be instrumented for bandit feedback at all. The price is deliberately serving some traffic suboptimally plus real logging discipline; the payoff is ever being able to compare a new policy against the incumbent on evidence rather than opinion.

## Two feedback regimes In ordinary supervised learning, every training example arrives with the answer attached. For a six-way classification problem you are told the true class, which means you can score all six candidate answers and compute a loss — cross-entropy, say — that is defined over *every* option, whether or not your model would have picked it. This is called **full-information feedback**: for each example, the payoff of every action is knowable. **Bandit feedback** is the weaker regime. Each round the world presents a situation, your system picks exactly one action out of a set, and the world returns a single number — the reward — for that action alone. The payoffs of the actions you did not take are never revealed, not late, not in aggregate, not ever. The name comes from the multi-armed bandit framing, where each option is one arm of a slot machine and pulling one arm tells you nothing about the others. ## The concrete shape of it Take a news homepage with one hero slot and six candidate headlines. On each visit the system serves one headline and records: the visitor context, which headline was served, and whether it was clicked. What is missing from that record is what any of the other five would have done **for that same visitor**. Scale does not repair this. After a billion visits you have click estimates for the headlines the policy served often, thin estimates for the ones it served rarely, and nothing at all for a headline it stopped serving on day one. ## Three consequences that matter in practice **1. There is no loss over all actions.** You cannot form a target vector across the six headlines and minimise the distance to it, because five of the six entries are unknown. Learning reduces to estimating a reward function `r(context, action)` from the (context, action, reward) triples you happened to log, then acting greedily or near-greedily on it. **2. The dataset is caused by the model.** In a supervised dataset the examples come from the world; here, which action gets an observation is decided by the policy that was running. If the old policy showed the sports headline mostly to visitors arriving from a sports referrer, then the sports headline's observed click rate mixes headline quality with audience composition. Comparing raw observed rates across actions is therefore a confounded comparison, not a measurement of the actions. **3. Exploration becomes a data-collection requirement, not a nicety.** An action that is never taken never gets an observation, so its estimate never updates. A policy that always serves what currently looks best can lock onto a mediocre option and remain self-confirming indefinitely, because the evidence that would refute it is exactly the evidence it declines to collect. ## Partial is not the same as noisy A common muddle is to call bandit feedback *noisy labels*. Noise means the observation you do get is a random draw around a true value: one visitor's click is a Bernoulli sample, not the headline's true rate. Missingness means there is no observation at all. More traffic averages away the noise on the actions you serve; it never fills in the actions you do not serve. Both problems exist at once, but only the second one defines the regime. It is also not the same as ordinary missing labels. In a normal dataset with gaps, the gaps arise from causes outside the model and each example still has one true label. Under bandit feedback each (context, action) pair has its own outcome, exactly one per round is revealed, and the choice of which one is made by the system itself — which is why imputation alone cannot rescue it. ## A ladder of feedback strength It helps to see the regimes as a ladder. Strongest is full information: every action's payoff is revealed. Next is bandit feedback: only the chosen action's payoff. Weaker still are settings where the reward is a comparison rather than a value, or an aggregate over many rounds. There is also an asymmetric case worth recognising: sometimes every action produces *some* observation (any headline shown is clicked or not), while in other problems one particular action produces no observation at all — a rejected loan applicant never repays and never defaults. The one-sided case is strictly harder, because the missing evidence is concentrated exactly where the policy is most confident. ## What a practitioner does about it Log the context, the action served, and the reward together as one record, and keep the identity of the action served rather than only the eventual winner. Keep a slice of traffic whose action is chosen independently of the model, so some evidence exists that the policy did not shape. And treat offline accuracy at predicting the reward of the logged action as weak evidence about a candidate policy that would take *different* actions, since that policy's behaviour lands in exactly the region where the log is thinnest.

  • Is bandit feedback a noise problem or a missingness problem?
    Primarily missingness, though noise is present too. The reward you do observe is a noisy draw — one visitor's click is not the headline's true rate — and more traffic averages that out. But for the actions you did not take there is no observation at all, noisy or otherwise. Traffic fixes the noise; nothing but taking the action fixes the missingness.
  • How does this differ from a supervised dataset that simply has missing labels?
    In an ordinary dataset each example has one true label and the gaps come from causes outside the model. Under bandit feedback each context has a separate outcome per candidate action, exactly one is revealed per round, and which one is revealed is chosen by the system itself. The missingness is generated by the model's own behaviour, so imputing from neighbours quietly copies the policy's bias.
  • Can logged bandit data be turned into a supervised training set at all?
    Partly. Predicting the reward of the action that was taken, given the context, is honest supervised learning on the triples you have. What you cannot do is treat the action taken as the correct label — it was never shown to be best, it was merely what the policy chose. Scoring all actions from such a model extrapolates into regions where the log holds no evidence.

It is like ordering one dish from a menu of six. You learn how that dish tasted, and you never find out about the other five, no matter how many times you come back and order the same one.

saying these in an interview costs you the question

  • Calls it ordinary supervised learning with fewer labels
  • Treats the action that was taken as the correct label
  • Believes more traffic eventually reveals the unshown options
  • Confuses partial reward with label noise
  • Compares raw observed rates across actions as if unconfounded

context

open as a page

Why can't a loan-approval model learn anything from the applicants it rejected?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Repayment is observed only for approved applicants, so rejected ones carry no outcome to learn from. The training set is the slice the previous policy approved, and nothing in it says whether a rejection would have repaid.

open as a page

How do you decide whether in-app upsell placement is a contextual bandit or full reinforcement learning?

level: principalimportance: nice to knowfreq 32%

basics

~20 s

Ask whether the action changes the situation the next decision faces. If each impression pays its own reward and the next starts from the same state, it is a contextual bandit; only real carry-over justifies reinforcement learning.

open as a page