skip to content

In a priority-inbox ranker, what does promoting a small random share of arriving mail give the next refresh?

level: seniorimportance: should knowfreq 45%

answer

  1. exposure the model did not choose
  2. two jobs: training rows and measurement
  3. displace the marginal promotion, not add
  4. tag the draw at decision time
  5. re-randomise each period, spread thinly

basics

~20 s

Exposure the ranker did not choose. Those messages produce opens for sender classes the live model suppresses, so the next training set carries uncensored positives, and the same slice doubles as the only measurement of interest the model has not shaped.

solid answer

~40 s

The randomised slice buys two things that are easily confused. First, **training data**: because the slot was filled without reference to the score, a demoted sender class gets seen, and its opens enter the next refresh as genuine positives rather than as the near-zero the loop would otherwise produce. Second, **measurement**: an open rate computed on randomly promoted mail is an estimate of interest the ranker did not influence, which is what makes a collapsing class readable at all. To be usable it has to be tagged in the placement log at decision time, re-randomised each period, and taken from the existing promotion slots rather than added on top — otherwise the promoted volume changes and the comparison is confounded.

go deeper

for a junior

Recall that a model can only learn from mail somebody saw, so a system that shows only what it already likes needs some exposure decided by chance.

for a middle

Separate the two jobs the random slice does — supplying uncensored training rows and supplying an honest measurement — and say why the slot has to be recorded as random when it is served.

for a senior

Argue the placement details: displace the marginal promotion rather than adding slots, re-randomise each period, spread the cost across recipients, and keep the slice from feeding the features it is meant to check.

for a principal

Own the trade the slice represents and its review cadence, including whether a stratified allocation is worth the added weighting complexity downstream.

## Why anything random is needed at all A retraining loop fitted on its own placements can only learn about mail it placed. The fix cannot come from the model, from more of its log, or from a longer history of its log, because all three were produced under the same rule. The only way new evidence enters is if some exposure is decided **without reference to the score**. That is what the randomised promotion slice is: a small, deliberate hole in the ranker's authority over what recipients see. ## The two things it buys They are usually conflated, and they have different sizing arguments: 1. **Uncensored training rows.** A sender class the ranker demotes to near-invisibility still appears in the prominent view occasionally, purely at random. Whatever recipients do with those messages is recorded, so the next refresh sees positives from a class that the closed loop would have reported as dead. 2. **An honest measurement.** Open rate on the randomised slice is an estimate of interest that the current ranker did not shape. It is the control described in any detection scheme for this failure: without it, a class's collapse in the training data cannot be told apart from recipients genuinely losing interest. A slice sized only for measurement can be far smaller than one sized to re-seed learning for a rare class, so name which job you are sizing for. ## Where the slots come from Take the random promotions **out of the existing promotion budget**, not in addition to it — typically by replacing the lowest-scoring promotions in the slate. Two reasons: - **Comparability.** If the promoted volume grows when exploration is switched on, every downstream metric moves for a second reason and the effect of exploration cannot be read. - **Cost control.** The recipient's prominent view has a fixed size in practice; expanding it to make room for random mail spends attention rather than ranking quality, and attention is the scarcer budget. Displacing the *marginal* promotion also keeps the cost at its minimum: the slot surrendered is the one the model was least sure about. ## The logging contract The slice is worthless to the next refresh unless it can be identified later, so at decision time the placement record carries: - a flag saying this promotion was random rather than scored; - the probability the slot was filled that way, since the refresh weights scored and random rows differently; - the model version that produced the slate, so a later analysis knows which ranker's censoring it is looking through. All three are cheap at write time and impossible to reconstruct afterwards — a score can be recomputed from a stored model, but the draw that happened at serving time cannot. ## Design details that decide whether it works - **Re-randomise every period.** Freezing the random choice makes it a fixed extra allowlist, which is another placement rule producing its own censored log. - **Spread it across recipients.** The same total share concentrated on a few recipients buys the same data and hands the whole cost to those few. - **Consider stratifying by sender class.** A slice drawn uniformly over all mail gives the rarest classes almost nothing, which is exactly where the loop bites hardest; allocating some of the slice per class is a pipeline decision, at the price of a more complicated weighting downstream. - **Keep the slice out of the personalisation features it is meant to test.** If a random promotion feeds an engagement feature that raises the class's score next refresh, the control has started to influence the thing it measures. ## What it does not do It does not make the rest of the log unbiased, and it does not repair the months already collected under a closed loop. It gives the next refresh evidence it would otherwise not have, and gives the operators a number they can trust. Whether the quality spent on it is worth those two things is a business call, not an engineering one.

  • Should the random promotions be extra slots or replace scored ones?
    Replace scored ones, taking the lowest-scoring promotions in the slate. Adding slots grows the promoted volume, so every downstream metric shifts for a second reason and the cost of exploration cannot be separated from it. Displacing the marginal promotion also makes the surrendered slot the one the ranker was least confident about, which is the cheapest place to spend.
  • Does the randomised slice have to be uniform over all arriving mail?
    Not necessarily. Uniform draws give the rarest sender classes almost no exposure, which is where the loop does the most damage, so allocating part of the slice per class often buys far more usable positives for the same cost. The price is that rows then arrive with different selection probabilities, which the next refresh has to account for.
  • What breaks if the random flag is not written into the placement log?
    The refresh cannot tell a random promotion from a scored one, so the uncensored rows are mixed back into the censored ones and the control disappears. The draw cannot be reconstructed later either — a score can be recomputed from a stored model version, but the coin flip that happened at serving time is gone unless it was recorded.

saying these in an interview costs you the question

  • Adds random promotions on top, inflating the promoted volume
  • Freezes one random selection and reuses it every period
  • Forgets to tag the random slot in the placement log
  • Concentrates the whole exploration cost on a few recipients
  • Claims exploration makes the rest of the logged traffic unbiased