A priority-inbox ranker is refit each month on the mail recipients opened — why does each refresh harden its demotion of a sender class?
answer
- the log is not a sample
- exposure decides what gets labelled
- demotion suppresses opens, not interest
- positives per class shrink each refresh
- next model writes the next training set
basics
~20 sDemoted mail is rarely seen, so it is rarely opened, so the next training set holds almost no positives for that sender class. Each refresh learns a lower score for it, which suppresses exposure further, and the loop tightens.
solid answer
~40 sThe training set is not a sample of recipient preference; it is a sample of preference filtered through last month's placement decision. An open requires an exposure, and exposure was chosen by the ranker. Once a sender class is demoted, its exposures collapse, its opens collapse with them, and the refresh sees that class represented almost entirely by unopened messages labelled negative. The refit therefore scores it lower than the model that produced the log, and the next month's evidence is generated by that even more confident belief. The difference from an ordinary blind spot is that this one **grows each cycle**: the model writes the data its successor is fitted on. Only exposure the ranker did not choose — a randomised promotion slice — puts uncensored evidence back into the loop.
go deeper
Recall that a model can only be judged on mail it was shown. An unopened message sitting in a view the recipient rarely scrolls is weak evidence of disinterest, not a clean negative.
Be able to walk the loop in order: demotion cuts exposure, lost exposure cuts opens, the refit reads missing opens as negatives, the next score is lower. Name where in the pipeline the censoring enters.
Show how you would catch it from pipeline artifacts — compare the per-class composition of successive training snapshots, and distrust an offline number computed on a split carved from the same censored traffic.
The call you own is whether the product surrenders some promotion quality now to keep the training set usable later, and how that cost is expressed to the people who own the inbox's headline metric.
## The training set is a photograph of the model's own choices A priority-inbox ranker scores each arriving message and promotes a small share of them into a prominent view; the rest land in a secondary view that recipients scroll less often. A month later the refresh is assembled: a message the recipient opened becomes a positive, a message they did not becomes a negative. That looks like a sample of recipient preference. It is a sample of preference **filtered through last month's placement decision**. Opening a message requires seeing it, and seeing it was decided by the ranker. So the assembled training set answers a narrower question than the one the model is asked at serving time: not *would this recipient want this message*, but *did this recipient open this message, given that the previous ranker chose to show it prominently*. ## Why the second refresh is worse than the first The loop turns once per refresh, and it compounds: 1. The ranker assigns a low score to a sender class — shipping notifications, say — for whatever reason it had in month one: thin history, an unlucky feature value, a scoring quirk. 2. Mail from that class is demoted, so far fewer recipients look at it. 3. Opens from the class collapse — not because interest fell, but because the placement that produces opens was withdrawn. 4. The refresh assembles month two's data and finds the class represented almost entirely by unopened messages, labelled negative. 5. The refit learns a lower score for the class than month one's model held, and step 2 runs harder. The distinction worth stating in a design review is between a **static blind spot** and a **tightening one**. A model fitted once on censored data has a fixed hole. A model that produces the log its own successor is fitted on has a hole that widens every cycle, because each refresh's evidence was generated by the previous refresh's beliefs. Nothing inside the pipeline pushes back on it. ## The mirror on the promoted side The same mechanism runs with the sign flipped. A sender class promoted in month one accumulates exposures, therefore opens, therefore positives, therefore a higher score at the next refit, therefore more promotions. Two sender classes of genuinely equal interest can end up several refreshes later with a wide score gap that records which one was promoted first rather than which one the recipient prefers — the accumulation the **Matthew effect** names. This side is easier to miss because the headline metric looks healthy: the promoted view's open rate is high, precisely because it is measured on the mail the ranker was already most sure about. ## What moves and what does not | Quantity | Across successive refreshes | Why | |---|---|---| | recipient interest in the demoted class | stable, in the scenario as described | nothing about the recipient changed | | exposures logged for that class | falls sharply after the first demotion | placement decides exposure | | positives in the next training set | falls with exposure, approaching zero | an open needs an exposure first | | the class's learned score | falls further at each refit | the class is now almost all negatives | | model confidence on that class | tends to rise | it is scored on data it selected | | the offline evaluation number | can rise while live quality is flat | the held-out set is censored the same way | That last row is the trap. If the evaluation split is carved out of the same logged traffic, it inherits the same censoring, so the refresh reports an improvement on exactly the population the model has stopped doubting. ## What does not break the cycle - **More data.** Ten times the volume of the same censored log is ten times as much evidence that demoted mail goes unopened. - **A longer training window.** Older months were produced by earlier models in the same loop; the history is censored too, just less tightly. - **A larger model or richer features.** Capacity fits the censored distribution more sharply, which usually makes the demotion crisper, not softer. - **A faster refresh.** Turning the loop more often tightens it sooner; cadence changes the speed, not the direction. - **Freezing the ranker for a while.** Freezing stops further hardening, but the log still contains only what the frozen model chose to show, so the training set does not repair itself. ## What does Evidence the current ranker did not select. In practice that is a small randomised share of the promotion slots, filled without reference to the score, so some mail from suppressed classes is seen and its opens are recorded. Alongside it, recording the selection probability at decision time lets the next refresh re-weight classes that were under-exposed rather than absent. A third, weaker signal is a rescue action — a recipient digging a message out of the secondary view — which produces an occasional positive from a class the ranker had written off, though far too few of them to re-seed a training set on their own.
- Does slowing the monthly refresh to quarterly break the cycle?No. Cadence changes how fast the belief hardens, not whether it hardens. Between refreshes the live ranker still decides what gets exposed, so the quarter's log is censored by the same placement rule — a slower loop simply arrives at the same narrowed training set later. A slower cadence also delays the moment you can measure the problem from the refit at all.
- What is happening on the promoted side of the same loop?The mirror spiral. Promoted senders keep earning exposures, so they keep earning opens and positives, so each refit scores them higher and promotes them more. Two classes of equal true interest can separate purely on which one was promoted first. It is easy to miss because the promoted view's open rate — the metric most dashboards show — looks excellent while it happens.
- The refresh reports a better offline number every month. Is that reassuring?Not by itself. If the evaluation split comes from the same logged traffic, it carries the same censoring, so the number measures agreement with the previous model's choices. A reading that means something has to come from exposure the ranker did not choose — a randomised promotion slice — or from outcomes on mail the recipient reached some other way.
A kitchen that restocks only what sold yesterday, having hidden one dish at the back of the menu. Each week the hidden dish sells less, so it is stocked less, and by month three the sales report proves nobody ever wanted it.
saying these in an interview costs you the question
- Treats an unopened demoted message as a confirmed true negative
- Blames concept drift when recipient behaviour is stipulated unchanged
- Expects more history to restore a class whose history is censored too
- Assumes a bigger model or richer features will notice the gap
- Says freezing the ranker for a month repairs the training data
- Reads a rising offline number as evidence the loop is healthy