Reweighting a priority inbox's next training set by logged promotion probability fixes part of the censoring — which part survives?
answer
- weights rescale, they do not create
- zero exposure leaves nothing to weight
- tiny probability, enormous influence
- deterministic top-k records no probability
- random slice supplies the floor
basics
~20 sThe part with no exposure at all. Weighting can re-inflate sender classes that were shown rarely, but a class the ranker gave essentially no chance of promotion contributes no rows, and no weight applied to nothing produces evidence.
solid answer
~40 sWeighting by the recorded selection probability corrects **under-representation**: a class promoted one time in a thousand has its surviving rows counted as though far more of them had been seen. What it cannot do is manufacture rows. A class whose promotion probability was effectively zero leaves nothing in the log to weight, so the correction is silent exactly where the loop has bitten hardest. Two more limits follow: rows with tiny probabilities carry enormous weights, so a handful of messages can dominate the refit, and a strictly deterministic top-k rule has no probability to record in the first place. That is why reweighting and randomised exposure are complementary — the weights handle classes that were shown rarely, and the random slice handles classes that stopped being shown at all.
go deeper
Know that counting a rarely seen message more heavily is a way to compensate for it being rarely seen, and that it only works if it was seen at all.
Explain the requirement behind the correction: every class needs some chance of exposure, and a strictly deterministic placement rule leaves no probability to weight by.
Show where it fails in production — invisible zero-exposure classes, a refit dominated by a handful of high-weight rows — and pair it with the randomised slice that supplies the exposure floor.
Decide whether weights are capped and what that bias costs, and insist the refresh report raw per-class counts so absence is visible rather than averaged away.
## What the weight is doing Every promotion decision can be recorded with the probability that the slot was filled the way it was. When the next refresh reads the log, a row that had a one-in-a-thousand chance of being observed is counted as though it stood for many unobserved ones — the idea behind **inverse propensity scoring**. Applied to a priority inbox, it partly undoes the fact that the ranker, not chance, decided which mail the recipient ever saw. Rare exposures stop being drowned out by the classes the ranker promotes constantly. That is a real correction, and it is cheap once the probability is recorded. It is also strictly weaker than it looks. ## The requirement the weighting cannot manufacture The correction rests on every class having had **some** chance of exposure. Where that chance is effectively zero, the log contains no rows for the class, and a weight multiplies something that is not there. The consequence is uncomfortable in practice: - The classes with the largest weights are the ones nearest the edge of being dropped entirely. - The classes the loop has already suppressed completely are invisible to the correction, silently. - The report of the correction looks the same either way — a weighted dataset does not announce which classes are missing from it. ## Variance, and why a handful of messages can take over A row observed with probability 0.001 carries a thousand times the influence of one observed with probability 1. If a suppressed class survives in the log through a dozen such rows, the refit's belief about that whole class rests on those dozen messages and on whatever idiosyncrasies they carry. The correction is unbiased in expectation and can still be wildly unstable in the single sample you actually have. Capping the weights trades that instability back for bias, which is a defensible choice as long as it is a stated one. ## A deterministic rule records no probability If the placement rule is "promote the top eight scores", the probability of promotion given the slate is 1 for those eight and 0 for the rest. There is nothing meaningful to weight by: the log contains certainties, not draws. A system that intends to weight has to introduce the randomness that makes a probability exist — which is the same randomised slice that fixes the coverage problem. The two remedies turn out to need the same mechanism in the serving path. ## What each remedy covers | Situation in the log | Reweighting | A randomised promotion slice | |---|---|---| | class shown rarely but shown | corrects the under-representation | adds more observations of it | | class shown essentially never | no rows to weight, no effect | the only source of observations | | deterministic placement rule | no probability exists to use | creates the probability | | labels themselves are wrong | out of scope | out of scope | The last row is worth saying aloud: neither remedy touches what an open *means*. They change which mail is represented, not whether the recorded outcome was the right target. ## How to use both Run the randomised slice to guarantee every class you care about has a floor on its exposure probability, then weight the scored rows so the rare-but-real exposures are not swamped by the constantly promoted classes. Record the probability at decision time along with the model version, because it cannot be reconstructed later — a score can be recomputed from a stored model, but the draw that happened at serving time is gone unless it was written down. And keep reporting raw per-class exposure counts beside the weighted numbers, so a class that has fallen out of the data entirely is visible as a zero rather than hidden inside a weighted average.
- Why do very small logged probabilities destabilise the next refit?Because influence scales inversely with them. A row seen with probability 0.001 counts a thousand times as heavily as one seen with certainty, so a suppressed class represented by a dozen such rows has its entire learned behaviour decided by those dozen messages. Capping the weights steadies the refit at the price of some bias, which is an acceptable trade when it is recorded as a deliberate choice.
- What should the refresh report alongside the weighted training set?Raw per-class exposure and positive counts. A weighted average hides absence — a class with no rows contributes nothing and produces no warning — while a raw count shows it as a zero. Reporting both is what makes a collapsed class visible at the moment the refresh runs, rather than months later when a recipient complains.
saying these in an interview costs you the question
- Believes weighting recovers a class with no logged exposures
- Ignores that tiny probabilities give a few rows enormous influence
- Assumes a deterministic top-k rule has a probability to log
- Treats reweighting as a substitute for randomised exposure
- Expects the weights to fix what the recorded outcome means