A nightly ad job keeps every click and one non-click in a hundred — what must ship with that dataset to keep scores calibrated?
answer
- sampling changes the base rate
- thinning one class multiplies the odds
- ranking survives, calibration does not
- carry the rate with the rows
- weight of w on each kept negative
basics
~20 sThe negative sampling rate, recorded per stratum and per build. Thinning non-clicks a hundredfold multiplies the odds by about a hundred, and without that factor nobody can convert the model's scores back into real click probabilities.
solid answer
~40 sDownsampling the negatives changes the base rate, and the base rate is what a calibrated score reproduces. Keeping every click and 1% of non-clicks multiplies the odds of a click by about 100, so the raw predicted rate is roughly 100 times the true one. The correction is arithmetic and equivalent either way: carry a **per-row weight** of 100 on each kept non-click and 1 on each click so the weighted loss sees the original population, or train unweighted and rescale the predicted odds by dividing by the sampling factor. What must not happen is that the factor is lost. Ranking survives sampling untouched — it is any absolute number read off the score, such as expected spend per impression, that is wrong by that factor.
go deeper
The takeaway: if you throw away most of one class, the remaining data shows a different mix than reality, and the model predicts that different mix unless something puts the ratio back.
Be able to do the arithmetic both ways — a weight equal to the inverse sampling rate on kept negatives, or dividing the predicted odds by the same factor afterwards — and say which parts of the dataset description must record it.
Demonstrate what breaks downstream when the factor is lost: expected-value bidding, pacing and any fixed decision threshold, while rank metrics keep looking healthy and hide the problem for weeks.
The design choice is where the correction lives — baked into the artefact through weights, or applied at the boundary — and which teams inherit the obligation to know about it when several datasets with different rates feed one model.
## Why the assembly job samples at all At a billion impressions a day and a click rate near 0.1%, a faithful daily training set is a billion rows, of which about a million are positives. Keeping every click and one non-click in a hundred yields about 1,000,000 positives and about 9,990,000 kept negatives — roughly 11 million rows, a **91-fold shrink** — while discarding almost none of the rare signal. That is the whole bargain: negatives are abundant and highly redundant, positives are scarce and each one is informative. ## What the sampling does to the base rate Sampling the negative class at rate `1/w` multiplies the **odds** of a positive by `w`: - true rate `0.001` → true odds `0.001001` - after keeping all positives and one negative in a hundred, the sampled rate is `1 / 10.99 ≈ 0.091`, about **91 times** the true rate The model fits what it is shown, so it learns to predict roughly nine per cent, not one per thousand. Nothing about this is a bug — it is the arithmetic consequence of the sampling, and it is fully reversible **provided the factor is known**. ## Two equivalent corrections | correction | where it applies | what it changes | when to prefer it | |---|---|---|---| | per-row weight `w` on kept negatives | training time | the weighted loss reproduces the population's class balance, so the fitted scores come out already on the true scale | the trainer supports sample weights and you want one artefact with no post-processing | | rescale the predicted odds by `1/w` | scoring time | the model stays fitted on the sampled distribution; every prediction is mapped back afterwards | the trainer ignores weights, or several datasets with different rates feed one model | The odds rescaling is exact: with `p_s` the score on the sampled distribution and `w` the negative sampling factor, `p_true = p_s / (p_s + w * (1 - p_s))`. Substituting the worked numbers, `0.091 / (0.091 + 100 * 0.909) ≈ 0.001`, recovering the true rate. ## What the dataset has to carry The factor is not a property of the model; it is a property of the **rows**. It therefore belongs with them: 1. **The per-stratum sampling rate for this build.** If negatives are sampled at different rates by placement or by country, one global number is wrong and each stratum's rate must travel separately. 2. **The per-row weight**, materialised as a column rather than recomputed downstream. A consumer that does not know the rates cannot reconstruct them from the rows. 3. **The true population counts** the sample was drawn from, so anyone can check that the weighted totals reproduce them. A dataset that lost its sampling rates is not repairable by inspection: the sampled file looks like a perfectly ordinary set with a 9% click rate, and nothing in it announces that the real rate is a hundred times smaller. ## What is affected and what is not - **Unaffected:** the ranking of one impression against another, and therefore any rank-based metric. Uniform thinning of one class is a monotone transformation of the odds, so the order of scores is preserved. - **Affected:** anything that reads the score as a probability. An expected-value calculation that multiplies a bid by a predicted click rate is off by the sampling factor; pacing that spends against predicted clicks overspends by the same factor; a fixed decision threshold means something entirely different on the two scales. - **Also affected:** loss values and calibration measurements. A calibration curve drawn on the sampled validation set says nothing about production unless the same correction is applied to it. ## The related mistake worth naming Duplicating positives until the classes balance is not the same operation. It adds no information about the negative class, distorts the base rate further, and — because the duplicated rows are identical — tends to make the model more confident on exactly the rows it already fitted. Downsampling the abundant class keeps every distinct positive and drops redundant negatives, which is why assembly jobs at this scale reach for it first.
- Why is negative downsampling usually a better trade than duplicating positives to balance the classes?Because the negatives are the redundant class. Dropping ninety-nine in a hundred of them costs very little information and shrinks the job enormously, while every distinct positive is kept. Duplicating positives adds no new information, pushes the base rate even further from reality, and repeats identical rows the model then fits harder. Both distort the base rate, but only one of them also buys a large reduction in build cost.
- Negatives are sampled at 1% in one country and 10% in another. What changes?One global factor is now wrong for both. Each stratum needs its own rate recorded and its own per-row weight, or the corrected predictions will be biased in opposite directions by country. Any aggregate computed from the set — the overall click rate, a per-slice calibration curve — must be computed with the weights applied; the unweighted file no longer represents any real population at all.
A pollster who interviews every lottery winner and one passer-by in a hundred can still report the true win rate — but only if the two sampling rates are written on the form. Lose the form and the survey reads as though winning were common.
saying these in an interview costs you the question
- Thinks downsampling negatives needs no correction because the order of scores is unchanged.
- Keeps the sampling rate in the job's configuration rather than with the dataset.
- Applies one global rate when negatives were sampled differently per slice.
- Duplicates positives to balance the classes and calls that equivalent.
- Reads calibration off the sampled validation set without correcting it.
- Believes more training epochs will recover the true base rate.