A feed's ranking stage blends predicted finish, like and share probabilities into one score - why must each head be calibrated?
answer
- order survives any monotone stretch
- a blend adds magnitudes, not ranks
- an inflated head buys extra weight
- cutoffs and value arithmetic break first
- check predicted against observed per slice
basics
~20 sBecause a blend adds magnitudes across heads. Each head can order its own predictions perfectly while its numbers are systematically inflated, and the inflated head then dominates the sum - so the weighted order, any fixed cutoff and any value arithmetic come out wrong.
solid answer
~60 sOrdering one list by one head only needs the head to be **monotone**: any stretched or squashed version of the scores produces the same order. The moment several heads are combined - `value_finish x p_finish + value_like x p_like + value_share x p_share` - the arithmetic compares heads against each other, and a head whose probabilities read four times too high contributes four times the weight you assigned it. Clips strong on the inflated objective out-rank clips with genuinely higher expected value, and no re-tuning of the weights fixes it cleanly, because the weights are now absorbing a scale error instead of expressing business value. The same is true of anything else that reads the number rather than the order: a fixed cutoff below which nothing is shown, a quota expressed as a probability, or any downstream value computation. The fix is a monotone calibration map fitted on a recent held-out window - Platt scaling or isotonic regression - plus a standing check that mean predicted probability matches the observed rate per slice.
code
pseudocode · 14 linesvalue = { finish: 1.0, like: 3.0, share: 10.0 }
function blended(p):
return value.finish * p.finish
+ value.like * p.like
+ value.share * p.share
// calibrated heads: B is genuinely the better clip
A = { finish: 0.20, like: 0.12, share: 0.005 } // blended -> 0.61
B = { finish: 0.60, like: 0.04, share: 0.010 } // blended -> 0.82 B first
// like head reads 4x high; its own order is unchanged (0.48 > 0.16)
A = { finish: 0.20, like: 0.48, share: 0.005 } // blended -> 1.69 A first
B = { finish: 0.60, like: 0.16, share: 0.010 } // blended -> 1.18go deeper
Recall the distinction: a score that orders clips correctly is not necessarily a score whose value means anything. Only the second kind can be added, compared to a cutoff or multiplied by a business value.
Explain why a monotone distortion is invisible to a sorted list and to order-only metrics, and why a blend is different because it puts several heads on one axis and sums them.
Show the operating loop: refit the calibration map with every promoted model version, watch predicted against observed per slice, and treat a flat ranking metric as no evidence that the scale is intact.
Decide what the scores are a contract for. Once anything downstream consumes them as values, calibration becomes a platform guarantee with an owner and a check, not a step inside one team's model.
A modern feed rarely optimises one thing. The second stage predicts several outcomes for each shortlisted clip - the viewer finishes it, likes it, shares it - and combines them into one number that the slate is sorted by. That combination is where calibration stops being a nicety and becomes a correctness requirement. ## Order needs monotonicity; arithmetic needs the scale A score is **calibrated** when its value matches the observed frequency: among the clips predicted at 0.20, about 20 percent are finished. A score is merely **ordered** when higher means more likely without the value meaning anything. Sorting one list by one head needs only order. Apply any increasing function to every prediction of a head and the sorted list is identical - which is exactly why a head can look excellent on an order-only metric such as the area under the ROC curve while its numbers are far from the truth. A blend is arithmetic, not sorting. Adding `3 x p_like` to `1 x p_finish` puts the two heads on one axis and asserts that a like is worth three finishes. If the like head reads four times too high, the assertion the system actually enforces is that a like is worth twelve finishes, and the weights no longer mean what the document says they mean. ## A worked inversion Take value weights of 1 for a finish, 3 for a like and 10 for a share, and two clips: | | finish | like | share | blended value | |---|---|---|---|---| | clip A, calibrated | 0.20 | 0.12 | 0.005 | 0.61 | | clip B, calibrated | 0.60 | 0.04 | 0.010 | 0.82 | | clip A, like head inflated 4x | 0.20 | 0.48 | 0.005 | 1.69 | | clip B, like head inflated 4x | 0.60 | 0.16 | 0.010 | 1.18 | Calibrated, B is the better clip and is shown first. With the like head reading four times high, A wins - even though **the inflated head still ranks its own two predictions in the same order** (0.48 above 0.16, exactly as 0.12 was above 0.04). Nothing about the head's internal ordering changed. The blend inverted anyway, because the blend reads magnitudes. ## What else reads the number rather than the order - **A fixed cutoff.** A rule such as do not show anything below a 2 percent finish probability means nothing if the numbers are off by a factor; it silently becomes a much stricter or much looser rule. - **Cross-request comparisons.** Deciding whether today's best clip is good enough to be worth a slot at all compares scores across different shortlists, which order within one shortlist cannot support. - **Anything expressed in real units.** The moment a score is multiplied by a value in currency or minutes, a scale error becomes a value error. ## Why heads drift out of calibration 1. **Resampled training data.** If negatives are downsampled to make training tractable, the raw head predicts the resampled rate, not the live one, and every prediction is inflated until it is corrected. 2. **A shifted serving mix.** Calibration is fitted on one distribution of traffic; when the mix of viewers, surfaces or clip types moves, the fitted map is no longer right even though the model has not changed. 3. **A retrain.** A new model version can land with a better order and worse calibration, so the calibration step is refitted with every promotion rather than inherited. ## The fix, and how it is checked Fit a **monotone calibration map** on a recent held-out window: Platt scaling, a single logistic fit from raw score to probability, is stable on little data; isotonic regression, a monotone step function, is more flexible but needs far more data per bin and can overfit thin ones. Because both maps are monotone, neither changes any single head's own ordering - they change what the numbers mean, which is exactly the property the blend needs. Then keep a standing check. Compare mean predicted probability against the observed rate **per slice** - surface, viewer tenure, clip age - not just globally, because a global figure averages an over-predicting slice against an under-predicting one and reports health. Expected calibration error summarises the same comparison across score buckets. And watch for the trap this whole subject exists to name: an order-only metric will not move when calibration rots, so a dashboard built from ranking metrics alone reports a healthy model while every blended decision downstream is wrong.
- Could you just re-tune the blend weights instead of calibrating the heads?You can chase the symptom, but the weights then carry two jobs at once: the business value of an outcome and a correction for one head's scale. The next retrain changes the scale again and the weights are wrong in a new way, and nobody can read the configuration and say what the system values. Calibrate the heads so the weights mean only what they claim to mean.
- Which monitoring signal catches a calibration problem that ranking metrics miss?Mean predicted probability against the observed rate, computed per slice rather than globally, with expected calibration error as the summary. Order-only metrics are invariant to any monotone distortion, so they stay flat while the scale rots. Slicing matters because a global average can hide an over-predicting segment cancelling an under-predicting one.
- When would isotonic regression be the wrong calibration choice?When the held-out window is small or thin in the region that matters. Isotonic regression fits a monotone step function and needs enough observations per bin, so on sparse data it chases noise and produces flat plateaus that destroy resolution. A single logistic fit is the safer default there, and the refit cadence matters more than the choice between them.
saying these in an interview costs you the question
- Says calibration cannot matter because only the order is shown
- Believes an uncalibrated head is harmless once several heads are blended
- Uses an order-only metric to conclude the scores are trustworthy
- Re-tunes blend weights to absorb one head's scale error
- Reports one global calibration number across all slices
- Inherits the old calibration map after promoting a new model version