skip to content

A feed's ranking stage blends predicted finish, like and share probabilities into one score - why must each head be calibrated?

level: seniorimportance: must knowfreq 58%

answer

  1. order survives any monotone stretch
  2. a blend adds magnitudes, not ranks
  3. an inflated head buys extra weight
  4. cutoffs and value arithmetic break first
  5. check predicted against observed per slice

basics

~20 s

Because a blend adds magnitudes across heads. Each head can order its own predictions perfectly while its numbers are systematically inflated, and the inflated head then dominates the sum - so the weighted order, any fixed cutoff and any value arithmetic come out wrong.

solid answer

~60 s

Ordering one list by one head only needs the head to be **monotone**: any stretched or squashed version of the scores produces the same order. The moment several heads are combined - `value_finish x p_finish + value_like x p_like + value_share x p_share` - the arithmetic compares heads against each other, and a head whose probabilities read four times too high contributes four times the weight you assigned it. Clips strong on the inflated objective out-rank clips with genuinely higher expected value, and no re-tuning of the weights fixes it cleanly, because the weights are now absorbing a scale error instead of expressing business value. The same is true of anything else that reads the number rather than the order: a fixed cutoff below which nothing is shown, a quota expressed as a probability, or any downstream value computation. The fix is a monotone calibration map fitted on a recent held-out window - Platt scaling or isotonic regression - plus a standing check that mean predicted probability matches the observed rate per slice.

code

pseudocode · 14 lines
pseudocode
value = { finish: 1.0, like: 3.0, share: 10.0 }

function blended(p):
    return value.finish * p.finish
         + value.like   * p.like
         + value.share  * p.share

// calibrated heads: B is genuinely the better clip
A = { finish: 0.20, like: 0.12, share: 0.005 }   // blended -> 0.61
B = { finish: 0.60, like: 0.04, share: 0.010 }   // blended -> 0.82  B first

// like head reads 4x high; its own order is unchanged (0.48 > 0.16)
A = { finish: 0.20, like: 0.48, share: 0.005 }   // blended -> 1.69  A first
B = { finish: 0.60, like: 0.16, share: 0.010 }   // blended -> 1.18

go deeper

for a junior

Recall the distinction: a score that orders clips correctly is not necessarily a score whose value means anything. Only the second kind can be added, compared to a cutoff or multiplied by a business value.

for a middle

Explain why a monotone distortion is invisible to a sorted list and to order-only metrics, and why a blend is different because it puts several heads on one axis and sums them.

for a senior

Show the operating loop: refit the calibration map with every promoted model version, watch predicted against observed per slice, and treat a flat ranking metric as no evidence that the scale is intact.

for a principal

Decide what the scores are a contract for. Once anything downstream consumes them as values, calibration becomes a platform guarantee with an owner and a check, not a step inside one team's model.

A modern feed rarely optimises one thing. The second stage predicts several outcomes for each shortlisted clip - the viewer finishes it, likes it, shares it - and combines them into one number that the slate is sorted by. That combination is where calibration stops being a nicety and becomes a correctness requirement. ## Order needs monotonicity; arithmetic needs the scale A score is **calibrated** when its value matches the observed frequency: among the clips predicted at 0.20, about 20 percent are finished. A score is merely **ordered** when higher means more likely without the value meaning anything. Sorting one list by one head needs only order. Apply any increasing function to every prediction of a head and the sorted list is identical - which is exactly why a head can look excellent on an order-only metric such as the area under the ROC curve while its numbers are far from the truth. A blend is arithmetic, not sorting. Adding `3 x p_like` to `1 x p_finish` puts the two heads on one axis and asserts that a like is worth three finishes. If the like head reads four times too high, the assertion the system actually enforces is that a like is worth twelve finishes, and the weights no longer mean what the document says they mean. ## A worked inversion Take value weights of 1 for a finish, 3 for a like and 10 for a share, and two clips: | | finish | like | share | blended value | |---|---|---|---|---| | clip A, calibrated | 0.20 | 0.12 | 0.005 | 0.61 | | clip B, calibrated | 0.60 | 0.04 | 0.010 | 0.82 | | clip A, like head inflated 4x | 0.20 | 0.48 | 0.005 | 1.69 | | clip B, like head inflated 4x | 0.60 | 0.16 | 0.010 | 1.18 | Calibrated, B is the better clip and is shown first. With the like head reading four times high, A wins - even though **the inflated head still ranks its own two predictions in the same order** (0.48 above 0.16, exactly as 0.12 was above 0.04). Nothing about the head's internal ordering changed. The blend inverted anyway, because the blend reads magnitudes. ## What else reads the number rather than the order - **A fixed cutoff.** A rule such as do not show anything below a 2 percent finish probability means nothing if the numbers are off by a factor; it silently becomes a much stricter or much looser rule. - **Cross-request comparisons.** Deciding whether today's best clip is good enough to be worth a slot at all compares scores across different shortlists, which order within one shortlist cannot support. - **Anything expressed in real units.** The moment a score is multiplied by a value in currency or minutes, a scale error becomes a value error. ## Why heads drift out of calibration 1. **Resampled training data.** If negatives are downsampled to make training tractable, the raw head predicts the resampled rate, not the live one, and every prediction is inflated until it is corrected. 2. **A shifted serving mix.** Calibration is fitted on one distribution of traffic; when the mix of viewers, surfaces or clip types moves, the fitted map is no longer right even though the model has not changed. 3. **A retrain.** A new model version can land with a better order and worse calibration, so the calibration step is refitted with every promotion rather than inherited. ## The fix, and how it is checked Fit a **monotone calibration map** on a recent held-out window: Platt scaling, a single logistic fit from raw score to probability, is stable on little data; isotonic regression, a monotone step function, is more flexible but needs far more data per bin and can overfit thin ones. Because both maps are monotone, neither changes any single head's own ordering - they change what the numbers mean, which is exactly the property the blend needs. Then keep a standing check. Compare mean predicted probability against the observed rate **per slice** - surface, viewer tenure, clip age - not just globally, because a global figure averages an over-predicting slice against an under-predicting one and reports health. Expected calibration error summarises the same comparison across score buckets. And watch for the trap this whole subject exists to name: an order-only metric will not move when calibration rots, so a dashboard built from ranking metrics alone reports a healthy model while every blended decision downstream is wrong.

  • Could you just re-tune the blend weights instead of calibrating the heads?
    You can chase the symptom, but the weights then carry two jobs at once: the business value of an outcome and a correction for one head's scale. The next retrain changes the scale again and the weights are wrong in a new way, and nobody can read the configuration and say what the system values. Calibrate the heads so the weights mean only what they claim to mean.
  • Which monitoring signal catches a calibration problem that ranking metrics miss?
    Mean predicted probability against the observed rate, computed per slice rather than globally, with expected calibration error as the summary. Order-only metrics are invariant to any monotone distortion, so they stay flat while the scale rots. Slicing matters because a global average can hide an over-predicting segment cancelling an under-predicting one.
  • When would isotonic regression be the wrong calibration choice?
    When the held-out window is small or thin in the region that matters. Isotonic regression fits a monotone step function and needs enough observations per bin, so on sparse data it chases noise and produces flat plateaus that destroy resolution. A single logistic fit is the safer default there, and the refit cadence matters more than the choice between them.

saying these in an interview costs you the question

  • Says calibration cannot matter because only the order is shown
  • Believes an uncalibrated head is harmless once several heads are blended
  • Uses an order-only metric to conclude the scores are trustworthy
  • Re-tunes blend weights to absorb one head's scale error
  • Reports one global calibration number across all slices
  • Inherits the old calibration map after promoting a new model version