A feed ranker trains on the clips its own funnel showed - what does that training set never contain, and what does the gap cost?
answer
- the log is the funnel's output
- unobserved is not the same as negative
- dense evidence where it already agrees
- log the shortlist, not just impressions
- position is part of every outcome
basics
~20 sIt never contains outcomes for clips retrieval missed or the ranker buried, because nobody saw them. Those are unobserved, not negative. The model ends up confident where the funnel already agreed and evidence-free exactly where it must extrapolate.
solid answer
~50 sLabels exist only where an impression happened, so the training set is the funnel's own output, not a sample of the catalogue. Two absences follow: clips retrieval never returned, and clips that were shortlisted and scored but placed too low to be seen. Recording the second group as negatives is the common and damaging mistake - the viewer made no decision about them, so the row asserts a rejection that never occurred, and the model learns to reproduce the current funnel's habits rather than the viewer's preferences. Meanwhile at serving time the scorer is asked to score every shortlisted clip, including many it has almost no evidence about. The architectural remedies are logging ones: record the whole shortlist with each clip's stage scores and slot, capture the feature values as they were at scoring time, and keep a small randomised fraction of traffic whose slate is not fully score-ordered so some otherwise-buried clips get honest outcomes.
go deeper
Remember that outcomes only exist for clips that were actually shown. Everything the funnel dropped has no result attached, and treating silence as a rejection puts a claim into the data that was never observed.
Explain the three populations - never retrieved, scored but unshown, shown - and why only the third carries a label, then say what happens to a model trained as if the second were negative.
Name the logging contract you would put in place: the full shortlist with stage scores and slot, feature values captured at scoring time, and exposure probability recorded so rows can be weighted.
Argue the trade explicitly. Unbiased evidence about the funnel's margins has to be bought with a small amount of deliberately unordered traffic, and a system that refuses to spend it can only reproduce itself.
The second-stage scorer of a short-video feed is trained on logs from the feed it ranks. That sounds obviously right and quietly biases everything, which is why interviewers reach for it. ## Three populations, one of them labelled For a single request the catalogue splits into three groups: 1. **Never retrieved.** No stage saw the clip. No score, no row, no outcome. 2. **Retrieved and scored, never shown.** The clip was in the shortlist, the heavy scorer produced a number, and the slate stopped above it. The viewer made no decision about it. 3. **Shown.** The viewer had a chance to finish, like or share, so an outcome exists. Only group three carries a label. Group two is the one that causes trouble, because a row for it exists somewhere in the system and it is tempting to file it as a negative. ## Unobserved is not negative A negative example asserts a fact: the viewer had the chance and declined. A clip that was never rendered supports no such assertion. Writing it down as a negative injects a claim the data does not contain, and the claim has a direction - it says that whatever the current ranker placed low is bad. Train on that and the model is rewarded for agreeing with its predecessor, which is the one thing that produces no improvement at all. The cost shows up as a mismatch between what the model saw and what it is asked to do: - At training time it sees mostly clips the funnel already liked, so its evidence is dense in the region it already handles well. - At serving time it must score every clip on the shortlist, including many whose kind it has almost never observed with a real outcome. - Confidence is therefore highest exactly where it matters least, and the scores that decide marginal clips rest on the least evidence. There is a second distortion inside group three. An outcome on the top slot is partly an outcome about the top slot: the same clip is finished at very different rates at the top of a feed and ten positions down. Unless position is recorded and handled, the model learns that whatever was shown first is good. ## The logging that makes the gap fixable Every remedy here is an architecture decision made before any model is trained, and it is what the interviewer is checking for. | record this | why it matters | |---|---| | the whole shortlist, not just impressions | separates never-retrieved from scored-but-unshown | | each clip's stage scores | shows what the funnel believed about the clips it dropped | | the slot the clip would have occupied | lets the model separate the clip from its position | | feature values as read at scoring time | the row matches what the scorer saw, not a later recomputation | | the exposure probability of the slate | supports weighting rows by how likely they were to be seen | The fourth row deserves emphasis. Recomputing a feature from the offline store after the fact produces a value the scorer never saw - the counters have moved on - so the training row describes a decision nobody made. Logging the value at the moment of scoring is the only way the row is true. ## The one deliberate cost The structural remedy is to spend a little traffic on honesty: a small randomised fraction of requests whose slate is not strictly score-ordered produces outcomes for clips the ranker would have buried. Those rows are the only unbiased evidence the funnel generates about its own margins, and the cost is a slightly worse slate for that fraction of traffic. How much to spend is a business decision; that some is spent is what makes the next model better than the current one rather than a copy of it. ## The compact answer Say it as a chain: the log is the funnel's output, not a sample; unshown is unobserved, not negative; the model therefore knows most about what the funnel already agreed on and least about what it must decide; and the fixes are logging decisions - full shortlist, stage scores, slot, serving-time feature values - plus a small deliberate randomisation. A candidate who describes the model instead of the logging has answered a different question.
- Why is recomputing a feature value later, rather than logging it at scoring time, a problem?Because the counters and aggregates have moved since the decision. A recomputed row pairs the outcome with a value the scorer never saw, so the model is fitted on a decision nobody made, and the error is systematic rather than noisy - it usually reflects the very activity the outcome caused. Capturing the served value at scoring time is the only way the row is true.
- What does logging the slot alongside the outcome buy you?It lets the model separate the clip's appeal from the advantage of its position, since the same clip is finished at very different rates high and low in a feed. Position is included as an input during training and pinned to a fixed value at serving, so the served score compares clips rather than the positions they happened to get.
A shop that only records sales learns a great deal about what it stocked and nothing about what customers wanted and never found on the shelf.
saying these in an interview costs you the question
- Files every unshown shortlisted clip as a negative example
- Calls the impression log a random sample of the catalogue
- Recomputes feature values after the fact for training rows
- Ignores position when reading an outcome on the top slot
- Assumes more impression data alone removes the bias