In a short-video feed's two-stage ranker, how does a clip's retrieval score differ in job from its ranking score?
answer
- two numbers, two different jobs
- one admits, the other orders
- membership versus position
- scales not comparable across sources
- admitted first can be shown last
basics
~20 sThe retrieval score decides membership - whether a clip enters the shortlist at all. The ranking score decides order - where a shortlisted clip sits in the slate. The second stage may reverse retrieval's order entirely.
solid answer
~50 sThey answer different questions. The **retrieval score** is an admission ticket: computed cheaply over a very large catalogue, usually as a similarity between a user representation and a precomputed clip representation, it only has to get the right few hundred clips into the shortlist. The **ranking score** is a prediction about this user, this clip and this moment, computed one pair at a time with interaction and context features, and it is the number the slate is sorted by. Because the stages are optimised for different things - getting good clips into the shortlist versus ordering the shortlist - their orders routinely disagree, and a clip admitted first can be shown last. Similarities from two different retrieval sources are not even on a common scale, so the final order is never taken from them; at most the retrieval score enters the second stage as one more input feature.
go deeper
Recall that a feed request produces two scores with two jobs: one decides which clips enter the shortlist, the other decides the order of what is shown. They are not the same number computed twice.
Explain why the first stage's score cannot answer the ordering question: it is computed over the whole index from a representation fixed before the request, so it never sees the specific user-clip pair.
Show the operational consequence: the second stage looks healthy when the first stage silently stops returning a class of clips, so log what each stage received as well as what it emitted.
Frame the split as where the quality ceiling is set. Shortlist size and source coverage bound how good the final order can ever be, and that bound is bought with latency and compute, not with a better ranker.
A short-video home feed cannot run an expensive model over a catalogue of hundreds of millions of clips inside a request. It runs a funnel instead: a cheap stage cuts the catalogue down to a few hundred clips, and an expensive stage puts those few hundred in order. Both stages attach a number to a clip, and interviewers ask about the two numbers because candidates routinely treat them as one quantity measured twice. ## The retrieval score is an admission ticket The first stage answers a set question: which few hundred clips are worth spending real compute on? Its score exists only to cut the catalogue down, and it is shaped entirely by that job. - It is evaluated over an enormous number of items, so its per-item cost has to be minuscule - typically a similarity read out of an approximate nearest-neighbour index, or a position inside a simpler source such as a recency or popularity list. - It is usually **not comparable across retrieval sources**. A cosine similarity from one index and a rank inside a freshness list are different quantities, and putting them in one sorted list compares apples with millimetres. - Its information about the user arrives through a precomputed representation, never through a feature that names the specific pair. Once the shortlist exists, the admission ticket has been used. It is either discarded or carried forward as one input feature among many. ## The ranking score is a per-pair prediction The second stage answers a different question: given this user, this clip and this moment, what will happen if we show it? It scores each shortlisted clip individually, with features that name the pair - how many clips from this creator the viewer finished this week, how the last few clips of this length held them, time of day, connection quality, the slot the clip would occupy. - It runs over hundreds of items rather than millions, so it can afford a far heavier model and far richer inputs. - Its output is a prediction about an outcome (a finish, a like, a share), not a geometric similarity. - It is the number the slate is sorted by and, when the scores are calibrated, the number any downstream arithmetic consumes. ## Why the two orders are allowed to disagree | | retrieval score | ranking score | |---|---|---| | question answered | is this clip in the shortlist? | where does it go in the slate? | | items scored per request | the whole index | a few hundred survivors | | sees the user-clip pair | only through a precomputed representation | directly, with interaction features | | unit | a similarity or a source-local rank | a predicted outcome for this pair | | after the stage | discarded, or kept as one feature | sorts the slate, feeds any blend | A clip can sit first on the retrieval list and last in the slate. That is the cascade working as designed. The first stage was tuned to get good clips into the shortlist at all; it cannot see the signals that decide order. The reverse also holds: a clip the first stage never returned gets no ranking score, ever. ## Three rules a design round expects you to state 1. **Do not sort the final slate by the retrieval score.** It was never fitted to the ordering question, and scores from different sources share no scale. 2. **Do not read a low ranking score as a retrieval failure.** The shortlist is supposed to contain clips that lose; a shortlist where everything scores highly is usually too small. 3. **Do not average the two numbers.** They are different quantities in different units. The retrieval score can be a feature the second stage learns a weight for; it cannot be a term you add by hand. ## What the split costs, and what it hides The two scores also explain the shape of the funnel's cost. First-stage cost scales with index size and is largely paid whether or not anyone is watching; second-stage cost scales with shortlist size and is paid per surviving candidate on every request. Shortlist size is therefore the dial that trades quality against the ranking stage's latency budget. The split also hides one failure mode worth naming out loud. If the first stage stops returning a whole class of clips - a creator, a language, a freshness band - the second stage shows no symptom at all. Its scores stay well behaved, its offline metrics on the shortlist look normal, and the loss is invisible unless the funnel logs what each stage received as well as what it emitted. The order the user sees is produced by the second stage, but the ceiling on how good that order can be was set by the first.
- Is it ever legitimate for the retrieval score to influence the final order?Yes, but only as an input the second stage learns a weight for. Passing it in as a feature lets the ranker discover that a strong similarity carries signal, and lets it discount the value for sources where it does not. What is illegitimate is adding it to the ranking score by hand, or using it to break ties, because the two numbers are in different units and the similarity is not comparable across sources.
- The shortlist is 300 clips and almost all of them score highly. What does that suggest?Usually that the shortlist is too small or too narrow. A healthy shortlist is meant to contain plenty of clips the ranker rejects; that is how the second stage earns its cost. If nearly everything survives, the first stage is probably over-filtering - the good clips outside the shortlist never get a ranking score, and no second-stage metric will show their absence.
The door list decides who gets into the room; the seating chart decides where they sit. Being first on the door list buys no seat at the front.
saying these in an interview costs you the question
- Sorts the final slate by the retrieval similarity score
- Treats similarities from two retrieval sources as comparable numbers
- Says a low ranking score means retrieval returned the wrong clip
- Expects the second stage to only lightly adjust retrieval's order
- Assumes the shortlist should contain only clips worth showing