skip to content

questions

5

In a short-video feed, which features can the second-stage scorer use that the retrieval stage structurally cannot?

level: middleimportance: must knowfreq 68%

answer

  1. the first stage cannot hold a pair
  2. clip side is fixed before the request
  3. interaction, context and slate features
  4. no stored value per viewer-clip pair
  5. budget, not model shape, is the limit

basics

~20 s

Features that mix user, clip and request context in one value - this viewer's past finishes of this creator, the slot, the time of day, the connection. Retrieval must fix a clip's representation before the request arrives, so it cannot hold any of them.

solid answer

~50 s

The retrieval stage has a structural constraint: to search a very large catalogue quickly it must precompute one representation per clip, offline, without knowing who is asking. Any feature that mixes a viewer with a clip therefore cannot exist there - it would need a separate index per viewer. The second stage has no such constraint. It runs once per (viewer, clip, context) pair over a few hundred survivors, so it can use **interaction features** (finishes of this creator this week, skips of this topic today), **request context** (hour, connection quality, whether the session just started) and **slate context** (the slot the clip would occupy). That is the whole economic argument for the cascade: you pay for the expensive inputs only on candidates that survived the cheap filter. The binding limit on the second stage is not model shape, it is whether every one of those values can be fetched for every survivor inside the stage's latency budget.

go deeper

for a junior

Remember that the expensive stage sees the viewer and the clip together, while the cheap stage had to describe each clip before knowing who would ask. That difference, not model size, is why there are two stages.

for a middle

Be able to explain the mechanism: a viewer-by-clip value cannot be stored on an indexed clip without one entry per pair, so it can only be computed after the shortlist shrinks the pair count to a few hundred.

for a senior

Demonstrate the operating judgment: each added feature is a latency decision, so name the batched read, the per-candidate cost, the scoring slots it removes, and the default used when the read times out.

for a principal

Treat the feature surface as a budget to allocate. Extra per-candidate cost buys interaction signal but shrinks the shortlist, and past a point the funnel gains more from a wider shortlist than from a richer scorer.

A design round almost always reaches the question of why a recommendation feed bothers with two models rather than one good one. The honest answer is not that the second model is bigger. It is that the second model can see something the first one structurally cannot: the pair. ## The constraint that shapes the first stage To select a few hundred clips out of a very large catalogue in a couple of milliseconds, the first stage must do almost no work per clip at request time. Every architecture that achieves this - two-tower retrieval over an approximate nearest-neighbour index, a popularity list, a recency list, a co-visitation list - gets there the same way: **the clip side is computed ahead of the request and stored**. The request contributes a query vector or a key, and the index does the rest. That storage decision has a consequence that candidates often miss. A value that depends on both the viewer and the clip cannot be stored on the clip. There is no place to put it: you would need one stored value per (viewer, clip) pair, which is the cross product you built the index to avoid. So the first stage learns about the viewer only through a representation that is combined with the clip's at the last possible moment, by a similarity function with no learned interaction inside it. ## What the second stage gains The second stage runs over a few hundred survivors, one pass per pair, after the request has arrived. Every input the first stage could not hold becomes available: - **Interaction features** - how many clips from this creator the viewer finished this week, how many they skipped in the first two seconds, when they last saw this topic, whether this clip was already shown and ignored yesterday. - **Request context** - hour of day, day of week, whether the session just opened or is twenty clips deep, connection quality, whether the sound is on. - **Slate context** - the slot the clip is being considered for, and what has already been placed above it. - **Very fresh signals** - counters from the last few minutes that no offline index build could have baked in. - **Learned crosses** - representations of viewer and clip combined by the model itself rather than by a fixed similarity. | | retrieval stage | second-stage scorer | |---|---|---| | when the clip side is computed | before the request, offline | during the request | | items touched per request | the whole index | a few hundred survivors | | can hold a viewer-by-clip value | no - it would need one per pair | yes - it is computed per pair | | request context available | barely, through the query side | fully | | cost model | amortised over index builds | paid per candidate, per request | ## The limit that replaces it Removing the representation constraint does not make the second stage free. It swaps one constraint for another: **every feature must be fetchable for every survivor inside the stage's budget**. In practice that means three questions before any feature is added. 1. Where does the value come from at request time, and is that read batched across the whole shortlist rather than issued per clip? 2. What does it add to the per-candidate cost, and therefore how many scoring slots does it take away? A feature that doubles per-candidate cost halves the shortlist. 3. What happens when the read fails or times out? A feature with no defined default turns a slow dependency into a broken slate. That third question is where real systems break. A feature whose value is silently absent at request time is not a neutral loss: the model was fitted on rows where the value was present, so an absent value pushes the score somewhere arbitrary rather than somewhere safe. The architectural answer is to give every request-time feature an explicit default and to score with the defaults rather than fail the request - a degraded slate beats an empty one. ## Saying this well in an interview The compact form is: **the first stage cannot hold a feature that names both sides, the second stage exists precisely to hold those features, and the price is per-candidate cost**. Everything else in the cascade falls out of that sentence - why the shortlist is a few hundred and not a few million, why the expensive inputs are fetched after the filter and not before, and why adding a feature to the scorer is a latency decision as much as a quality one.

  • If pair features are so valuable, why not push a few of them into the retrieval stage?
    Because storing one value per viewer-clip pair is the cross product the index exists to avoid. What you can do instead is add cheap viewer-independent signals to the clip side, or run several retrieval sources whose selection rules differ, and let the second stage do the pairing. Anything genuinely per-pair belongs after the shortlist, where the number of pairs is a few hundred rather than the catalogue size.
  • A new interaction feature improves the scorer offline. What do you check before serving it?
    That the value can be read for every shortlisted clip in one batched request inside the stage's budget, what it adds to per-candidate cost and therefore how many scoring slots it costs, and what the scorer does when the read times out. Give it an explicit default and score with defaults rather than failing the request, so a slow dependency degrades the slate instead of emptying it.
  • Why does the slot a clip would occupy count as a feature the scorer can use?
    Because it is known at scoring time and it materially changes the outcome - the same clip is finished at different rates at the top of a feed and ten positions down. Including it lets the model separate the clip's appeal from the advantage of the position, provided the slot is logged for training and pinned to a fixed value at serving so the score compares clips rather than positions.

saying these in an interview costs you the question

  • Says the retrieval stage could just use the same features
  • Describes the second stage as only a bigger first stage
  • Forgets that clip representations must be indexable without the viewer
  • Adds a feature without asking whether it exists at request time
  • Assumes every offline feature can be fetched inside the stage budget
  • Leaves a missing feature value undefined instead of giving it a default
open as a page

A feed's ranking stage blends predicted finish, like and share probabilities into one score - why must each head be calibrated?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Because a blend adds magnitudes across heads. Each head can order its own predictions perfectly while its numbers are systematically inflated, and the inflated head then dominates the sum - so the weighted order, any fixed cutoff and any value arithmetic come out wrong.

open as a page

In a short-video feed's two-stage ranker, how does a clip's retrieval score differ in job from its ranking score?

level: juniorimportance: should knowfreq 54%

basics

~20 s

The retrieval score decides membership - whether a clip enters the shortlist at all. The ranking score decides order - where a shortlisted clip sits in the slate. The second stage may reverse retrieval's order entirely.

open as a page

A feed ranker trains on the clips its own funnel showed - what does that training set never contain, and what does the gap cost?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It never contains outcomes for clips retrieval missed or the ranker buried, because nobody saw them. Those are unobserved, not negative. The model ends up confident where the funnel already agreed and evidence-free exactly where it must extrapolate.

open as a page

Your feed's ranking stage has 30 ms and a heavy scorer costing 0.15 ms per clip - how do you decide the shortlist size it serves?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Subtract the stage's fixed cost from its budget and divide by the per-candidate cost. With 6 ms of batched feature fetch and setup, 24 ms remain, so about 160 clips fit. Shortlist size is arithmetic, not taste.

open as a page