skip to content

questions

27

In a jobs-for-you rail over four million postings, why retrieve a few hundred candidates before scoring?

level: juniorimportance: must knowfreq 74%

answer

  1. millions in, hundreds out
  2. cost per candidate times candidates
  3. index lookup, not a full scan
  4. the shortlist size is a budget
  5. retrieval sets the recall ceiling

basics

~20 s

Scoring every posting is impossible inside the rail's budget: four million postings at ten microseconds each is forty seconds of CPU per request. Retrieval trades exhaustive coverage for a few hundred plausible candidates found in milliseconds.

solid answer

~40 s

The two stages have opposite cost shapes. A feature-rich scorer costs tens of microseconds per candidate, so its bill is linear in how many candidates it sees; four million of them is tens of seconds of CPU for a single rail impression. Retrieval uses an approximate nearest-neighbour index over posting vectors that were computed in batch, so a query touches a small fraction of the corpus and its latency barely moves as the catalogue grows. So the design fixes a shortlist size as a budget — a few hundred — and asks the retrieval stage to fill it. The price of the split is that the shortlist caps everything: within one request the scorer can only reorder what retrieval handed it, so a good posting that retrieval missed is simply not in the rail.

go deeper

for a junior

Recall the shape: a cheap stage narrows millions of postings to hundreds, then an expensive stage ranks those. Be able to say why scoring everything is impossible using a per-candidate cost times the catalogue size.

for a middle

Explain the cost curves: retrieval is roughly flat in catalogue size because it searches an index, scoring is linear in candidate count. Name the shortlist size as a deliberate budget rather than an accident.

for a senior

Show that the shortlist caps achievable recall for that request, and that this is why the retrieval stage is measured on recall rather than on its own ordering. Tie the budget to the latency and spend it buys.

for a principal

Frame the funnel as where the organisation chooses to spend compute: widening retrieval buys recall the scorer can use, deepening the scorer buys precision on what arrived. The split is a cost allocation, and it should be revisited when either curve moves.

## The arithmetic that forces two stages A "jobs for you" rail has a page budget measured in tens of milliseconds and a catalogue of roughly four million open postings. A scoring function worth having — one that fetches per-seeker and per-posting features and runs a model over the pair — costs on the order of ten microseconds per posting. Four million postings times ten microseconds is **forty seconds of CPU for one request**. Spread across forty cores that is still a second of wall clock, and the rail serves thousands of requests a second. No fleet size fixes this, because the cost is incurred per request, not once per day. The funnel splits the work where the two cost curves cross: - **Retrieval cost is roughly independent of catalogue size.** An approximate nearest-neighbour index over precomputed posting vectors inspects a small fraction of the corpus per query, so doubling the catalogue costs memory and very little latency. - **Scoring cost is linear in candidates.** Each candidate means a feature fetch and a forward pass, so the shortlist size *is* the serving bill. - The architecture therefore treats the shortlist size as a **fixed budget** — a few hundred — and makes retrieval's job to fill that budget with the most plausible postings it can find. ## What the retrieval stage is allowed to know Retrieval has to rank candidates it has not enumerated, which constrains what it can compute: - Its scoring function must be **precomputable and searchable**: a seeker vector produced on the request, posting vectors produced in batch, one similarity metric the index can search over. A signal the index cannot search over cannot influence retrieval. - Beside the learned source sit **cheap lookup sources** that need no model at all: postings created in the last few hours inside the seeker's commute radius, postings matching a saved alert. They are ordinary queries against ordinary stores. - Hard constraints that would waste the budget — an expired posting, one the seeker already applied to — are applied as filters around the search rather than as a learned signal. ## The two stages side by side | | retrieval stage | scoring stage | |---|---|---| | input size | the whole catalogue, millions | the shortlist, hundreds | | cost per candidate | a fraction of a microsecond, amortised over the index | tens of microseconds plus a feature fetch | | what it may use | precomputed vectors and cheap lookups | features fetched per candidate | | output | a set of eligible candidates | an ordered list | | its failure mode | a good posting never enters the set | a good posting is placed too low | ## Recall is capped at stage one The funnel is strictly narrowing, and that is the consequence candidates most often miss: 1. Retrieval returns a shortlist. 2. The scorer permutes that shortlist; it is a ranking function over a given set, not a search over the catalogue. 3. Whatever edits the final slate operates inside the scored list. Nothing in that chain introduces a posting the first stage did not supply. Two practical consequences follow. First, **the ordering produced by retrieval barely matters** — membership is everything, because the scorer will re-order anyway. Second, **the retrieval stage's own quality metric is recall of the shortlist**, not any notion of precision: it is measured by how many of the postings that deserved a slot were present, not by how tidy its own ranking was. ## What an interviewer is listening for - The **cost asymmetry stated with a number**, not as "it would be slow". - That retrieval emits a **set under a size budget**, and that the budget is chosen deliberately. - That a miss at stage one is **unrecoverable within that request**, which is why retrieval is measured on recall. - That several cheap sources can feed the same shortlist; the learned index is not the only way in. The answer that fails is "score everything, just add machines": it misreads a per-request cost as a capacity problem, and it scales with catalogue size times traffic, which is exactly the product the funnel exists to avoid.

  • Why must the retrieval stage's scoring function be a vector comparison rather than a feature-rich model?
    Because retrieval has to rank candidates it has not enumerated. The only way to find the best few hundred out of millions without touching all of them is to make the score a geometric relation between a query point and indexed points, which an index can search. A model that needs features fetched per candidate can only run once you already have the candidate list.
  • Does adding a second retrieval source raise the scoring stage's cost?
    Not by itself. Sources are merged, de-duplicated and truncated to the same shortlist size, so a new source changes the composition of the shortlist rather than its length. Cost rises only if you also raise the depth, which is a separate decision with its own latency and spend.
  • If retrieval's own ordering barely matters, why not return a random few hundred eligible postings?
    Because recall is what the stage is judged on, and a random sample of four million has almost no chance of containing the postings that deserve the rail's 25 slots. Retrieval's ordering is throwaway, but its membership decision is the entire product of the stage.

A recruiter with four million CVs does not read them all. They pull the one drawer that matches the role, then read the fifty CVs in it properly. The pull is coarse and fast; the reading is slow and careful, and anything left in the cabinet never gets read.

saying these in an interview costs you the question

  • Just score the whole catalogue in parallel and add machines
  • The scorer can pull in a posting retrieval missed
  • Retrieval should return the final ranking, not a set
  • Retrieval and scoring can use the same per-candidate features
  • Catalogue growth hurts retrieval latency as much as scoring cost
open as a page

In a marketplace serving both a browse feed and a typed product search, what does the typed query change about the candidate set?

level: juniorimportance: must knowfreq 64%

basics

~20 s

A typed query is a retrieval constraint applied before any model runs, so the ranker only sees items that already match the stated intent. A browse feed has no such constraint and must manufacture candidates from the viewer's profile and context.

open as a page

For a listener with no listening history, what decides which rung of a podcast feed's fallback ladder serves the request?

level: middleimportance: must knowfreq 68%

basics

~10 s

A signal test on the request, not the account's age. Each rung declares a precondition; the funnel serves the first rung whose precondition holds, re-checked on every request, and logs which rung answered.

open as a page

A news feed's scored candidates hold forty near-identical stories on one event - how does the list-editing layer avoid showing ten?

level: middleimportance: must knowfreq 70%

basics

~20 s

Cluster the near-duplicates first, then select greedily. Content similarity over title and lead text groups the forty filings into one event cluster; the list-editing pass then picks one representative and, using maximal marginal relevance, scores each further slot on relevance minus similarity to what is already placed.

open as a page

On a marketplace search surface, why is applying an in-stock facet inside retrieval different from filtering the ranked list afterwards?

level: middleimportance: must knowfreq 57%

basics

~20 s

Inside retrieval, the shortlist comes back full of eligible items. Filtering afterwards spends the shortlist on ineligible ones, so a selective facet leaves too few results to fill the page and wastes the scoring stage's budget.

open as a page

In a short-video feed, which features can the second-stage scorer use that the retrieval stage structurally cannot?

level: middleimportance: must knowfreq 68%

basics

~20 s

Features that mix user, clip and request context in one value - this viewer's past finishes of this creator, the slot, the time of day, the connection. Retrieval must fix a clip's representation before the request arrives, so it cannot hold any of them.

open as a page

How do you measure whether a jobs-for-you retrieval stage is dropping postings the scorer would have ranked highly?

level: seniorimportance: must knowfreq 60%

basics

~20 s

Measure two different recalls: the index's recall against exact vector search, and the funnel's recall against the scorer's own top-25 over a much deeper candidate pool replayed offline. Clicks cannot measure it — a posting never retrieved was never shown.

open as a page

How does a podcast feed guarantee a freshly published episode gets its first impressions before any engagement data exists?

level: seniorimportance: must knowfreq 60%

basics

~20 s

By reserving impressions rather than hoping for them: a fixed share of slate positions goes to items below a confidence bar, metered by a per-item impression quota so the budget spreads across inventory, and each item graduates once its estimate is trustworthy.

open as a page

An outlet retracts a story already live in a news feed - at which stage of the funnel do you enforce the takedown?

level: seniorimportance: must knowfreq 63%

basics

~20 s

At the last stage before the response: a suppression check against the final slate, reading a takedown list from a low-latency key-value store on the serving path. Removing the article upstream is also right, but candidate sources and precomputed slates refresh too slowly to be the guarantee.

open as a page

A feed's ranking stage blends predicted finish, like and share probabilities into one score - why must each head be calibrated?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Because a blend adds magnitudes across heads. Each head can order its own predictions perfectly while its numbers are systematically inflated, and the inflated head then dominates the sum - so the weighted order, any fixed cutoff and any value arithmetic come out wrong.

open as a page

A news feed's ranked top ten carries at most two stories per outlet - which post-scoring rule produced that slate?

level: juniorimportance: should knowfreq 57%

basics

~20 s

A per-outlet slot quota in the list-editing pass that runs after scoring. The ranker orders candidates by relevance score; the quota then walks that order and skips any outlet that already holds two slots on the page.

open as a page

In a short-video feed's two-stage ranker, how does a clip's retrieval score differ in job from its ranking score?

level: juniorimportance: should knowfreq 54%

basics

~20 s

The retrieval score decides membership - whether a clip enters the shortlist at all. The ranking score decides order - where a shortlisted clip sits in the slate. The second stage may reverse retrieval's order entirely.

open as a page

A jobs-for-you retrieval stage draws candidates from three sources — how do you merge and de-duplicate them?

level: middleimportance: should knowfreq 55%

basics

~20 s

Merge on rank, not on raw score: cosine similarity, recency order and an alert match are incomparable numbers. Reciprocal rank fusion or fixed per-source quotas produce one list; collapse duplicates on a canonical posting id before truncating.

open as a page

A newly published podcast has no listens and so no behavioural vector — what makes it retrievable by the funnel?

level: middleimportance: should knowfreq 52%

basics

~20 s

A representation built from what exists at publish time — declared category, description text, host and network links, transcript topics. It is either projected into the retrieval space alongside behavioural items or run as a separate candidate source with its own quota in the merge.

open as a page

A news feed caps each outlet at two items per page, yet one outlet fills four consecutive slots - how did that happen?

level: middleimportance: should knowfreq 44%

basics

~20 s

The cap's window is the page, and the window resets at the page boundary. Slots nine and ten of the first page plus slots one and two of the second are four in a row to a scrolling reader, because the second request counted from zero.

open as a page

Job postings expire hourly but the posting-vector index rebuilds nightly — how does the retrieval stage stay fresh?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Split additions from removals. New postings are embedded on arrival into a small fresh segment searched beside the nightly-built base index; filled or expired postings are dropped on the read path from a tombstone set, because the structure only reclaims them at rebuild.

open as a page

How would you shard an approximate nearest-neighbour index of job postings across locales and private employer boards?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Two partitions for two different reasons. Tenant is a correctness boundary: a private board gets its own index selected before the search, never a filter applied afterwards. Locale is a performance partition inside the public index, routed by the seeker's commute area.

open as a page

A new listener picks three topics during onboarding — how long should that stated signal steer the podcast feed?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Only until observed behaviour replaces it. Treat the picks as a low-confidence prior that seeds retrieval while there is nothing else, then decay their weight against completed listens — not against elapsed days — and keep them attributable rather than written in as fake events.

open as a page

A marketplace blends keyword matches and embedding neighbours into one candidate list — why not normalise the two scores and add them?

level: seniorimportance: should knowfreq 46%

basics

~20 s

The two scores share no scale. A keyword score has no fixed ceiling and moves with query length and term rarity, while a similarity score is bounded but query-dependent, so per-query normalisation makes a weak list look exactly like a strong one. Rank fusion merges on position instead.

open as a page

Why does a marketplace's typed product search need a relevance cutoff that can return nothing, when its browse feed does not?

level: seniorimportance: should knowfreq 41%

basics

~20 s

A typed query asserts an intent the catalogue may not satisfy, and retrieval will hand back its nearest candidates regardless, so without a cutoff the surface silently answers a different question. A browse feed makes no such assertion to contradict.

open as a page

A feed ranker trains on the clips its own funnel showed - what does that training set never contain, and what does the gap cost?

level: seniorimportance: should knowfreq 40%

basics

~20 s

It never contains outcomes for clips retrieval missed or the ranker buried, because nobody saw them. Those are unobserved, not negative. The model ends up confident where the funnel already agreed and evidence-free exactly where it must extrapolate.

open as a page

Your feed's ranking stage has 30 ms and a heavy scorer costing 0.15 ms per clip - how do you decide the shortlist size it serves?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Subtract the stage's fixed cost from its budget and divide by the per-candidate cost. With 6 ms of batched feature fetch and setup, 24 ms remain, so about 160 clips fit. Shortlist size is arithmetic, not taste.

open as a page

A news feed's post-scoring rule layer has grown to forty rules from four teams - what governance keeps it honest?

level: principalimportance: should knowfreq 37%

basics

~20 s

Make precedence a property of the system, not of insertion order; give every rule an owner, a reason and an expiry; measure each rule's cost with holdouts and per-rule attribution logging; and ship and retire a rule through the same ramp and flag a model change gets.

open as a page

A two-tower jobs rail ships a new seeker encoder but keeps the existing posting-vector index — what breaks?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Retrieval silently returns near-random postings. Each training run produces its own coordinate space, so a new seeker vector compared against the previous encoder's posting vectors is a meaningless comparison — with no error, unchanged latency and a full result set.

open as a page

Why apply a news feed's recency boost after the ranker instead of feeding article age into the ranking model?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Because the two change on different timescales. A post-scoring multiplier is an operations lever an editor can turn during a breaking event and reverse in minutes; article age as a model input learns a better-shaped response from data but only moves when the model is retrained.

open as a page

A marketplace trains one ranker on clicks pooled from its browse feed and its typed search results. What goes wrong?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

The two surfaces label different things. A feed click signals interest; a search click signals a match to a stated intent, and an abandoned query is a failure on search but an ordinary session on the feed. Pooled, the higher-traffic surface's objective wins.

open as a page

What would justify raising the share of a podcast feed's impressions reserved for items with no engagement history?

level: principalimportance: nice to knowfreq 32%

basics

~10 s

Evidence that the current reserve no longer bootstraps the inventory arriving: a graduation rate below the publishing rate, p95 time-to-first-impression drifting out, falling catalogue coverage, or supply-side churn among creators who never get shown.

open as a page