skip to content

In a jobs-for-you rail over four million postings, why retrieve a few hundred candidates before scoring?

level: juniorimportance: must knowfreq 74%

answer

  1. millions in, hundreds out
  2. cost per candidate times candidates
  3. index lookup, not a full scan
  4. the shortlist size is a budget
  5. retrieval sets the recall ceiling

basics

~20 s

Scoring every posting is impossible inside the rail's budget: four million postings at ten microseconds each is forty seconds of CPU per request. Retrieval trades exhaustive coverage for a few hundred plausible candidates found in milliseconds.

solid answer

~40 s

The two stages have opposite cost shapes. A feature-rich scorer costs tens of microseconds per candidate, so its bill is linear in how many candidates it sees; four million of them is tens of seconds of CPU for a single rail impression. Retrieval uses an approximate nearest-neighbour index over posting vectors that were computed in batch, so a query touches a small fraction of the corpus and its latency barely moves as the catalogue grows. So the design fixes a shortlist size as a budget — a few hundred — and asks the retrieval stage to fill it. The price of the split is that the shortlist caps everything: within one request the scorer can only reorder what retrieval handed it, so a good posting that retrieval missed is simply not in the rail.

go deeper

for a junior

Recall the shape: a cheap stage narrows millions of postings to hundreds, then an expensive stage ranks those. Be able to say why scoring everything is impossible using a per-candidate cost times the catalogue size.

for a middle

Explain the cost curves: retrieval is roughly flat in catalogue size because it searches an index, scoring is linear in candidate count. Name the shortlist size as a deliberate budget rather than an accident.

for a senior

Show that the shortlist caps achievable recall for that request, and that this is why the retrieval stage is measured on recall rather than on its own ordering. Tie the budget to the latency and spend it buys.

for a principal

Frame the funnel as where the organisation chooses to spend compute: widening retrieval buys recall the scorer can use, deepening the scorer buys precision on what arrived. The split is a cost allocation, and it should be revisited when either curve moves.

## The arithmetic that forces two stages A "jobs for you" rail has a page budget measured in tens of milliseconds and a catalogue of roughly four million open postings. A scoring function worth having — one that fetches per-seeker and per-posting features and runs a model over the pair — costs on the order of ten microseconds per posting. Four million postings times ten microseconds is **forty seconds of CPU for one request**. Spread across forty cores that is still a second of wall clock, and the rail serves thousands of requests a second. No fleet size fixes this, because the cost is incurred per request, not once per day. The funnel splits the work where the two cost curves cross: - **Retrieval cost is roughly independent of catalogue size.** An approximate nearest-neighbour index over precomputed posting vectors inspects a small fraction of the corpus per query, so doubling the catalogue costs memory and very little latency. - **Scoring cost is linear in candidates.** Each candidate means a feature fetch and a forward pass, so the shortlist size *is* the serving bill. - The architecture therefore treats the shortlist size as a **fixed budget** — a few hundred — and makes retrieval's job to fill that budget with the most plausible postings it can find. ## What the retrieval stage is allowed to know Retrieval has to rank candidates it has not enumerated, which constrains what it can compute: - Its scoring function must be **precomputable and searchable**: a seeker vector produced on the request, posting vectors produced in batch, one similarity metric the index can search over. A signal the index cannot search over cannot influence retrieval. - Beside the learned source sit **cheap lookup sources** that need no model at all: postings created in the last few hours inside the seeker's commute radius, postings matching a saved alert. They are ordinary queries against ordinary stores. - Hard constraints that would waste the budget — an expired posting, one the seeker already applied to — are applied as filters around the search rather than as a learned signal. ## The two stages side by side | | retrieval stage | scoring stage | |---|---|---| | input size | the whole catalogue, millions | the shortlist, hundreds | | cost per candidate | a fraction of a microsecond, amortised over the index | tens of microseconds plus a feature fetch | | what it may use | precomputed vectors and cheap lookups | features fetched per candidate | | output | a set of eligible candidates | an ordered list | | its failure mode | a good posting never enters the set | a good posting is placed too low | ## Recall is capped at stage one The funnel is strictly narrowing, and that is the consequence candidates most often miss: 1. Retrieval returns a shortlist. 2. The scorer permutes that shortlist; it is a ranking function over a given set, not a search over the catalogue. 3. Whatever edits the final slate operates inside the scored list. Nothing in that chain introduces a posting the first stage did not supply. Two practical consequences follow. First, **the ordering produced by retrieval barely matters** — membership is everything, because the scorer will re-order anyway. Second, **the retrieval stage's own quality metric is recall of the shortlist**, not any notion of precision: it is measured by how many of the postings that deserved a slot were present, not by how tidy its own ranking was. ## What an interviewer is listening for - The **cost asymmetry stated with a number**, not as "it would be slow". - That retrieval emits a **set under a size budget**, and that the budget is chosen deliberately. - That a miss at stage one is **unrecoverable within that request**, which is why retrieval is measured on recall. - That several cheap sources can feed the same shortlist; the learned index is not the only way in. The answer that fails is "score everything, just add machines": it misreads a per-request cost as a capacity problem, and it scales with catalogue size times traffic, which is exactly the product the funnel exists to avoid.

  • Why must the retrieval stage's scoring function be a vector comparison rather than a feature-rich model?
    Because retrieval has to rank candidates it has not enumerated. The only way to find the best few hundred out of millions without touching all of them is to make the score a geometric relation between a query point and indexed points, which an index can search. A model that needs features fetched per candidate can only run once you already have the candidate list.
  • Does adding a second retrieval source raise the scoring stage's cost?
    Not by itself. Sources are merged, de-duplicated and truncated to the same shortlist size, so a new source changes the composition of the shortlist rather than its length. Cost rises only if you also raise the depth, which is a separate decision with its own latency and spend.
  • If retrieval's own ordering barely matters, why not return a random few hundred eligible postings?
    Because recall is what the stage is judged on, and a random sample of four million has almost no chance of containing the postings that deserve the rail's 25 slots. Retrieval's ordering is throwaway, but its membership decision is the entire product of the stage.

A recruiter with four million CVs does not read them all. They pull the one drawer that matches the role, then read the fifty CVs in it properly. The pull is coarse and fast; the reading is slow and careful, and anything left in the cabinet never gets read.

saying these in an interview costs you the question

  • Just score the whole catalogue in parallel and add machines
  • The scorer can pull in a posting retrieval missed
  • Retrieval should return the final ranking, not a set
  • Retrieval and scoring can use the same per-candidate features
  • Catalogue growth hurts retrieval latency as much as scoring cost