skip to content

How do you measure whether a jobs-for-you retrieval stage is dropping postings the scorer would have ranked highly?

level: seniorimportance: must knowfreq 60%

answer

  1. two recalls, two references
  2. clicks cannot measure a miss
  3. exact search versus approximate search
  4. the scorer's top-25 over a deeper pool
  5. depth curve flattens; price the knee

basics

~20 s

Measure two different recalls: the index's recall against exact vector search, and the funnel's recall against the scorer's own top-25 over a much deeper candidate pool replayed offline. Clicks cannot measure it — a posting never retrieved was never shown.

solid answer

~50 s

Replay a sample of yesterday's rail requests offline. For each, run the production retrieval and also a much deeper pool — exhaustive over the eligible catalogue if the sample is small enough. Two numbers come out. **Index recall@k** compares the approximate search's results with exact search over the same vectors: it says whether the index is losing neighbours. **Funnel recall** scores the deep pool with the production scorer, takes its top 25 as the reference slate, and asks how many of those the production shortlist contained: it says whether the stage is losing postings that would have been shown. Clicks are not a usable reference, because a posting that retrieval missed was never rendered and so could never be clicked. The diagnostic is the pair: high index recall with low funnel recall means the towers or the source mix are wrong, not the index.

code

pseudocode · 15 lines
pseudocode
sample = 2000 rail requests replayed from yesterday's logs
index_recall = 0
funnel_recall = 0

for each request in sample:
    shortlist = production_retrieval(request)              // depth 900, filters applied
    deep_pool = all_eligible_postings(request)             // exhaustive on the sample

    exact_top = top_k_by_similarity(deep_pool, k = 900)    // same encoder, same metric
    index_recall = index_recall + size(shortlist ∩ exact_top) / 900

    reference = top_25(score_offline(deep_pool))           // what a perfect stage would feed
    funnel_recall = funnel_recall + size(reference ∩ shortlist) / 25

report index_recall / 2000, funnel_recall / 2000

go deeper

for a junior

Know that retrieval is judged on whether the good postings were in its output at all, and that you cannot learn that from clicks, because a posting that was never retrieved was never on screen.

for a middle

Distinguish the approximate index's recall against exact search from the stage's recall against what the scorer would have shown, and describe how a replayed offline sample produces each.

for a senior

Read the pair as a diagnostic, set depth from a measured recall-versus-depth curve rather than a round number, and derive the over-retrieval factor from the live survival rate after eligibility filters.

for a principal

Treat depth as a spend decision: each extra point of funnel recall costs scoring capacity, and the right depth is where that trade stops paying against the rail's actual objective.

## Two recalls, and they answer different questions "Recall" in a retrieval tier is ambiguous and the ambiguity hides bugs. Separate them explicitly: | | index recall@k | funnel recall | |---|---|---| | reference set | exact nearest neighbours of the seeker vector under the same encoder and metric | the scorer's top 25 over a much deeper candidate pool | | question it answers | is the approximate search losing neighbours it should have found? | is the stage losing postings that would have been shown? | | moves when | index parameters, shard layout, tombstone build-up change | the towers, the source mix or the depth change | | typical healthy value | high, often 0.95 and above | markedly lower, and that is normal | A tier can have excellent index recall and poor funnel recall, and that combination is the most informative reading available: the search is faithfully returning the neighbours of a seeker vector that is simply not pointing where the scorer's notion of a good posting lives. ## Building the reference set without clicks The tempting label is engagement — the postings the seeker clicked or applied to. It cannot work here, because engagement only exists for postings that were **shown**, and the whole question is about postings that were never retrieved. Any recall computed from clicks is a measurement of the funnel's own output, which is exactly the thing being audited. The workable reference is the scorer's own judgment applied to a much larger pool: 1. Sample a few thousand rail requests from logs, with the seeker features as they were at the time. 2. Build a deep pool: exhaustive over eligible postings if the sample is small, otherwise retrieval at ten or twenty times production depth. 3. Score the entire deep pool offline with the production scorer. 4. Take its top 25 — what the rail would have shown with a perfect retrieval stage — as the reference. 5. Report the fraction of that reference the production shortlist actually contained. This is expensive per request and cheap in aggregate: it runs on a sample, offline, on a schedule, not on the serving path. ## Choosing the depth: find the knee, then price it Funnel recall is a function of depth, and the curve flattens. A typical shape measured this way: - depth 400 → funnel recall 0.72 (18 of the reference 25 present) - depth 900 → 0.83 - depth 4,000 → 0.88 The first step buys eleven points of recall; the second buys five for four times the scoring cost. The depth decision is where that curve meets the latency and spend the scoring stage is willing to absorb — which is a budget conversation, not a modelling one. ## The over-retrieval factor Depth at the index is not depth at the scorer, because eligibility filters run **after** the search: postings that expired between the index build and the request, roles the seeker already applied to, employers the seeker blocked. If those remove 55% of what comes back, then handing the scorer 400 candidates requires retrieving about 900, because 400 divided by 0.45 is roughly 889. So the over-retrieval factor is measured, not guessed: - track the **survival rate** — eligible candidates out over candidates retrieved — as a live metric, per source; - set depth from the current survival rate with headroom, and alert when it moves, because a drop means the scorer is being starved without any error appearing; - where a constraint is stable and cheap to express, pushing it into the search as a pre-filter removes the attrition instead of paying for it. ## Reading the diagnostic - **Index recall high, funnel recall low** — the search is fine; the seeker tower, the posting tower or the source mix is the problem. Widening depth helps only marginally. - **Both low** — the index configuration or its freshness is the first suspect. - **Both high, rail still weak** — retrieval is supplying the right candidates and the ordering stage is the constraint; this is where the retrieval tier's owner hands the investigation on. - **Funnel recall drifting down week over week with no deploy** — look at survival rate and index freshness before looking at the model.

  • Eligibility filters remove 55% of retrieved postings. How deep must retrieval go to hand the scorer 400?
    About 900. Only 45% survive, and 400 divided by 0.45 is roughly 889, so depth is set from the measured survival rate with headroom. Track that rate live: if it falls, the scorer is starved even though every stage still returns successfully and no error is raised.
  • Why is retrieval measured on recall rather than on the precision of its own ordering?
    Because the stage that follows re-orders everything it receives, so retrieval's internal ordering is discarded. What survives from retrieval is only the membership of the shortlist, and membership is exactly what recall measures.
  • Index recall is 0.98 and funnel recall is 0.55. What do you investigate first?
    Not the index. The approximate search is faithfully returning the neighbours of the seeker vector, so the problem is what that vector means: the towers' training, or a source mix that never queries where the scorer's preferred postings live. Tuning index parameters here buys two points of a number that is already fine.

saying these in an interview costs you the question

  • Measure retrieval recall from what seekers clicked
  • Index recall against exact search is the only recall that matters
  • Deeper retrieval always improves the rail, so retrieve as deep as possible
  • Retrieve exactly the shortlist size, since filters run inside the index
  • A falling survival rate is harmless because no request errors