A jobs-for-you retrieval stage draws candidates from three sources — how do you merge and de-duplicate them?
answer
- three lists, one shortlist
- different units, same positions
- fuse on rank, not on value
- collapse ids before truncating
- carry provenance for attribution
basics
~20 sMerge on rank, not on raw score: cosine similarity, recency order and an alert match are incomparable numbers. Reciprocal rank fusion or fixed per-source quotas produce one list; collapse duplicates on a canonical posting id before truncating.
solid answer
~50 sEach source speaks its own units — the embedding source returns a similarity, the recent-in-radius source returns an ordering by posting time, the saved-alert source returns an unordered match set. Comparing those numbers directly lets whichever scale happens to be largest dominate the merged list. Two policies work. **Rank fusion** ignores the values and sums `1 / (K + rank)` across the sources that returned a candidate, so agreement between sources lifts a posting and no scale dominates. **Quotas** give each source a fixed share of the shortlist, which guarantees a weak source still contributes. Then de-duplicate: the same posting can come from all three, and the same job is often reposted under a new id, so collapse on a canonical id **before** truncating to the shortlist size, and keep the set of contributing sources on each candidate for logging.
code
pseudocode · 17 linesSHORTLIST = 400
K = 60 // rank-fusion smoothing constant
fused = empty map keyed by canonical posting id
for each source in [embedding_neighbours, recent_in_radius, saved_alert]:
results = source.fetch(seeker, depth = 900, deadline = 15 ms)
rank = 1
for each posting in results:
id = canonical_id(posting) // collapses reposts of one job
entry = fused.get_or_create(id)
entry.score = entry.score + 1 / (K + rank)
entry.sources.add(source.name) // provenance for attribution
rank = rank + 1
candidates = sort(fused.values) by score descending
return first SHORTLIST of candidatesgo deeper
Recall that a retrieval stage usually queries several sources and has to produce one list. Know that the same posting can come back from more than one of them and must be collapsed.
Explain why raw scores from different sources cannot be compared, and describe one concrete merge policy end to end — rank fusion's sum over positions, or a quota split — including where de-duplication sits in the order of operations.
Show the operational consequences: per-source deadlines with partial results, provenance logged on each candidate, and de-duplication before truncation so the shortlist budget is not quietly spent on copies of one job.
Treat the source mix as a portfolio: each source costs latency and shortlist slots and should justify them by measured contribution to the final slate, with a way to retire one that only duplicates the learned source.
## Why a retrieval stage has more than one source A learned two-tower source is good at "postings like the ones this seeker engages with" and bad at everything that has no history behind it. So a production retrieval stage for a jobs rail typically queries three things in parallel: - the **approximate nearest-neighbour index** of posting vectors, around the seeker vector computed on the request; - a **recent-in-radius source**: postings created in the last few hours inside the seeker's commute area, ordered by posting time; - a **saved-alert source**: postings matching criteria the seeker explicitly stored, such as a title family and a salary floor. Each one is a cheap query with its own latency. The design question is not which to keep but how three result lists become one shortlist. ## The scores are not comparable The embedding source returns a similarity in a bounded range; the recency source returns a position in a time ordering; the alert source returns a boolean match. These measure different things, and none of them is a probability of anything. Merging on raw values means the source whose numbers happen to sit highest wins every tie, and the balance between sources shifts silently whenever a model or a source is changed. So merge on **position**, which every source produces honestly: | policy | how it combines | what it guarantees | where it hurts | |---|---|---|---| | reciprocal rank fusion | sums `1 / (K + rank)` over the sources that returned the candidate | agreement across sources is rewarded; no scale dominates | a source that is right but always ranks late contributes little | | per-source quotas | fixed slot counts, filled from each source's own order | every source has a floor in the shortlist | the split is a hand-tuned constant that ages | | weighted quotas | quotas adjusted by each source's measured contribution to the final slate | adapts to what actually gets shown | needs the attribution logging to be trustworthy | `K` in rank fusion is a smoothing constant, commonly around 60; it flattens the difference between the very top ranks so a single source cannot own the head of the merged list. ## De-duplication is part of the merge, not an afterthought Two distinct collapses happen here: 1. **The same posting from several sources.** This is the good case: it is evidence, and in rank fusion it is exactly what raises the candidate. Merge the entries, sum the contributions, keep one candidate. 2. **The same job under several identifiers.** Boards re-post a role weekly to appear fresh, and one role can arrive through several feeds. These are different ids pointing at one job, so they need a canonical id — a fingerprint over employer, normalised title and location — computed when the posting is ingested, not at query time. Both collapses must happen **before truncation**. Truncating first and de-duplicating after means a shortlist nominally of 400 hands the scorer 300 distinct postings, and the missing hundred are silently paid for in the latency budget anyway. Note the boundary: trimming a final slate that reads as five near-identical warehouse roles is a later list-editing concern; what the retrieval stage owns is not paying twice for one job. ## Slow sources and partial results The three sources are queried in parallel, so the stage's latency is the slowest of them. Give each source its own deadline and take what has arrived: - a missing embedding source is a serious quality event and should be visible as one; - a missing alert source usually costs a handful of candidates and should not fail the request; - log which sources were present, because a quiet timeout on one source looks exactly like a model regression when the rail's metrics move. ## Keep provenance on every candidate Carry the set of contributing sources through the funnel. It costs almost nothing and it is what makes three later questions answerable: which source supplied the postings that ended up in the final slate, whether a source can be switched off, and whether a source is merely duplicating what the embedding index already found. Without it, source attribution has to be reconstructed by replaying every request.
- When are per-source quotas a better merge policy than rank fusion?When a source must be guaranteed a share for a reason outside relevance — a newly launched source that needs exposure to generate any data at all, or a source whose ranks are not comparable to the others because it returns an unordered match set. Quotas make the contract explicit at the cost of a constant somebody has to maintain.
- Why compute the canonical posting id at ingestion rather than during the merge?Because the fingerprint needs the full posting body — employer, normalised title, location — and the retrieval stage only holds identifiers and vectors. Doing it at ingestion also means the index itself never holds two vectors for one job, which saves both memory and shortlist slots.
saying these in an interview costs you the question
- Normalise every source's score to 0-1 and average them
- Duplicates are harmless because the scorer ranks copies identically
- De-duplicate after truncating to the shortlist size
- One source timing out should fail the whole rail request
- A posting in two sources should be counted only once, at its better rank