skip to content

In RAG, when relevance alone picks the wrong passages, how do you design selection policy?

level: principalimportance: should knowfreq 33%

answer

  1. similarity is a proxy for the real goal
  2. newest wording is not newest truth
  3. ranking answers a different question than selection
  4. constraints, caps and reserved slots
  5. log why each candidate was dropped

basics

~20 s

Add an explicit policy layer after ranking that turns scores into a chosen set under constraints: recency windows, source authority, per-document caps, required coverage slots and a relevance floor. Then evaluate the policy on answer outcomes, not ranking order.

solid answer

~50 s

Ranking answers "what resembles the query?", which is often not the same as "what should the model read?". Make selection a distinct, reviewable stage that consumes ranked candidates and applies stated constraints: a relevance floor so nothing weak sneaks through, freshness rules where stale material is actively harmful, authority preference when a canonical source and a stale forum post both match, per-document and per-source caps so one document cannot fill the set, and reserved slots when the answer must span several facets. In a shipment-exceptions assistant, the current carrier contract must outrank a superseded one that matches the wording better. Keep the policy declarative and versioned rather than scattered through prompt text, because it encodes business rules that legal, compliance and support will want to inspect. Judge it on end-to-end answer quality and on incident review, since ranking metrics will look worse by construction.

go deeper

for a junior

Know that the highest-scoring passage is not always the right one to use — an outdated document can match the wording better than the current one. Be able to name recency and source trust as extra criteria.

for a middle

Explain how a selection stage differs from ranking: hard filters, a relevance floor, per-document caps and freshness rules applied over already-scored candidates, and why those cannot be learned by the ranker.

for a senior

Show you can diagnose which stage failed — never retrieved, ranked low, or rejected by policy — by logging drop reasons, and can slice the answer eval by the cases a new rule targets rather than trusting aggregate scores.

for a principal

Own the policy as a business artifact: define who reviews it, version it with the corpus, accept that ranking metrics will regress by design, prune rules on a schedule, and make "why did the system cite that document?" answerable from logs for audit.

## Relevance is a proxy, not the objective Every scoring stage in a retrieval pipeline optimises textual and semantic similarity to the query. That is a proxy for the real objective, which is: *which passages, read together, let the model produce a correct, current, appropriately sourced answer?* Wherever those diverge, relevance ranking confidently picks the wrong thing, and no amount of reranker quality fixes it — because the reranker is doing its job correctly. The divergences recur across domains: - **Staleness.** A superseded policy or an expired contract often matches a question *better* than its replacement, because the replacement is written in newer language. Similarity has no notion of "current". - **Authority.** A canonical runbook and a three-year-old chat thread can be near-identical in wording; only one should ground an answer. - **Coverage.** A question with several facets — cause, remedy, and who to notify — can be filled entirely by passages about the cause, because those are the most on-topic. - **Entitlement.** The best passage may be one this user is not permitted to see, which is a correctness and compliance issue rather than a ranking one. - **Confidence.** Sometimes the right selection is the empty set. ## Make selection its own stage The design move is to stop treating "take the top m" as a line of code and start treating selection as a stage with its own inputs, rules and tests. It consumes ranked, scored candidates plus their metadata, and emits the final set plus a record of why each passage was kept or dropped. Typical rules in such a layer: - **A relevance floor**, so a slot is never filled just because it exists. - **Hard filters**, applied before anything else: access control, tenancy, jurisdiction, document status. These are correctness constraints, not preferences, and pushing them into the retrieval query is usually better than filtering afterwards. - **Recency policy**: a hard cutoff where old material is dangerous, or a soft decay that discounts age when it is merely less useful. - **Authority tiers**: a preference ordering over sources, or a tie-break that prefers the canonical copy when two candidates say the same thing. - **Caps**: at most so many passages per document or per source, which prevents one long document from occupying the whole set. - **Reserved slots**: for multi-facet questions, require the set to include at least one passage of each required kind. ## The costs, stated honestly A policy layer is not free and a principal-level answer says so. **Ranking metrics will get worse.** You are deliberately demoting the highest-scoring candidate. If the team's dashboard is a ranking metric, the policy will look like a regression forever. Move the headline metric to end-to-end answer quality before shipping the policy, or it will be reverted by someone reading the dashboard. **Rules accumulate and interact.** Each incident tempts a new rule. Six months later, no one can predict what the selector will do, and the rules contradict each other. Mitigate with a fixed evaluation order, a bounded rule count, an owner, and a periodic review that deletes rules whose motivating incident is gone. **Rules encode assumptions that expire.** "Prefer documents from the compliance space" is right until the team reorganises. Version the policy alongside the corpus, and treat it as a configuration artifact with a changelog rather than as prompt text. **Hard-coding beats the model's judgment — sometimes wrongly.** A strict recency cutoff makes historical questions unanswerable. Where the right behaviour depends on the query ("what did the policy say last year?"), the policy must branch on query intent rather than apply a blanket rule, which means intent classification becomes a dependency with its own error rate. ## Who owns it Selection policy is where business rules meet the pipeline, so it should be legible to people who are not on the retrieval team. Compliance cares that superseded documents cannot ground an answer; support leadership cares that the canonical runbook wins; legal cares about jurisdiction filters. Expressing the policy declaratively — a small, ordered, readable set of constraints with a version — lets those stakeholders review it, and lets you answer the audit question "why did the system cite that document?" from a log rather than from a re-run. ## Evaluating it Evaluate outcomes, not order. - **Slice the answer eval by the cases the policy targets**: questions with a superseded predecessor in the corpus, multi-facet questions, questions with a permission-restricted best match. Aggregate scores will hide the exact behaviour you built the policy for. - **Guard against collateral damage** with a broad regression set, since a rule aimed at one failure mode routinely degrades unrelated queries. - **Log the drop reason per candidate.** When an answer is wrong, you need to know whether the passage was never retrieved, was ranked low, or was retrieved, ranked well and then rejected by policy — three different fixes. - **Feed incidents back.** Each wrong-source incident should end with either a new rule and its eval case, or an explicit decision not to add one. The strongest version of this answer accepts that the policy is a living business artifact: it will grow, it must be pruned, its ranking metrics will look bad, and its justification lives in end-to-end outcomes and incident review.

  • A rule preferring documents updated in the last 12 months fixes stale answers but breaks historical questions. How do you resolve it?
    Make the rule conditional on query intent rather than global. Detect that the question is historical — explicit dates, past-tense framing, "what did X used to say" — and switch to a policy that permits or even prefers superseded documents, labelling them as such in the injected context. Blanket recency rules are a symptom of treating one query class as the whole traffic; branch, and evaluate both branches separately.
  • How do you keep a selection policy from turning into an unmaintainable pile of rules?
    Treat it as versioned configuration with an owner, a fixed evaluation order, and an eval case attached to every rule. Review periodically and delete rules whose motivating incident no longer exists or whose eval case now passes without them. Cap the rule count deliberately — if a new rule cannot displace an old one, the underlying problem probably belongs in chunking, metadata or the corpus itself.
  • Should access-control filtering happen in the selection stage?
    Preferably earlier, in the retrieval query, so restricted documents never enter the candidate pool. Filtering after ranking is fragile — one code path that forgets it leaks — and it also distorts the cascade, since restricted documents consume candidate slots and skew score distributions. Treat entitlement as a hard pre-filter and correctness constraint, not as one preference among many in the selection policy.

saying these in an interview costs you the question

  • Assuming a better reranker can encode business rules like recency or authority
  • Judging a selection policy by ranking metrics that it intentionally degrades
  • Burying selection rules in prompt text rather than versioned configuration
  • Applying a blanket recency cutoff that makes historical questions unanswerable
  • Filtering permissions after ranking instead of inside the retrieval query

context