How do you build a query router over three RAG indexes, and how does each approach fail?
answer
- rules, centroids, classifier
- cheap and brittle to costly and general
- a miss costs more than an extra search
- the output is a set, not a choice
- route on the rewritten query
basics
~20 sThree common builds: deterministic rules, embedding centroids over labelled route examples, and an LLM classifier. Rules are exact but brittle, centroids are cheap but miss unfamiliar phrasing, classifiers generalise but cost a call and drift.
solid answer
~50 sSay you have product docs, release notes and a community forum. **Deterministic rules** — pattern matches on version strings, error codes, explicit route words — are free, instant and auditable, but cover only phrasings you anticipated. **Embedding centroids** (a semantic router) embed a set of labelled example questions per route and match the query to the nearest centroid above a threshold: sub-millisecond, no LLM call, but it needs curated examples per route and degrades on out-of-distribution phrasing, and its threshold is a real tuning burden. **An LLM classifier** generalises best and handles multi-label selection naturally — "does this new flag work?" legitimately hits docs *and* release notes — at the cost of a call, latency, non-determinism and prompt drift. The mature pattern is layered: rules first for high-precision known cases, a cheap semantic router next, an LLM only for the residual, and fan-out to all indexes when confidence is low.
code
json · 6 lines{
"query": "Does the new export flag work with SSO?",
"routes": ["product_docs", "release_notes"],
"confidence": 0.62,
"low_confidence_fallback": "search_all"
}go deeper
Know that with several indexes something has to choose which one to search, and that the choice can be made by simple rules, by comparing the query to example questions, or by asking a model.
Compare the three builds on cost, latency, determinism and generalisation, and explain why a router should be able to return more than one index for a question that legitimately spans two.
Show how you would evaluate and operate it: a labelled routing set, a per-route confusion matrix, recall weighted above precision, a low-confidence fan-out fallback, and logging misroutes as a backlog.
Own whether routing should exist. Argue the baseline of searching everything with reranking, decide when index size, cost or semantic separation justifies a classifier, and set the layering and distillation path from LLM router to cheap router.
## The decision the router makes With one index there is no routing problem. With three — product docs, release notes, a community forum — every query needs a choice, and the choice has asymmetric costs. Searching one index too many costs a little latency and some prompt tokens. Missing the only index that holds the answer costs the answer. That asymmetry should shape every design decision below. ## Three implementations **Deterministic rules.** Regex and keyword matching, sometimes on structured signals rather than the text: a semantic-version string or a "since which release" phrasing routes to release notes; an error code known to appear in the docs routes there; a question typed from inside the forum's own widget carries its origin as a feature. Rules are free, sub-millisecond, deterministic, unit-testable and explainable to a colleague. They are also brittle — they cover exactly the phrasings someone thought of, they accumulate into an unmaintainable pile, and they generalise to nothing. Their proper role is a high-precision front layer, not the whole router. **Embedding centroids (a semantic router).** For each route, collect a set of representative questions, embed them, and keep either all the vectors or their centroid. At query time, embed the query once and pick the nearest route above a similarity threshold. Cost is one embedding call and a tiny nearest-neighbour comparison; latency is negligible next to an LLM call, and behaviour is deterministic given fixed vectors. Its failures are specific: quality depends entirely on how representative the example sets are; routes whose examples are semantically close blur together; out-of-distribution phrasing — a new product name, a customer's idiosyncratic wording — sits far from every centroid; and the threshold is a genuine tuning burden, since too high sends everything to the fallback and too low routes noise confidently. It also has no reasoning, so it cannot notice that a question needs two routes for two different reasons. **An LLM classifier.** Give a model the route descriptions and the query — ideally the rewritten, standalone query — and have it return the routes to search. This generalises to phrasings nobody anticipated, handles multi-label selection naturally, can emit a confidence or a short justification, and is updated by editing route descriptions rather than by relabelling a dataset. Costs: an extra call on the critical path, per-query spend, non-determinism run to run, and the fact that it is a prompt — so it drifts with model upgrades and needs its own eval set. A common maturation step is to run the LLM router in production, log its decisions, and distill them into a small classifier or into centroid examples once you have volume. ## Multi-label routing and the confidence escape hatch With three indexes, single-label routing is often simply wrong. "Does the new export flag work with SSO?" plausibly needs release notes (when the flag shipped) and product docs (how SSO interacts), and the forum may hold the practical caveat. Design the router's output as a **set**, not a choice, and cap the set size so you do not fan out to everything by default. Give it an explicit low-confidence behaviour. Because a miss is worse than an extra search, the sensible default when the router is unsure is to search more, not less: fall back to all indexes and let ranking sort it out. That fallback is also your safety net during cold start, before you have enough labelled traffic to trust the router at all. ## Measuring a router A router is a classifier, so evaluate it like one. Build a labelled set from real queries — each labelled with the index or indexes that actually contain the answer — and report a confusion matrix rather than a single accuracy number, because the per-route error pattern is what you act on. Track precision and recall per route, and weight recall higher for the reasons above. Then measure the thing that matters end to end: does end-to-end answer quality with routing beat the naive baseline of searching everything and reranking? On small corpora it often does not, and that is a legitimate finding — routing earns its place when indexes are large, expensive, or semantically distinct enough that mixing them pollutes ranking. Online, log the chosen routes with the query and the eventual outcome. The highest-value monitoring signal is queries where the router picked one index and the answering passage turned out to live in another; those are your retraining and rule-writing backlog. ## Where the router sits In a conversational system, routing must run on the **rewritten** standalone query, not the raw turn — "what about the later one?" carries no signal for any classifier. Because both steps are cheap model calls on the same input, a common optimisation is to fold rewriting and routing into a single call that returns both the standalone question and the chosen routes, paying one round trip instead of two.
- When is routing not worth building at all?When the corpora are small, cheap to search and semantically compatible. Searching everything and letting a reranker sort it out is a strong baseline with no classifier to maintain and no misroute failure mode. Routing earns its place when indexes are large or costly, when they use different backends, or when mixing them pollutes ranking.
- How should the router behave when it is not confident?Search more, not less. A miss removes the answer entirely, while an extra index costs a little latency and some prompt tokens. So low confidence should fall back to a broader set — often all indexes — with reranking to sort the merged results. Reserve narrow routing for cases the router is confident about.
- Why must routing run after conversational rewriting?Because the raw follow-up turn carries almost no routing signal. "What about the later one?" has no product name, no version, no intent markers, so a centroid router matches noise and a classifier guesses. The rewritten standalone question contains the entities and intent the router needs — and folding both into one model call keeps the latency cost to a single round trip.
- What is the single most useful production signal for improving a router?Queries where the router chose one index and the passage that actually answered lived in another. Logging chosen routes alongside the eventual grounding source turns that into a concrete backlog: new centroid examples, a new rule, or a sharpened route description. It also tells you which route pairs are being confused, which a single accuracy number hides.
saying these in an interview costs you the question
- Assuming a router must pick exactly one index
- Treating routing accuracy as a single number rather than per-route
- Routing on the raw follow-up turn instead of the rewrite
- Believing an LLM classifier is always the right choice
- Falling back to the smallest index when confidence is low