skip to content

Why should a RAG chatbot route greetings and "can you repeat that" to no retrieval?

level: juniorimportance: should knowfreq 42%

answer

  1. not every turn is a question
  2. similarity search never returns nothing
  3. irrelevant context competes for attention
  4. the two errors are not equal
  5. keep the skip list narrow

basics

~10 s

Greetings, thanks and "can you repeat that" have no answer in the corpus. Retrieving anyway costs an embedding call and search latency, and injects irrelevant passages the model may weave into its reply.

solid answer

~50 s

Not every turn is an information request. "Hi", "thanks, that helped" and "can you say that again?" are answered from the conversation itself, not from documents. Routing them to a no-retrieval short-circuit saves the embedding call, the vector search, any reranking and a large chunk of prompt tokens — but the stronger argument is precision. Forced retrieval always returns its top-k, so a greeting pulls in whatever passages sit nearest to it in embedding space, and a model instructed to ground its answers in context may drag that irrelevant material into a reply that should have been one friendly line. The escape hatch must be conservative, though: wrongly skipping retrieval on a real question produces an ungrounded, possibly hallucinated answer, which is far worse than one wasted search. Keep the no-retrieve class narrow and explicit, and retrieve whenever unsure.

go deeper

for a junior

Be ready to say that greetings and requests about the previous answer have no answer in the documents, so searching for them just wastes time and drags in irrelevant text.

for a middle

Explain that similarity search always returns its top-k, so a forced retrieval injects arbitrary passages into a grounded-generation prompt, and describe a cheap rule or classifier that short-circuits the obvious cases.

for a senior

Argue the asymmetry: a false skip destroys grounding while a false retrieve costs milliseconds, so bias toward retrieving, keep the skip class explicit, and monitor skipped turns for rephrasing as the false-skip signature.

for a principal

Own where the decision lives — a deterministic classifier, a route label on the existing router, or a capability the model may decline — and set the policy for how much cost you are willing to spend to keep grounding guaranteed.

## The class of turns that has no answer in your corpus A RAG pipeline built as a straight line — embed the turn, search, stuff the results into the prompt, generate — treats every user message as an information request. Real conversations are not like that. A meaningful share of turns are social ("hi", "thanks, that's exactly what I needed"), meta ("can you repeat that?", "say it shorter", "in French please"), or corrective ("no, I meant the other booking"). None of them has an answer sitting in the document corpus. Their answer, where they have one, lives in the conversation that already happened. A router with a no-retrieval short-circuit recognises this class and skips straight to generation with the conversation as context. ## What skipping actually saves **Latency and cost.** You avoid an embedding call, an ANN search, optional reranking and, in a conversational pipeline, often the rewriting call too. On a chat where a noticeable fraction of turns are social or meta, that is real money and a visibly snappier reply to "thanks". **Precision — the bigger win.** Similarity search does not have a concept of "nothing matched". Given "thanks!" it returns its k nearest passages regardless, and those passages are essentially arbitrary. Now the generator, usually under an instruction to answer from the provided context and prefer it over its own knowledge, receives irrelevant material. Sometimes it ignores it. Sometimes it produces a reply that awkwardly volunteers refund policy in response to a thank-you, which reads as broken. Injecting unrelated context is not free — it competes for attention with the instructions you actually care about. **Correctness on meta-requests.** "Can you repeat that?" is the sharpest example. The correct answer is a restatement of the previous assistant turn. Retrieval cannot supply it, and passages fetched by similarity to the phrase "can you repeat that" actively point the model away from the right behaviour. ## The asymmetry that governs the design The two errors are not equally bad. A **false retrieve** — searching for a greeting — costs a few hundred milliseconds and some tokens, and usually produces a fine answer anyway. A **false skip** — deciding a genuine question needs no documents — removes grounding entirely and hands the user whatever the model remembers, which is the failure RAG exists to prevent, delivered with full confidence and no citation. So the escape hatch should be conservative by construction: an explicit, narrow allowlist of no-retrieve intents rather than a general "does this need retrieval?" judgement; a bias toward retrieving whenever the classifier is unsure; and monitoring for the false-skip case, since it is the one that silently harms users. ## How it is implemented Because the class is narrow and the phrasings are conventional, this rarely needs a large model. Common implementations, roughly in order of cost: - **Rules and patterns** for the obvious cases: short turns matching greeting, thanks and farewell templates, or explicit meta-commands about the previous answer. Fast, testable, and adequate for a surprising share of traffic. - **A small classifier or embedding centroid** with a `no_retrieval` route alongside the content routes, so the short-circuit is just one more label the existing router can emit — this is the tidiest integration, since one component makes one decision. - **A single combined call** in pipelines that already rewrite conversationally: the same model returns the standalone question, the chosen indexes and a flag for "no retrieval needed", for one round trip total. A related pattern worth mentioning is making retrieval a capability the model can decline to invoke rather than a fixed pipeline stage, which reaches the same outcome by a different route; the deterministic classifier is simply the cheaper and more predictable version when the no-retrieve class is well defined. ## Measuring it Label a sample of production turns as retrieve or no-retrieve and score the classifier as a binary decision, reporting the two error rates separately rather than as one accuracy figure — they have very different costs. In production, track the share of turns short-circuited (if it drifts upward, something is over-triggering) and audit skipped turns that were followed by user rephrasing or dissatisfaction, which is the observable signature of a false skip. Keep a per-turn log of the routing decision so any bad answer can be traced back to whether it was grounded at all.

  • Which error is worse — retrieving unnecessarily or skipping retrieval on a real question?
    Skipping is far worse. An unnecessary search costs a little latency and a few tokens, and the answer is usually fine anyway. A wrong skip strips grounding from a genuine question, so the user gets an unsourced answer from parametric memory — exactly the failure RAG was built to prevent. Design the short-circuit to retrieve whenever unsure.
  • Why can't you just let the retriever return nothing for a greeting?
    Because similarity search has no notion of "nothing matched" — it returns its k nearest neighbours regardless of absolute distance. You can approximate a null result with a score threshold, but thresholds are corpus-dependent and brittle, and you have already paid the embedding and search cost by the time you apply one.
  • How would you notice the short-circuit misfiring in production?
    Log the routing decision per turn. Then watch two signals: the share of turns skipped, which should be stable, and skipped turns followed by a user rephrasing or complaining — the observable signature of a real question answered without grounding. Sampling skipped turns for manual review catches drift after a prompt or model change.

saying these in an interview costs you the question

  • Assuming every user turn requires a document lookup
  • Expecting vector search to return nothing when nothing is relevant
  • Making the no-retrieve classifier aggressive to save cost
  • Treating retrieve and skip errors as equally costly
  • Trying to answer "can you repeat that" from the corpus

context