skip to content

When scoped retrieval returns zero documents after metadata filtering, what should the system do?

level: seniorimportance: should knowfreq 38%

answer

  1. empty is a state, not zero rows
  2. two empties: no match vs all below threshold
  3. relax preferences, never boundaries
  4. fixed fallback ladder, recorded in the trace
  5. empty-scope rate signals ingestion regressions

basics

~20 s

Return an explicit empty-scope signal rather than silently proceeding. Widening is permitted only on preference predicates such as recency or document type, and the relaxation must be labelled; security predicates like tenant and ACL are never relaxed, and the generation step must know the context is empty.

solid answer

~60 s

First, make emptiness explicit. An empty result must reach the generation step as a distinct state — "no documents in scope" — not as an empty string spliced into a prompt, because a model handed no context will answer from parametric memory and sound exactly as confident as when it was grounded. Second, separate the two kinds of empty. Nothing matched the predicate at all is different from documents matched but everything scored below the relevance threshold, and they lead to different fixes: the first is a scope or metadata problem, the second is a retrieval-quality problem. Third, widen only what is safe. Relaxing a recency window on a release-notes index from the last two minor versions to older ones is legitimate if the results are labelled as out-of-window. Relaxing tenant, ACL or region is never legitimate — that converts an unhelpful answer into a disclosure. Retry with fallbacks in a fixed, auditable order, never by dropping the whole predicate set. Finally, treat the empty-scope rate as a monitored metric; a spike usually means a metadata field stopped being populated.

go deeper

for a junior

Know that an empty scoped result must be handled deliberately — the generation step has to be told there was no context, rather than receiving an empty block and answering from memory.

for a middle

Explain the two distinct empties and which predicates are safe to widen: recency, language and document type yes, with the results labelled; tenant, ACL and region never.

for a senior

Design the fallback as a fixed, auditable ladder recorded in the trace, and operate empty-scope and widening rates as metrics that catch ingestion regressions and wrong default windows before users complain.

for a principal

Own the policy: what the product promises when it has nothing in scope, which relaxations are permitted for which customer segments, and how the audit trail evidences that boundary predicates were never escaped.

## Emptiness is a state, not an absence Once scope is enforced, some queries legitimately match nothing: the user asked about a region they have no documents for, the version window excludes everything relevant, an ACL admits no documents on that subject. The dangerous thing is not the empty result — it is the pipeline treating it as an ordinary result that happens to have zero rows. A retrieval function that returns an empty list, whose caller joins that list into a context block and interpolates it into a prompt, produces a prompt with an empty context section and instructions that assume grounding. The model then answers from parametric memory, in the same register and with the same confidence as a grounded answer, and often with a fabricated-looking citation because the prompt asked for citations. The empty result never surfaced as a decision point. So the first rule is structural: retrieval returns a *typed outcome* — results, empty-scope, or below-threshold — and every caller must handle all three explicitly. ## Two different empties **Nothing matched the predicate.** No document in the collection satisfies the scope conditions. This is a metadata or scope problem: a field that was never populated for the older half of the corpus, a version window that has drifted past all available documents, a tenant with no ingested content, or a predicate combining conditions no document satisfies together. **Documents matched, but none scored above threshold.** The scope is populated and healthy; the query simply has no good answer in it. This is a retrieval-quality problem, and the diagnosis and fix are entirely different — different chunking, different query formulation, a different corpus. Collapsing them into one "no results" branch loses the single most useful diagnostic signal you have. Log the matched-count separately from the returned-count. ## What may be widened, and what may not The split follows the boundary/preference distinction. **Preference predicates may be relaxed** — a recency window, a document-type restriction, a language filter, a source-system restriction. A release-notes assistant scoped to the last two minor versions can, on an empty result, legitimately widen to older versions, provided the retrieved passages are marked as coming from outside the intended window so the downstream step can qualify what it says. **Boundary predicates may never be relaxed** — tenant, ACL group, region, classification. Retrying without them, softening them into score penalties, or returning "the nearest documents anyway, with a warning" all convert an unhelpful answer into a disclosure. The correct behaviour when a boundary predicate empties the result is to stop with an empty scope, every time. When you do widen, do it as a small fixed ladder — one or two predefined relaxations, in a deterministic order, each recorded in the trace so an audit can reconstruct which scope actually produced the answer. Unbounded or model-driven widening ("the agent decides what to loosen") is how boundary predicates get dropped by accident, because the model is optimising for producing an answer, not for staying in scope. ## Do not paper over it Several popular reflexes make things worse. Silently falling back to an unfiltered search is the disclosure case above. Falling back to the model's own knowledge without saying so produces an ungrounded answer indistinguishable from a grounded one. Lowering k or the score threshold to "get something back" fills the context with the least relevant material in the corpus, which is worse than nothing because it looks like evidence. Retrying the same query repeatedly costs latency and changes nothing, since the predicate is deterministic. ## Operating it Empty-scope rate is a first-class metric, tracked per predicate field and per tenant. A step change almost always means an ingestion regression — a producer stopped setting a version field, a reindex dropped a payload attribute, a tenant's connector broke — and it is visible in the empty-scope rate long before anyone files a quality complaint. Alert on the derivative, not the absolute level, since a healthy system has a nonzero baseline. Also track the widening rate. If half of all queries are hitting the recency fallback, your default window is wrong and should be changed rather than repeatedly escaped at runtime. And sample empty-scope queries for review: they are the cheapest available list of what your corpus does not cover. What the user is finally told — how the shortfall is phrased and whether an answer is attempted at all — belongs to the generation and attribution step, not to retrieval. Retrieval's job ends at handing that step an unambiguous, labelled statement of what was in scope and what was not.

  • How do you distinguish 'nothing matched the filter' from 'matches existed but all scored poorly'?
    Have the retrieval layer report matched-count separately from returned-count. A nonzero matched-count with an empty returned set means the scope is healthy and the query has no good answer in it — a retrieval-quality problem. A zero matched-count means the scope itself is empty, which points at metadata: an unpopulated field, a drifted version window, or a tenant with no ingested content.
  • An agent decides its own retrieval filters. Can it loosen them when a search comes back empty?
    Only within predicates you have designated as relaxable, and ideally only by choosing from a fixed menu the tool exposes. A model optimising to produce an answer will happily drop whatever is blocking it, and it cannot reliably tell a recency window from an ACL. Intersect any model-supplied filter with the session's boundary predicates server-side so loosening is structurally impossible for the ones that matter.
  • What does a sudden spike in empty-scope rate usually indicate?
    An ingestion regression, not a change in user behaviour. A producer stopped populating a field used in predicates, a reindex dropped a payload attribute, or a tenant's connector broke — and every scoped query for the affected slice now matches nothing. Because a healthy system has a nonzero baseline, alert on the rate of change per predicate field and per tenant rather than on an absolute level.

saying these in an interview costs you the question

  • Retries the query without the ACL or tenant predicate
  • Splices an empty context block into the prompt unchanged
  • Lowers the score threshold until something comes back
  • Treats zero matches and zero above-threshold as one case
  • Lets the model choose which filters to loosen

context