Why does a marketplace's typed product search need a relevance cutoff that can return nothing, when its browse feed does not?
answer
- one surface carries a stated intent
- nearest neighbours are always returned
- non-empty is not the same as relevant
- raw magnitude is not comparable across queries
- relax and label before showing nothing
basics
~20 sA typed query asserts an intent the catalogue may not satisfy, and retrieval will hand back its nearest candidates regardless, so without a cutoff the surface silently answers a different question. A browse feed makes no such assertion to contradict.
solid answer
~50 sEmbedding retrieval returns its k nearest items for any query, however unusual, and a relaxed keyword match returns partial hits — so "retrieval returned results" is not evidence that the catalogue contains what was asked for. On the search surface that matters, because the shopper stated an intent: showing the least-bad items as if they were answers is worse than showing none. The browse feed has no stated intent to contradict, so it has no equivalent obligation; it fills the page with the best available items and is judged on engagement. The cutoff has three requirements: a **relevance score comparable across queries** (raw keyword magnitude is not — it moves with query length and term rarity), a **relaxation ladder** tried before the empty state, and **impression logging that stops at the cutoff**, so items never shown are not recorded as rejected.
go deeper
Remember that retrieval hands back its nearest candidates for any query, so getting results back does not mean the catalogue actually has what was asked for.
Explain why a raw score threshold fails across queries and what calibrated or per-query-normalised quantity you would apply the cutoff to instead.
Lay out the full ladder — correct, relax, label, then empty — and state the logging rule that stops impressions at the cutoff so the next model is not poisoned.
Own the trade explicitly: withheld results are lost revenue now, shown near-misses are lost trust later, and the position should be measured by empty-rate against reformulation-rate.
## The asymmetry between the two surfaces A browse feed and a typed search page can run the same funnel and still owe the shopper different things. | | browse feed | typed product search | |---|---|---| | what the shopper asserted | nothing explicit | a specific intent, in words and facets | | what a weak item costs | a wasted slot | a visible wrong answer to a stated question | | is an empty result meaningful? | rarely; there is almost always something plausible | yes — "we do not sell this" is a true and useful answer | | judged by | engagement on the session | whether results match the stated intent | The feed can under-fill in edge cases — a brand-new account where every retrieval source is empty — but it has no assertion to violate, so it has no reason to refuse to show items. The search surface does. ## Why retrieval cannot decide this for you The trap is treating a non-empty candidate set as proof of relevance. Embedding retrieval is a nearest-neighbour operation: ask for twenty and you get the twenty closest vectors, whether the closest is an exact match or the least-distant item in a catalogue that contains nothing of the kind. Keyword retrieval can return zero, but a funnel tuned to avoid empty pages usually relaxes the match — any-of instead of all-of, stemmed or fuzzy terms — which brings partial hits back in. So both sources have a strong bias toward handing you *something*, and the funnel needs an explicit decision about whether that something is good enough to show. ## Why the cutoff cannot be a raw score threshold The naive implementation is a fixed threshold on whatever score the retrieval or scoring stage produced. It does not survive contact with real traffic: - **Keyword score magnitude moves with the query.** More terms and rarer terms produce larger values, so a threshold tuned on two-word queries misfires on six-word ones. - **Similarity bands are region-dependent.** A given similarity value means "very close" in a sparse region of the embedding space and "unrelated" in a dense one. - **The scoring stage is usually trained to rank, not to measure.** A model trained to order items produces scores whose absolute value carries little meaning; two items can be correctly ordered with both being wrong. So the cutoff needs a quantity that is comparable across queries. The practical options are: calibrate the scoring stage's output into a probability against relevance labels (isotonic regression or Platt scaling on held-out judgments); train a separate lightweight relevance classifier on query-item pairs whose output is a probability by construction; or use a per-query normalised quantity such as the gap between the top item and the tail, which detects "nothing stands out" without needing an absolute scale. ## What has to exist before the empty state Returning nothing is the last rung, not the first: 1. **Correct the query** — spelling, folding, synonym expansion — and try again. 2. **Relax the constraint** — drop the weakest free-text term, loosen a soft attribute; keep ticked facets intact. 3. **Offer adjacent results, labelled as such** — "no matches for X; shoppers also looked at Y". Labelling is what separates a helpful fallback from a silent substitution. 4. **Return nothing, and say why**, with the query echoed and the facets shown so the shopper can see which constraint to loosen. ## The logging rule that goes with the cutoff Once a cutoff exists, the impression log must stop where the page stops. Items scored below the cutoff were never rendered, so recording them as impressions with no click teaches the next model that those items are unattractive for that query, when in truth the shopper never saw them. The same rule covers the relaxed rung: results shown under a "similar items" label were presented in a different context, and should be logged with that context rather than pooled with exact matches. ## Where the honest disagreement lies The cutoff is a business decision as much as a technical one. Every result withheld is revenue that certainly will not happen; every irrelevant result shown is a small deposit against trust in the search box, paid back later in queries not typed. Different catalogues land in different places, and the design should make the position **explicit and measurable** — track the empty-result rate and the reformulation-and-abandon rate together, because moving the cutoff trades one directly against the other.
- What quantity can the cutoff actually be applied to?Something comparable across queries: a calibrated relevance probability from isotonic regression or Platt scaling over held-out judgments, a small relevance classifier trained on query-item pairs, or a per-query normalised signal such as the top-to-tail score gap. A raw ranking score will not do, because a ranker is trained to order items, not to say whether any of them are right.
- How do you know whether the cutoff is set in the right place?Track the empty-result rate and the query reformulation-and-abandon rate together and watch them move in opposite directions as you shift the cutoff. Tightening it raises empty results and should lower reformulations; if both rise, the cutoff is not the problem and the retrieval tier's recall is.
- The browse feed reserves a slot for exploratory items. Why is the same slot harmful on the search page?Exploration costs an ambiguous impression on a feed, where intent was only inferred. On a query surface the shopper has stated exactly what they want, so an exploratory item is a visible wrong answer whose cost lands immediately as a reformulation or an abandon. Gather exposure for under-served inventory where intent is weak — the feed and broad head queries — not on precise tail queries.
A librarian asked for a specific title should say "we do not have it" rather than hand over a book from the next shelf. The display table by the entrance is under no such obligation — its job is simply to hold something worth picking up.
saying these in an interview costs you the question
- Treating a non-empty candidate set as evidence the catalogue has a match
- Setting the cutoff as a fixed threshold on a raw ranking score
- Filling the page with near-misses rather than showing an empty state
- Substituting adjacent results without labelling them as substitutes
- Logging below-cutoff items as impressions with no click
- Assuming a browse feed needs the same relevance cutoff