skip to content

Semantic-only search regressed on brand and SKU queries in an A/B — what do you do?

level: seniorimportance: must knowfreq 62%

answer

  1. identifiers denote, they do not mean
  2. rare exact tokens are where dense loses
  3. the two legs fail on different queries
  4. segment the experiment by query shape
  5. exact-match path can short-circuit

basics

~20 s

Add a lexical retrieval leg. Dense embeddings compress meaning and blur rare exact tokens like part numbers and brand names, so keep a keyword index (typically BM25) running alongside the vector search and merge both candidate lists before ranking.

solid answer

~50 s

That regression is the classic signature of dense-only retrieval: an embedding is a lossy summary of meaning, and a model code such as `WD-42X`, a brand name, or a rare proper noun carries almost no semantic content — it is an identifier. The model has seen it rarely, tokenises it into fragments, and places near-identical SKUs in nearly the same place, so the exact one the shopper typed is not reliably first. Lexical retrieval has the opposite bias: it is exact-match and rewards rare terms heavily, which is precisely the case dense search loses. The fix is **hybrid retrieval** — run both a vector search and a keyword search, then combine the two candidate lists into one shortlist (the fusion method itself is a separate design choice). Before shipping it, segment your A/B by query shape so you can prove the lexical leg recovers the regressed segment without giving back the natural-language wins.

go deeper

for a junior

Be able to say that embeddings capture meaning and are weak on exact strings like part numbers, and that real search usually keeps a keyword index alongside the vector one.

for a middle

Explain the mechanism: identifiers carry little semantic content, are split into subword fragments, and near-identical codes produce near-identical vectors. Explain that lexical scoring has the opposite bias — exact and rare-term-sensitive — so the two are complementary.

for a senior

Show the diagnosis before the fix. Segment the experiment by query shape, check whether the correct item was even retrieved versus merely mis-ranked, and choose between always running both legs and routing by query pattern with the failure modes of each.

for a principal

Own the cost side. Two indexes mean two ingestion paths, drift between them, a fusion layer with tunable parameters and a larger serving footprint. Be ready to argue when that operational surface is justified and to set the guardrail that no launch may regress a high-intent query segment.

## Why dense retrieval loses exactly here An embedding is a fixed-size, lossy compression of meaning. That compression is what makes semantic search work — it is why "walkable neighbourhood near good schools" can match a listing that uses none of those words. It is also why it fails on identifiers. Three mechanisms combine: 1. **Identifiers carry no semantics to compress.** `WD-42X-B` does not *mean* anything; it denotes. There is no meaningful neighbourhood for it to occupy, so where it lands in the space is close to arbitrary. 2. **Rare tokens are fragmented.** A model code gets split into several subword pieces, most of which are shared with thousands of unrelated strings. The resulting vector is dominated by whatever surrounding text existed, not by the code. 3. **Near-duplicates collapse.** Two SKUs differing in one character describe near-identical products, so their vectors are near-identical too. Dense retrieval will happily return the sibling variant. For a semantic question that is a fine answer; for someone typing a part number it is the wrong item, and users experience it as the search being broken. Brand names sit in a middle zone: they have semantics (a brand implies a category and a price tier), so a dense search returns plausible competitors — the single most annoying failure mode a catalogue search can have. ## What lexical retrieval contributes A sparse lexical scorer such as BM25 ranks on term overlap weighted by how rare each term is across the corpus. Its biases are the mirror image of dense retrieval's: it cannot match paraphrases at all, but it is exact on strings and it rewards rare terms strongly — so a query containing one very rare token is pulled hard toward the few documents containing that token. That is precisely the regressed segment. This is why hybrid retrieval is the default shape for production search rather than an optimisation: the two legs fail on disjoint query populations, so the union of their candidates is meaningfully better than either alone. You run both retrievals, then merge the two ranked lists into one shortlist. How you merge — rank-based fusion, score-weighted blending, learned weighting — is a real design decision with its own tradeoffs and is a topic in its own right; the point to make in an interview is that you need a principled combiner because the two legs produce scores on incomparable scales. ## The alternative: routing Instead of always running both, you can classify the query and route it. Queries that look like identifiers — mixed alphanumerics, hyphens, high digit ratio, an exact match against a known brand or SKU dictionary — go lexical-first; natural-language phrases go dense-first. Routing is cheaper at serving time and gives crisper behaviour on the two extremes, but it introduces a classifier that can be wrong, and mixed queries ("quiet WD-42X replacement") are exactly where classifiers are worst. In practice most teams run both legs and reserve routing for a small set of high-confidence patterns, sometimes as a short-circuit that returns an exact identifier match directly. ## Diagnosing before fixing The reason this scenario is a senior question is that the interesting work happens before the fix. An aggregate A/B number that reads "neutral" routinely hides a large win on long natural-language queries and a large loss on head identifier queries. The discipline is: - **Segment queries by shape** — natural-language phrase, single common noun, brand name, alphanumeric identifier, misspelling — and read the experiment per segment. - **Watch zero-result and reformulation rates**, not just click metrics. A user who searches, gets plausible wrong products, and immediately retypes the code has registered a failure that a click-through average may not show. - **Check candidate membership**, not just ordering: for a regressed query, was the exact product even retrieved? That determines whether you need a lexical leg or merely better ranking. ## Related non-embedding fixes A lexical leg is the general answer, but for identifiers there are cheaper complements worth naming: an exact-match lookup path that short-circuits when the query matches a SKU or ISBN pattern; normalising identifiers at index and query time (case, separators, whitespace) so `wd 42x` and `WD-42X` collide; and indexing the identifier as its own field so it is not diluted by a long product description. These are unglamorous and they solve a large share of the regression on their own. ## What not to claim Do not present hybrid retrieval as strictly better with no cost. You now maintain two indexes that can drift out of sync, two sets of candidates to fetch and fuse, and a combiner with parameters that need tuning and can be overfitted to your evaluation set. It is the right default for catalogue and enterprise search; it is not free.

  • Why does an aggregate A/B result often hide this regression entirely?
    Because the wins and losses live in different query populations. Long natural-language queries improve a lot under semantic search while short identifier queries get worse, and averaged over all traffic the two can roughly cancel. Reading the experiment per query segment — natural language, brand, alphanumeric code, misspelling — separates them, and the identifier segment is usually low-volume but high-intent, so its loss costs more revenue than its share of traffic suggests.
  • Would fine-tuning the embedding model on your catalogue fix the SKU problem instead?
    Only partially, and expensively. Training can pull product codes into better-separated regions, but near-identical codes still describe near-identical products, so the model is fighting its own objective. You also re-embed the entire corpus on every model version and re-face the problem for every newly added SKU. A lexical leg handles new identifiers on day one with no training, which is why it is the standard answer.
  • Your product manager asks to drop the keyword index once the reranker is good enough. What do you say?
    A reranker only reorders what was retrieved. If dense retrieval never returns the exact SKU, no reranker recovers it — the regression comes back. The keyword index affects candidate generation, which sits upstream of ranking. I would show the per-segment candidate-membership numbers with and without the lexical leg to make that concrete before changing anything.

saying these in an interview costs you the question

  • Claims a better embedding model alone will fix exact identifier matching
  • Treats one aggregate A/B number as the verdict for all query types
  • Adds a reranker to fix a candidate-generation problem
  • Says keyword search is legacy and semantic search supersedes it
  • Blends dense and lexical scores directly as if the scales were comparable

context