skip to content

How do you decide whether semantic search should replace an existing keyword search?

level: principalimportance: should knowfreq 42%

answer

  1. the incumbent carries years of hidden tuning
  2. measure the baseline before proposing anything
  3. segment the queries or ship a regression
  4. interleaving decides faster than an A/B
  5. augmenting usually beats replacing

basics

~20 s

Compare against the incumbent per query segment, not in aggregate. Semantic search wins on long natural-language queries and can lose on head, identifier and navigational ones, so the decision is which segments improve, which regress, and whether the regressions are acceptable.

solid answer

~50 s

Treat it as a migration decision with a named incumbent, not a technology adoption. First establish the keyword baseline's real numbers, because most teams have never measured it. Then evaluate offline on a query set sampled from actual traffic and **segmented** — natural-language phrases, head navigational queries, identifiers, misspellings, and the long tail — since a single average routinely hides a large win in one segment cancelling a large loss in another. Ship behind an online experiment or interleaving, with guardrails on the segments most likely to regress and on zero-result and query-reformulation rates. Weigh the operating cost honestly: an embedding pipeline, a vector index, a full re-embed on every model change, and a second failure domain. The usual honest outcome in 2026 is not replacement but addition — semantic retrieval as a second leg beside the keyword index, which keeps every existing win while adding the queries keywords could never serve.

go deeper

for a junior

Know that a new search system is judged by comparing it against the one already running, on real user queries, and that it can be better for some kinds of query and worse for others.

for a middle

Be able to describe the evaluation shape: a judged query set sampled from real traffic, results broken down by query type, and an online experiment to confirm offline gains rather than assuming they transfer.

for a senior

Demonstrate launch judgement — guardrail metrics agreed in advance, staged rollout with rollback, monitoring the segments most likely to regress, and refusing a launch whose aggregate looks fine but whose high-intent segment got worse.

for a principal

Own the decision economics. Weigh the recurring cost of an embedding pipeline, index rebuilds on model upgrades and a second failure domain against the tail queries currently unserved, and be willing to conclude that augmenting the incumbent beats replacing it.

## Frame it as a migration, not an upgrade The question contains a trap: "replace". A keyword search that has been in production for years has absorbed a great deal of accumulated tuning — synonym lists, boosts, curated results for head queries, spelling correction, business rules. That tuning is invisible and undocumented and it is why the incumbent performs better than anyone expects on the queries that matter most commercially. Replacing it wholesale discards all of it at once. The disciplined framing is: what does the new capability add, what does it cost, and what does it break. ## Establish the baseline before anything else Most teams proposing semantic search cannot state how the current system performs. So step one is instrumentation, not modelling: - Sample real queries with real frequencies from logs. Traffic is heavily head-weighted, and evaluating on a hand-written query list systematically over-represents the tail where semantic search looks best. - Build a judged set with known-good answers, either from human labels or from strong implicit signals such as the result users ultimately converged on. - Record the incumbent's offline quality on that set, plus its production behaviour: zero-result rate, reformulation rate, latency percentiles, and success on the commercially important query classes. Without this, any comparison later is a story rather than a measurement. ## Segment or you will ship a regression This is the central discipline. Query populations have wildly different characteristics and semantic search affects them in opposite directions: - **Navigational and head queries** — a brand, a category, a known product. The incumbent is heavily tuned here and semantic retrieval frequently makes them worse by returning plausible alternatives. - **Identifier queries** — SKUs, reference numbers, error codes. Reliably weaker under dense retrieval. - **Long natural-language queries** — descriptions of intent that contain none of the corpus's vocabulary. This is where semantic search wins decisively, often turning a zero-result page into a good one. - **Misspellings and morphological variants** — mixed, and depend on whether the incumbent already has spelling correction. - **The long tail** — high aggregate volume, individually rare, historically poorly served, and usually the strongest business case for the change. Read the experiment per segment and require each one to be evaluated on its own. A flat overall number with a large loss in the highest-intent segment is a worse outcome than a modest overall gain. ## Offline first, online decides Offline evaluation on a judged set is cheap and fast and is the right place to iterate. It is not the decision, because it cannot capture presentation, latency, user behaviour or the interaction with the rest of the page. Move to an online comparison for the decision: either a conventional A/B, or interleaving, where results from both systems are mixed into one list and user preference between them is measured directly — interleaving needs far less traffic to reach a conclusion, which matters when a search team has weeks rather than quarters. Set guardrail metrics before you start and agree what would stop the launch: a regression threshold on identifier and navigational segments, latency percentiles, zero-result rate, and reformulation or abandonment rate. Deciding those afterwards is how teams talk themselves into shipping a regression. ## Count the operating cost honestly The recurring costs are the part a principal is expected to raise unprompted: - **An ingestion path that must not drift.** Every document has to be embedded, re-embedded on change, and reconciled if the pipeline fails midway. Silent divergence between the source of truth and the index is the most common production defect in vector search. - **Rebuilds on model change.** Upgrading the embedding model invalidates the entire index. That means capacity for a full re-embed, a dual-index or shadow strategy, and a cutover plan — for a large corpus this is a project, not a deploy. - **A second failure domain.** A new store, new capacity planning, new latency characteristics, and a new dependency in the request path. - **Serving cost.** Embedding every query, plus vector search, plus any reranking tier. Against that, quantify the upside in the same currency: the tail and natural-language queries currently returning nothing, converted at your normal rate. ## The usual honest conclusion As of 2026 the mature answer is almost always **augment rather than replace**: keep the keyword index and add semantic retrieval as a second leg, so every existing tuned win is preserved and the new capability covers what keywords structurally cannot. That also de-risks the rollout, because the incumbent remains available as a fallback and as a comparison in production. Full replacement is defensible mainly for a genuinely new surface with no incumbent tuning to lose, or a corpus and query mix where lexical matching was never the right model in the first place. Stage the rollout regardless: internal traffic, then a small percentage, then per-segment expansion, with a fast rollback path and per-segment dashboards left in place afterwards — because model and corpus drift will move these numbers long after launch.

  • Why is interleaving often preferred over a conventional A/B test for comparing two rankers?
    Because it removes between-user variance. Each user sees a single list built from both systems' results, so their clicks express a direct preference between the two rankers rather than being compared across separate populations. That sensitivity means a conclusion in a fraction of the traffic and time an A/B needs. The tradeoff is that it measures ranking preference specifically, so end-to-end business outcomes still need a conventional experiment.
  • Leadership wants one number to decide the launch. What do you give them?
    A primary metric plus explicit guardrails, never a single number alone. I would name one headline outcome measure and pair it with hard stop conditions: no regression beyond an agreed threshold on identifier and navigational queries, latency percentiles within budget, and no rise in zero-result or reformulation rates. One aggregate number is exactly what lets a serious segment regression pass unnoticed.
  • How do you plan for the day the embedding model is upgraded?
    Treat it as a full index rebuild from the start. Version the index by model identifier, build the new index in parallel from the source of truth, validate it against the same judged query set as the original launch, then cut over and keep the old index until confidence is established. Anything incremental mixes incomparable vector spaces. Budget the re-embed cost and duration when you first size the system, not when the upgrade arrives.

saying these in an interview costs you the question

  • Compares only aggregate quality across all queries
  • Never measured the existing keyword system's performance
  • Ignores the cost of re-embedding on every model upgrade
  • Assumes semantic search strictly dominates keyword search
  • Decides on offline metrics alone with no online experiment

context