skip to content

Embeddings

Turning text into vectors and doing useful things with them: which model produces the vector, which metric compares them, how many dimensions you need, how an ANN index searches them, and what search or clustering you build on top. Interviewers reach for this area whenever semantic search or RAG comes up.

on this pageshow

explore

questions

page 2 of 2

Why must IVF and PQ indexes be trained, and what breaks when the data drifts?

level: seniorimportance: should knowfreq 40%

basics

~20 s

IVF centroids and PQ codebooks are fitted from a sample of the corpus before ingestion, then frozen. When later data drifts away from that sample, cells become unbalanced and codes reconstruct poorly, so recall degrades quietly while queries keep succeeding.

open as a page

How do you choose an embedding similarity threshold for auto-closing incidents?

level: principalimportance: should knowfreq 42%

basics

~20 s

Derive it from labelled pairs on that exact model and corpus, choose the operating point from the cost of a wrong auto-close, and prefer a margin over the runner-up to a bare number. Re-validate on every model or corpus change.

open as a page

How do you decide whether semantic search should replace an existing keyword search?

level: principalimportance: should knowfreq 42%

basics

~20 s

Compare against the incumbent per query segment, not in aggregate. Semantic search wins on long natural-language queries and can lose on head, identifier and navigational ones, so the decision is which segments improve, which regress, and whether the regressions are acceptable.

open as a page

Your HNSW build takes 11 hours on one machine — how do you plan rebuilds without downtime?

level: principalimportance: should knowfreq 36%

basics

~20 s

Treat the build as a batch job, not an online operation. Shard so builds run in parallel, keep serving the old index while the new one builds on separate capacity, then swap atomically behind the query path. Reserve headroom for two copies.

open as a page

When is exact flat vector search still the right choice over IVF or PQ?

level: principalimportance: should knowfreq 44%

basics

~20 s

Whenever the corpus fits in memory and exhaustive scanning meets the latency budget. Exact search gives perfect recall with no training, no tuning, no rebuild and trivial updates — approximation is a cost you should only pay once the arithmetic forces it.

open as a page

How do shared text-image embeddings let a description retrieve untagged photos?

level: middleimportance: nice to knowfreq 35%

basics

~20 s

A text encoder and an image encoder are trained together on image-caption pairs so both output into one shared vector space. A description and a matching photo land near each other, so ordinary similarity search over image vectors answers a text query.

open as a page

In vector search, why does one chunk appear in the top-10 for most queries?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

That chunk is a hub. In high-dimensional spaces a few vectors sit near the centre of the distribution and land in many nearest-neighbour lists regardless of the query. Generic boilerplate is the textbook hub.

open as a page

showing 31–37 of 37