Embeddings
Turning text into vectors and doing useful things with them: which model produces the vector, which metric compares them, how many dimensions you need, how an ANN index searches them, and what search or clustering you build on top. Interviewers reach for this area whenever semantic search or RAG comes up.
on this pageshowhide
explore
- Embedding Models5 questions
- Similarity Metrics4 questions
- Dimensionality4 questions
- Vector Indexes (ANN)10 questions
- Graph-Based Indexes (HNSW)5 questions
- Partitioned and Compressed Indexes (IVF, PQ)5 questions
- Semantic Search5 questions
- Clustering & Visualization4 questions
- Embedding Space Geometry5 questions
questions
page 2 of 2Why must IVF and PQ indexes be trained, and what breaks when the data drifts?
basics
~20 sIVF centroids and PQ codebooks are fitted from a sample of the corpus before ingestion, then frozen. When later data drifts away from that sample, cells become unbalanced and codes reconstruct poorly, so recall degrades quietly while queries keep succeeding.
How do you choose an embedding similarity threshold for auto-closing incidents?
basics
~20 sDerive it from labelled pairs on that exact model and corpus, choose the operating point from the cost of a wrong auto-close, and prefer a margin over the runner-up to a bare number. Re-validate on every model or corpus change.
How do you decide whether semantic search should replace an existing keyword search?
basics
~20 sCompare against the incumbent per query segment, not in aggregate. Semantic search wins on long natural-language queries and can lose on head, identifier and navigational ones, so the decision is which segments improve, which regress, and whether the regressions are acceptable.
Your HNSW build takes 11 hours on one machine — how do you plan rebuilds without downtime?
basics
~20 sTreat the build as a batch job, not an online operation. Shard so builds run in parallel, keep serving the old index while the new one builds on separate capacity, then swap atomically behind the query path. Reserve headroom for two copies.
When is exact flat vector search still the right choice over IVF or PQ?
basics
~20 sWhenever the corpus fits in memory and exhaustive scanning meets the latency budget. Exact search gives perfect recall with no training, no tuning, no rebuild and trivial updates — approximation is a cost you should only pay once the arithmetic forces it.
How do shared text-image embeddings let a description retrieve untagged photos?
basics
~20 sA text encoder and an image encoder are trained together on image-caption pairs so both output into one shared vector space. A description and a matching photo land near each other, so ordinary similarity search over image vectors answers a text query.
In vector search, why does one chunk appear in the top-10 for most queries?
basics
~10 sThat chunk is a hub. In high-dimensional spaces a few vectors sit near the centre of the distribution and land in many nearest-neighbour lists regardless of the query. Generic boilerplate is the textbook hub.
showing 31–37 of 37