skip to content

Embeddings

Turning text into vectors and doing useful things with them: which model produces the vector, which metric compares them, how many dimensions you need, how an ANN index searches them, and what search or clustering you build on top. Interviewers reach for this area whenever semantic search or RAG comes up.

on this pageshow

explore

questions

page 1 of 2

What is the difference between a text embedding model and a generative LLM?

level: juniorimportance: must knowfreq 70%

answer

  1. one vector out, not text out
  2. fixed length regardless of input
  3. comparison versus generation
  4. one forward pass, cacheable, cheap
  5. spaces are model-specific, not portable

basics

~20 s

An embedding model maps a piece of text to one fixed-length vector of numbers that can be compared with other vectors. A generative LLM produces new text token by token. Embeddings are for comparing meaning, not for writing.

solid answer

~50 s

An embedding model is a **representation** model: you give it a text and it returns a fixed-length vector — the same size for a three-word query and a three-paragraph passage — whose geometry encodes meaning, so two texts about the same thing land close together. A generative LLM is a **prediction** model: it samples the next token repeatedly to produce new text. That difference decides the use case. Search, deduplication, clustering, routing and recommendation all need a stable numeric handle on meaning that you can index once and reuse; answering, summarizing and rewriting need generation. Embedding models are also far cheaper: one forward pass, no decoding loop, small enough to run locally, and the vectors are cacheable because the same input gives the same output. Note that the two families have converged architecturally — several strong open embedding models are decoder-only LLM backbones adapted with contrastive training — but the *interface* distinction, one vector versus a token stream, is what matters in practice.

go deeper

for a junior

Be able to say plainly that an embedding model returns one fixed-length vector of numbers while a generative model returns text, and name one task for each: similarity search versus drafting a reply.

for a middle

Explain why comparison at scale needs a precomputed representation: embed the corpus once, embed only the query at request time, compare vectors. Mention that vectors are model-specific and not portable.

for a senior

Own the migration consequence — changing embedding model invalidates every stored vector, so model identity and version belong in the index metadata and a re-embedding plan belongs in the design.

for a principal

Frame it as where you spend inference budget: cheap representation over the whole corpus, expensive generation over a shortlist. Be ready to argue when a task genuinely needs generation rather than a retrieval or classification primitive.

## Two different jobs Both families are transformer neural networks trained on large text corpora, and both "understand" language in some loose sense. What separates them is what comes out the other end. An **embedding model** consumes a text and emits a single fixed-length array of floating-point numbers — commonly a few hundred to a few thousand of them. The length does not depend on the input length: a two-word query and a 400-word passage both come back as, say, 768 numbers. Those numbers are coordinates in a vector space, and the space is trained so that texts with similar meaning sit near each other. That is the entire product: a comparable numeric handle on meaning. A **generative LLM** consumes a text and emits a probability distribution over the next token, samples one, appends it, and repeats until it decides to stop. The product is new text. ## Why the distinction decides the tool You reach for embeddings whenever the task is fundamentally *comparison over a set*: - finding which of a million stored documents is closest in meaning to a query; - detecting that two bug reports filed by different teams describe the same defect; - grouping incoming support tickets into themes nobody labelled in advance; - routing a message to the right queue by similarity to past examples; - flagging near-duplicate listings in a marketplace. Every one of those needs a *stable, precomputable* representation. You embed the corpus once, store the vectors, and at query time you embed only the query and compare. That is a single cheap forward pass against a prebuilt index. You reach for a generative model when the output is language a human will read: drafting the reply, summarizing the incident, rewriting the clause, turning extracted fields into prose. A generative model cannot give you the comparison primitive directly — asking an LLM "are these two paragraphs about the same thing?" works, but it costs a full inference call per *pair*, which is quadratic in the corpus and unusable at scale. The standard architecture is therefore embeddings to narrow a million candidates down to a handful, then a generative model to do something intelligent with those few. ## Cost and operational profile Embedding models are typically far smaller than frontier generative models, run in one forward pass with no autoregressive decoding loop, and are cheap enough to run on commodity hardware or even on-device. Their output is deterministic for a fixed model and input, which means you can cache it indefinitely and treat the vector as derived data. That determinism has a sharp operational consequence: **vectors from different models are not comparable**. Each model learns its own coordinate system during training, so a vector produced by model A and a vector produced by model B are in unrelated spaces even if they have the same number of dimensions. Changing your embedding model is therefore not a config flip — it invalidates every stored vector and requires re-embedding the whole corpus. Plan for that migration cost before you pick a model, and keep the model identity and version stored alongside the vectors. ## The architectural convergence Historically the split mapped cleanly onto architecture: embedding models were encoder-only transformers (the BERT family, and the sentence-transformer models built on them) that read the whole input bidirectionally, while generative models were decoder-only transformers that read left to right. That mapping no longer holds strictly. Since roughly 2023 a strong line of embedding models has been built by taking a decoder-only LLM backbone and adapting it with contrastive training on paired texts; several of the highest-scoring open embedding models today are of that kind. They still emit one vector, they are still used exactly the same way, and they are usually heavier and slower than a classic encoder — which is why compact encoders remain the workhorse for large corpora. So do not define the difference by architecture in an interview. Define it by interface and purpose: one fixed-length vector for comparison, versus a token stream for reading. ## What weak answers get wrong The common mistake is treating the embedding model as a weaker LLM, or assuming any LLM can "give you the embedding" interchangeably. A second mistake is assuming vectors are portable between models because the dimension count matches. A third is reaching for a generative model to score similarity in a hot path — correct in principle, ruinous in latency and cost the moment the candidate set is bigger than a handful.

  • If both models output vectors of 1024 numbers, why can't I compare one model's vector with another's?
    Because each model learns its own coordinate system during training. Dimension 412 means something entirely different in each space, so a distance between them is arithmetic without meaning. Matching dimensionality is a coincidence of configuration, not compatibility. The practical consequence is that swapping embedding models forces a full re-embedding of the corpus, so store the model name and version next to every vector.
  • Why not just ask a generative model whether two texts mean the same thing?
    For a single pair, that works and is often more accurate than cosine similarity. It does not scale: comparing a query against a million documents means a million inference calls. Embeddings make the comparison a cheap vector operation over a precomputed index. The usual design uses embeddings to shortlist a handful of candidates and a stronger model only on that shortlist.
  • Are embedding outputs deterministic?
    For a fixed model version and input, yes in practice — one forward pass with no sampling, so the same text yields the same vector, which is what makes caching safe. Minor numeric drift can appear across hardware, precision settings or serving backends, so treat exact bit equality as unreliable while treating the vector as stable enough to store and reuse.

saying these in an interview costs you the question

  • Calling an embedding model a smaller or dumber LLM
  • Assuming vectors from different models are comparable if dimensions match
  • Thinking embedding output length grows with input length
  • Using a generative model to score similarity across a whole corpus
  • Believing embedding models must be encoder-only architectures

context

open as a page

Which parts of a search request should never be answered by vector similarity?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Anything with a truth condition: numeric and date constraints, sorting, counting, permission scoping, and exact identifier lookups. Vector similarity ranks by how alike two texts are, and "alike" cannot express "price under 800,000" or "only records this user may see".

open as a page

In vector search, what does cosine similarity measure, and when is it preferred over Euclidean distance?

level: juniorimportance: must knowfreq 80%

basics

~20 s

Cosine similarity measures the angle between two vectors and ignores their lengths, ranging from -1 to 1. Prefer it over Euclidean distance when only direction is meaningful and vector length reflects something you do not want scored, such as text length.

open as a page

What can a UMAP or t-SNE plot of embeddings actually tell you?

level: middleimportance: must knowfreq 66%

basics

~20 s

UMAP and t-SNE preserve which points sit near each other locally, so tight visible groups usually reflect real neighbourhoods. The gaps between blobs, the relative blob sizes and the axes carry no reliable meaning and shift with hyperparameters and seed.

open as a page

How do you choose between k-means and DBSCAN for clustering document embeddings?

level: middleimportance: must knowfreq 62%

basics

~20 s

k-means forces every document into one of k clusters you fix in advance, so it fits a corpus you want fully partitioned. DBSCAN groups by density, discovers how many clusters exist, and leaves sparse points unlabelled as noise.

open as a page

What do you lose when you truncate a Matryoshka embedding from 3072 to 256 dimensions?

level: middleimportance: must knowfreq 58%

basics

~20 s

Fine-grained discrimination. Matryoshka training packs coarse meaning into the leading dimensions, so a 256-value prefix still separates obviously different texts but blurs near-neighbours, costing recall on subtle and long-tail queries. Re-normalise the truncated vector before using it.

open as a page

Why do unrelated texts still score 0.8 cosine in an embedding space?

level: middleimportance: must knowfreq 55%

basics

~20 s

Learned embedding spaces are anisotropic: vectors crowd into a narrow cone instead of spreading over the sphere, so almost every pair shares a large common direction. That inflates all cosine scores, making 0.8 a baseline rather than a match.

open as a page

How does a sentence embedding model pool BERT token vectors into one vector?

level: middleimportance: must knowfreq 65%

basics

~20 s

A transformer encoder emits one vector per token, so a pooling step collapses them into one: take the [CLS] token's vector, or average all real token vectors (mean pooling). Which one works depends on how the model was trained.

open as a page

Why does a semantic search pipeline retrieve with a bi-encoder and rerank with a cross-encoder?

level: middleimportance: must knowfreq 70%

basics

~20 s

A bi-encoder embeds documents independently, so their vectors are built offline and searched in milliseconds. A cross-encoder reads query and document together: much more accurate, but impossible to precompute, so it only rescores the shortlist the first stage returns.

open as a page

Why do cosine, dot product and L2 rank identically once vectors are unit-normalized?

level: middleimportance: must knowfreq 62%

basics

~20 s

Unit-normalized vectors have length 1, so the dot product equals the cosine, and squared Euclidean distance equals 2 minus twice that cosine. All three are monotone functions of one quantity, so the ordering is identical even though the numbers differ.

open as a page

Which HNSW parameters are fixed at build time and which can you tune per query?

level: middleimportance: must knowfreq 64%

basics

~20 s

M and efConstruction are baked into the graph when vectors are inserted, so changing them means rebuilding. efSearch is a query-time knob: raise it per request to buy recall with latency, lower it for cheap traffic.

open as a page

How does an HNSW index search its layers to find nearest neighbours?

level: middleimportance: must knowfreq 72%

basics

~20 s

HNSW stacks proximity graphs. Search starts at one entry point in the sparse top layer and greedily hops to closer neighbours, dropping a layer at each local minimum, then runs a widened beam search over the dense bottom layer.

open as a page

How does an IVF index narrow a vector search, and what does nprobe trade off?

level: middleimportance: must knowfreq 72%

basics

~20 s

An IVF index clusters vectors into cells around centroids and files each vector under its nearest centroid. A query compares itself only against the nprobe nearest cells, so raising nprobe raises recall and latency together.

open as a page

What does product quantization compress in a vector index, and at what cost?

level: middleimportance: must knowfreq 62%

basics

~20 s

Product quantization splits each vector into subvectors and replaces every subvector with the id of the nearest centroid in a small per-subspace codebook. Storage drops to a few bytes per vector, but every distance computed from codes is an estimate.

open as a page

Why can an embedding model silently ignore half of a 900-token chunk?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Every encoder has a maximum sequence length, and the default behaviour when input exceeds it is to truncate rather than fail. The returned vector represents only the leading tokens; the tail never influences it, and nothing in the response says so.

open as a page

Semantic-only search regressed on brand and SKU queries in an A/B — what do you do?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Add a lexical retrieval leg. Dense embeddings compress meaning and blur rare exact tokens like part numbers and brand names, so keep a keyword index (typically BM25) running alongside the vector search and merge both candidate lists before ranking.

open as a page

How do you work out what a cluster of document embeddings actually represents?

level: juniorimportance: should knowfreq 38%

basics

~20 s

Read the documents nearest the cluster's centre, extract the terms that are common inside the cluster but rare outside it, and have a model draft a label from a random sample of members. Then check that label against members you did not look at.

open as a page

How much memory do 50 million 1536-dimension float32 embeddings require?

level: juniorimportance: should knowfreq 48%

basics

~10 s

Roughly 300 GB: 50,000,000 x 1536 x 4 bytes is about 307 GB of raw vector data, before any index structure, metadata, or replicas. Dimension count multiplies every byte you store, ship, and compare.

open as a page

Does the curse of dimensionality break nearest-neighbour search over 1536-dimension embeddings?

level: middleimportance: should knowfreq 42%

basics

~20 s

In practice, no. Distance concentration is proven for independent random coordinates, but learned embeddings lie on a much lower-dimensional manifold with correlated coordinates, so meaningful contrast survives. The real cost of high dimension is memory and index efficiency, not broken similarity.

open as a page

Why do embedding models like E5 and BGE need different prefixes for queries and passages?

level: middleimportance: should knowfreq 48%

basics

~20 s

They were trained asymmetrically: short questions and long documents are mapped into the shared space by different roles, and the prefix tells the model which side it is encoding. Encode both sides identically and retrieval quality degrades silently.

open as a page

Why does an English-trained embedding model degrade on Portuguese and Japanese text?

level: middleimportance: should knowfreq 48%

basics

~20 s

Its training data and tokenizer vocabulary were built for English, so other languages fragment into unfamiliar subword pieces and land in a poorly organized region of the space. Vectors still come back looking normal, so the degradation is invisible without a per-language evaluation.

open as a page

In semantic search, should you embed a 3,000-word listing or a generated summary of it?

level: middleimportance: should knowfreq 50%

basics

~20 s

Usually the summary, or another focused unit. A single vector for 3,000 words averages many topics into a blurry point that matches short queries poorly, while a tight summary of what the item actually is aligns far better with how people search.

open as a page

When one vector store returns a distance and another a similarity score, what bug follows?

level: middleimportance: should knowfreq 38%

basics

~20 s

Distance is better when lower and similarity is better when higher. Code that sorts both the same way returns the worst matches from one of them, and any minimum-score threshold inverts into a keep-only-the-worst filter. Nothing crashes; results merely become quietly wrong.

open as a page

Silhouette scores stay flat from k=2 to k=30 on document embeddings — what does that mean?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A flat curve means no value of k gives a materially better partition than any other: the corpus behaves more like a continuum than a set of separated groups. Do not take the argmax of noise — validate against a held-out labelling or a downstream task instead.

open as a page

How would you validate cutting a legacy 4096-dimension encoder down to 128 dimensions with PCA?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Judge it by retrieval quality, not by explained variance. Fit PCA on a corpus sample, project queries with the same transform, then sweep widths and compare recall@10 and MRR on labeled queries against the full-dimension ranking. Ship the smallest width before the cliff.

open as a page

What breaks when only half a vector index's embeddings are L2-normalized?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Dot-product ranking scales with vector length, so unnormalized entries score by magnitude as well as direction. A half-normalized collection splits into two score scales: one group crowds every result list while the other becomes effectively unreachable.

open as a page

In a dot-product recommender, popular titles dominate results — why, and is that a bug?

level: seniorimportance: should knowfreq 47%

basics

~20 s

Dot-product scores grow with vector length, and items with abundant training interactions often end up with longer vectors, so popular titles outrank better-matching niche ones. Whether that is a defect depends on whether popularity is a signal you deliberately want in the score.

open as a page

How do deletes degrade an HNSW index over time, and what fixes it?

level: seniorimportance: should knowfreq 46%

basics

~20 s

Most implementations tombstone deletes: the vector stays in the graph as a connector and is only filtered out of results. Queries then spend their candidate budget on dead nodes, so latency rises and effective recall falls until compaction or a rebuild.

open as a page

Why does an HNSW index need more RAM than the vectors alone, and how much more?

level: seniorimportance: should knowfreq 52%

basics

~20 s

HNSW stores a neighbour-id list per vector per layer — about 2M ids at layer 0. At M=32 that adds roughly 250 bytes per vector, so an 80M-vector catalogue pays around 20 GB for graph edges alone, on top of the vectors.

open as a page

Why do quantized vector indexes rescore a shortlist against full-precision vectors?

level: seniorimportance: should knowfreq 48%

basics

~20 s

Compressed codes give noisy distances, so their ordering is unreliable near the top. Retrieving an oversized shortlist with the cheap distance and re-ranking it against the original full-precision vectors recovers most of the lost recall for a small extra cost.

open as a page

showing 1–30 of 37