What is the difference between a text embedding model and a generative LLM?
answer
- one vector out, not text out
- fixed length regardless of input
- comparison versus generation
- one forward pass, cacheable, cheap
- spaces are model-specific, not portable
basics
~20 sAn embedding model maps a piece of text to one fixed-length vector of numbers that can be compared with other vectors. A generative LLM produces new text token by token. Embeddings are for comparing meaning, not for writing.
solid answer
~50 sAn embedding model is a **representation** model: you give it a text and it returns a fixed-length vector — the same size for a three-word query and a three-paragraph passage — whose geometry encodes meaning, so two texts about the same thing land close together. A generative LLM is a **prediction** model: it samples the next token repeatedly to produce new text. That difference decides the use case. Search, deduplication, clustering, routing and recommendation all need a stable numeric handle on meaning that you can index once and reuse; answering, summarizing and rewriting need generation. Embedding models are also far cheaper: one forward pass, no decoding loop, small enough to run locally, and the vectors are cacheable because the same input gives the same output. Note that the two families have converged architecturally — several strong open embedding models are decoder-only LLM backbones adapted with contrastive training — but the *interface* distinction, one vector versus a token stream, is what matters in practice.
go deeper
Be able to say plainly that an embedding model returns one fixed-length vector of numbers while a generative model returns text, and name one task for each: similarity search versus drafting a reply.
Explain why comparison at scale needs a precomputed representation: embed the corpus once, embed only the query at request time, compare vectors. Mention that vectors are model-specific and not portable.
Own the migration consequence — changing embedding model invalidates every stored vector, so model identity and version belong in the index metadata and a re-embedding plan belongs in the design.
Frame it as where you spend inference budget: cheap representation over the whole corpus, expensive generation over a shortlist. Be ready to argue when a task genuinely needs generation rather than a retrieval or classification primitive.
## Two different jobs Both families are transformer neural networks trained on large text corpora, and both "understand" language in some loose sense. What separates them is what comes out the other end. An **embedding model** consumes a text and emits a single fixed-length array of floating-point numbers — commonly a few hundred to a few thousand of them. The length does not depend on the input length: a two-word query and a 400-word passage both come back as, say, 768 numbers. Those numbers are coordinates in a vector space, and the space is trained so that texts with similar meaning sit near each other. That is the entire product: a comparable numeric handle on meaning. A **generative LLM** consumes a text and emits a probability distribution over the next token, samples one, appends it, and repeats until it decides to stop. The product is new text. ## Why the distinction decides the tool You reach for embeddings whenever the task is fundamentally *comparison over a set*: - finding which of a million stored documents is closest in meaning to a query; - detecting that two bug reports filed by different teams describe the same defect; - grouping incoming support tickets into themes nobody labelled in advance; - routing a message to the right queue by similarity to past examples; - flagging near-duplicate listings in a marketplace. Every one of those needs a *stable, precomputable* representation. You embed the corpus once, store the vectors, and at query time you embed only the query and compare. That is a single cheap forward pass against a prebuilt index. You reach for a generative model when the output is language a human will read: drafting the reply, summarizing the incident, rewriting the clause, turning extracted fields into prose. A generative model cannot give you the comparison primitive directly — asking an LLM "are these two paragraphs about the same thing?" works, but it costs a full inference call per *pair*, which is quadratic in the corpus and unusable at scale. The standard architecture is therefore embeddings to narrow a million candidates down to a handful, then a generative model to do something intelligent with those few. ## Cost and operational profile Embedding models are typically far smaller than frontier generative models, run in one forward pass with no autoregressive decoding loop, and are cheap enough to run on commodity hardware or even on-device. Their output is deterministic for a fixed model and input, which means you can cache it indefinitely and treat the vector as derived data. That determinism has a sharp operational consequence: **vectors from different models are not comparable**. Each model learns its own coordinate system during training, so a vector produced by model A and a vector produced by model B are in unrelated spaces even if they have the same number of dimensions. Changing your embedding model is therefore not a config flip — it invalidates every stored vector and requires re-embedding the whole corpus. Plan for that migration cost before you pick a model, and keep the model identity and version stored alongside the vectors. ## The architectural convergence Historically the split mapped cleanly onto architecture: embedding models were encoder-only transformers (the BERT family, and the sentence-transformer models built on them) that read the whole input bidirectionally, while generative models were decoder-only transformers that read left to right. That mapping no longer holds strictly. Since roughly 2023 a strong line of embedding models has been built by taking a decoder-only LLM backbone and adapting it with contrastive training on paired texts; several of the highest-scoring open embedding models today are of that kind. They still emit one vector, they are still used exactly the same way, and they are usually heavier and slower than a classic encoder — which is why compact encoders remain the workhorse for large corpora. So do not define the difference by architecture in an interview. Define it by interface and purpose: one fixed-length vector for comparison, versus a token stream for reading. ## What weak answers get wrong The common mistake is treating the embedding model as a weaker LLM, or assuming any LLM can "give you the embedding" interchangeably. A second mistake is assuming vectors are portable between models because the dimension count matches. A third is reaching for a generative model to score similarity in a hot path — correct in principle, ruinous in latency and cost the moment the candidate set is bigger than a handful.
- If both models output vectors of 1024 numbers, why can't I compare one model's vector with another's?Because each model learns its own coordinate system during training. Dimension 412 means something entirely different in each space, so a distance between them is arithmetic without meaning. Matching dimensionality is a coincidence of configuration, not compatibility. The practical consequence is that swapping embedding models forces a full re-embedding of the corpus, so store the model name and version next to every vector.
- Why not just ask a generative model whether two texts mean the same thing?For a single pair, that works and is often more accurate than cosine similarity. It does not scale: comparing a query against a million documents means a million inference calls. Embeddings make the comparison a cheap vector operation over a precomputed index. The usual design uses embeddings to shortlist a handful of candidates and a stronger model only on that shortlist.
- Are embedding outputs deterministic?For a fixed model version and input, yes in practice — one forward pass with no sampling, so the same text yields the same vector, which is what makes caching safe. Minor numeric drift can appear across hardware, precision settings or serving backends, so treat exact bit equality as unreliable while treating the vector as stable enough to store and reuse.
saying these in an interview costs you the question
- Calling an embedding model a smaller or dumber LLM
- Assuming vectors from different models are comparable if dimensions match
- Thinking embedding output length grows with input length
- Using a generative model to score similarity across a whole corpus
- Believing embedding models must be encoder-only architectures