In semantic search, should you embed a 3,000-word listing or a generated summary of it?
answer
- you choose the text on both sides
- one vector for 3,000 words is an average
- boilerplate drags long documents together
- what the summary omits becomes unfindable
- same model and version, both sides
basics
~20 sUsually the summary, or another focused unit. A single vector for 3,000 words averages many topics into a blurry point that matches short queries poorly, while a tight summary of what the item actually is aligns far better with how people search.
solid answer
~50 sThe comparison happens between two vectors, so you should choose the text unit on each side to make that comparison fair. A short query like "walkable neighbourhood near good schools" produces a focused vector; a 3,000-word listing that also covers the boiler, the HOA rules and the closing process produces one vector that is an average of all of it, and averages sit near nothing in particular. Embedding a generated summary — or a structured précis of the fields that people actually search on — gives a vector whose content matches the granularity of real queries, and it is cheap because you generate it once at index time. The tradeoff is that anything the summary omits becomes unsearchable, and the summariser is now part of your retrieval quality. Keep the full text for display and for any reranking stage; embed the unit that answers the query.
go deeper
Know that an embedding is one fixed-size vector no matter how long the text is, so a very long document produces a vague vector. Be able to say that shorter, focused units usually match short queries better.
Explain the dilution mechanism and the three remedies — split into passages, embed a derived summary, or expand the query — and name the hard alignment rules: same model, same version and the same preprocessing on both sides.
Show that you would decide it with an offline measurement on real queries rather than by preference, and that you have thought about what a summary makes unfindable, how to keep exact terms searchable anyway, and how a model upgrade forces a full corpus rebuild.
Treat the summariser as production infrastructure: a prompt change silently re-shapes retrieval across the whole corpus. Own the versioning, rebuild strategy and cost model for re-embedding, and the policy for what the indexed representation is contractually required to cover.
## The comparison is between two vectors, not two documents Everything in this question follows from one observation: retrieval compares a *query vector* to a *document vector*. Neither is the text. So the design question is not "how do I index my documents" but "what pair of texts do I want compared", and you control both sides. Real queries are short, focused and written in the language of intent: what someone wants, not how a document describes itself. A real document is long, multi-topic, and written in the language of the domain. Handing both to the same encoder and hoping the geometry works out is the default that produces the classic complaint — results that are topically in the neighbourhood but useless. ## Why a single vector over long text goes blurry An embedding is a fixed-size vector regardless of input length. A 3,000-word property listing that spends 300 words on the neighbourhood, 600 on the interior, 400 on the heating system, 500 on legal terms and 1,200 on agency boilerplate produces one vector that reflects all of those in proportion. The signal a searcher cares about is a small fraction of the input and is diluted accordingly. Two consequences follow: - **Long documents become homogeneous.** Because every listing contains similar boilerplate, their vectors drift toward a common centre and away from each other, so the ordering among them becomes noisy. - **Focused queries under-match.** A query about walkability is compared against a vector that is only fractionally about walkability, so the similarity is low even for a perfect match, and a shorter, sloppier document may beat it. ## The three standard remedies **Split the document into smaller units.** This is the usual first answer: index passages rather than whole documents, so each vector is about one thing. Choosing those boundaries and sizes is a substantial topic of its own; the relevant point here is that splitting changes what the vector is *about*, and that is the whole mechanism. **Embed a derived representation.** Generate a short summary — or a canonical field-based description assembled from structured data — at index time and embed that, keeping the original text stored alongside for display and for the reranking stage. This works especially well for listings, product pages and job postings, where the searchable essence is a handful of facts buried in marketing prose. It costs one generation per document at ingest and nothing at query time. **Move the query toward the documents instead.** Rather than compressing documents, expand the query: generate a hypothetical answer or a richer paraphrase and embed that. This trades index-time cost for query-time latency and adds a generation step in the hot path, which is why the index-side transformation is more common when the corpus is stable. These are not exclusive. A common production shape is: split into sections, prepend a short document-level summary to each section so the passage carries its own context, embed that combined unit. ## What the derived representation costs you Be honest about the downside in an interview. If the summary omits the boiler specification, no query about boilers will ever retrieve that listing — you have made an editorial decision about searchability, and it is invisible until a user complains. Mitigations: keep a lexical index over the full original text so exact terms remain findable; template the summary so it always covers the fields users search on rather than letting a model choose freely; and re-generate when the source changes. You have also introduced a dependency: a change to the summarisation prompt or model silently changes retrieval behaviour across the whole corpus, so it belongs under the same versioning discipline as the embedding model. ## Alignment rules that are non-negotiable Whatever units you choose, the two sides must be produced consistently: - **The same embedding model and version on both sides.** Vectors from different models are not comparable, even when the dimension matches. A model upgrade means re-embedding the entire corpus, so plan for a full rebuild with dual-write or a shadow index rather than a partial migration. - **The same preprocessing.** Case handling, whitespace, stripped markup, truncation rules — a mismatch between the ingest path and the query path is a bug that produces subtly worse results and no error message. - **The model's prescribed prefixes.** Several retrieval embedding models are trained expecting a short instruction prefix that differs for queries and for passages. If the model documents them, use them exactly, on the correct side. Omitting them, or applying the query prefix to documents, measurably degrades results while everything still appears to work. ## How to decide, concretely Take a sample of real queries with known correct answers, build two indexes — raw text and derived summaries — and measure how often the correct item is retrieved by each. This is a cheap experiment because it is offline and needs no ranking model. In practice the summary index usually wins for long marketing-style content and loses for short, information-dense records where there was nothing to compress in the first place.
- What breaks if the ingest pipeline uses a newer version of the embedding model than the query path?Everything appears to work and quality quietly collapses. Vectors from two model versions are not comparable — the same dimension does not mean the same space — so distances become close to meaningless and results look randomly plausible. There is no error to catch it. The discipline is to treat the model identifier as part of the index identity, rebuild the whole corpus on an upgrade, and refuse queries whose model version does not match the index.
- If you embed a generated summary, what do you hand to a reranking stage?Give the reranker the text that best supports a relevance judgement, usually the original passage rather than the summary, and be deliberate about it. The failure to avoid is the two stages judging different objects — retrieving on a summary and reranking on unrelated boilerplate from the same document. Store the mapping from each indexed unit back to the source span so the reranker sees the corresponding original text.
- When is embedding the raw document actually the right choice?When the document is already short and information-dense — a support ticket, a commit message, a product title with attributes — there is nothing to compress and a summary only adds a lossy step plus a model dependency. It is also right when queries are themselves long and document-shaped, because then both sides already sit at similar granularity.
saying these in an interview costs you the question
- Assumes longer input always makes a richer, better embedding
- Embeds documents with one model and queries with another
- Mixes vectors from two model versions in the same index
- Forgets a model's required query and passage prefixes
- Never checks which content the summary silently dropped