skip to content

In a RAG index, how do you embed five-word queries against 800-token chunks?

level: middleimportance: should knowfreq 55%

answer

  1. short question, long answer, unequal sides
  2. the model was trained for one regime
  3. some models expect a literal marker per side
  4. unit length makes dot product cosine
  5. index and query paths must agree exactly

basics

~20 s

Use a model trained for asymmetric search, apply its expected query and passage prefixes on the correct side, L2-normalize both sides so dot product equals cosine, and keep every convention identical at index time and query time.

solid answer

~50 s

Short-query-to-long-passage retrieval is an **asymmetric** task, and models are trained for one regime or the other. A symmetric model tuned on sentence-pair similarity will happily rank an 800-token policy chunk as dissimilar to "how do I reset MFA" simply because their lengths and registers differ. Asymmetric retrieval models are trained on exactly this shape, and many of them expect a short instruction or prefix on each side — for example the E5 family's `query:` and `passage:` prefixes — applied to the right side and never swapped. Two other index-time decisions follow from the model: whether to L2-normalize (if you do, dot product equals cosine, and the index metric must match), and never to exceed the model's maximum sequence length, since anything beyond it is silently truncated. Every one of these conventions must be identical when indexing and when querying; a mismatch does not error, it just quietly degrades ranking.

code

python · 8 lines
python
import numpy as np

def l2_normalize(v):
    return v / np.linalg.norm(v, axis=-1, keepdims=True)

q = l2_normalize(np.array([[3.0, 4.0]]))
d = l2_normalize(np.array([[6.0, 8.0], [-4.0, 3.0]]))
print(q @ d.T)  # [[1.0, 0.0]] - inner product now equals cosine

go deeper

for a junior

Know that queries and documents must go through the same embedding model with the same settings, and that text longer than the model's limit gets cut off rather than rejected.

for a middle

Explain asymmetric versus symmetric training regimes, name the prefix convention as a real requirement of some model families, and state why L2 normalization must match the index's distance metric.

for a senior

Treat the encoder's conventions as a deployable contract: pin them in index metadata, assert them on the query path, and show how you would detect a silent mismatch before users do.

for a principal

Own the decision of where enrichment happens — query expansion, contextual chunk prefixes, or a different encoder entirely — and weigh each against ingestion cost, re-index risk and the measurable recall it buys on your own query set.

## Symmetric versus asymmetric search Embedding models are trained on a task, and the task shows up in how they behave at retrieval time. Two regimes matter: - **Symmetric similarity:** both sides are the same kind of text of roughly the same length — duplicate-question detection, paraphrase matching, clustering. Training pairs look like sentence-to-sentence. - **Asymmetric retrieval:** a short query is matched against a much longer passage that *answers* it rather than resembles it. A five-word question like "how do I reset MFA" and an 800-token section of an identity-management policy share almost no surface form, and their relationship is answer-hood, not similarity. Using a symmetric model for asymmetric retrieval is one of the most common quiet failures in a RAG build. Nothing crashes; recall is simply mediocre, and the team blames chunking. ## Prefixes and instructions Many retrieval-tuned encoders were trained with a literal text marker distinguishing the two sides, and they expect it at inference. The E5 family is the clearest published example, prepending `query:` to queries and `passage:` to documents. Several BGE releases instead attach a short instruction to the query side only. Later instruction-tuned encoders generalize this into a task description prepended to the query. Three rules follow: 1. **Apply the prefix the model was trained with, verbatim.** Inventing your own wording is not equivalent. 2. **Apply it on the correct side.** Prefixing documents with the query marker is a real, observed bug and it degrades ranking without any signal. 3. **Apply it consistently across time.** If you re-embed part of the corpus after changing the convention, that part of the index lives in a slightly different region of the space than the rest. If a model's documentation says no prefix is needed, adding one is also a change, not a no-op. ## Normalization and the metric Most retrieval encoders are trained with cosine similarity, so the vector's *direction* carries the meaning and its *magnitude* mostly does not. L2-normalizing every vector to unit length makes the dot product numerically identical to cosine similarity, which is convenient: many indexes execute inner product faster than cosine, and normalized vectors also behave better under quantization and under some ANN structures. The trap is consistency between three places: whether you normalize documents, whether you normalize queries, and which distance metric the index is configured with. Normalizing at ingestion but not at query time, or storing normalized vectors in an index configured for raw Euclidean distance, produces a ranking that is wrong in a way no exception will reveal. Pick one convention, encode it in a single shared function that both the ingestion job and the query path call, and assert on it. ## Length, truncation and what the chunk should be Every encoder has a maximum sequence length. Feed it more and the excess is silently dropped — the tail of an 800-token chunk may simply not exist as far as the model is concerned, even though your chunker thought it was indexed. Check the model's limit against your chunk-size distribution before ingestion, not after a quality complaint. The converse also matters. A very long chunk is compressed into the same fixed-length vector as a short one, so its representation becomes a blur of several topics and matches everything weakly. That is a real reason to prefer moderate chunks even when the model could technically accept longer ones. ## Making the short side carry more signal When a five-word query is genuinely too thin, the usual moves are to enrich one side rather than change the model: expand the query into a fuller question or a hypothetical answer before encoding, or prepend contextual framing (document title, section heading, parent-document summary) to each chunk before embedding so the chunk's vector knows where it lives. Both are cheap and both are testable against the same in-domain query set you would use to compare encoders. ## What good answers include An interviewer is checking that you know the model imposes a *protocol*, not just an API call: the right training regime, the exact prefix convention, a normalization decision that matches the index metric, and a sequence-length budget — all pinned identically on the ingestion path and the query path. Candidates who treat embedding as "call the model on the text" miss every one of these, and every one of them fails silently.

  • How would you catch a prefix or normalization mismatch between the ingestion job and the query path in production?
    Store the contract in index metadata — model name, model version, prefix strings, whether vectors are normalized, and the index metric — then assert it at query time and fail loudly on a mismatch. Back that with a smoke eval: a handful of query/expected-chunk pairs run after every deploy, since a mismatch shows up as a sharp recall drop on known-good pairs long before users report vague answers.
  • When does prepending document titles or section headings to a chunk before embedding actually help?
    It helps most when chunks are fragments whose subject is stated only in an ancestor — a numbered clause, a table row, a step in a procedure — because the chunk's own text lacks the terms a user would search for. It helps least when chunks are already self-contained, where the added boilerplate is repeated across many chunks and pulls their vectors closer together, blurring rather than sharpening the index.

It is like a lost-property desk: the description you shout at the counter and the item on the shelf look nothing alike, so the clerk must be trained to match descriptions to objects, not objects to objects.

saying these in an interview costs you the question

  • Uses a sentence-similarity model for short-query-to-passage retrieval
  • Swaps the query and passage prefixes, or applies both to both sides
  • Normalizes at index time but not at query time
  • Assumes text longer than the model's limit is still embedded
  • Thinks a mismatched convention will raise an error

context