Why do embedding models like E5 and BGE need different prefixes for queries and passages?
answer
- queries and documents are not interchangeable
- role marker, not decoration
- trained on question/answer pairs, not paraphrases
- fails silently, scores still look normal
- same convention at index and query time
basics
~20 sThey were trained asymmetrically: short questions and long documents are mapped into the shared space by different roles, and the prefix tells the model which side it is encoding. Encode both sides identically and retrieval quality degrades silently.
solid answer
~50 sRetrieval embedding models are trained on (query, passage) pairs, and the two sides do not look alike — a query is short, often interrogative, and states a need; a passage is long, declarative, and states facts. To make them land near each other, the model has to treat them as distinct roles. Some families do that with two separate encoders (as in Dense Passage Retrieval); the E5 and BGE families do it with a single encoder plus a role marker, so E5 expects `query:` and `passage:` prefixes, and BGE v1.5 expects an instruction prefix on the query side only. The prefix is part of the input the model was trained on, not decoration. The dangerous property is that omitting or mismatching it throws no error: scores still look plausible, sitting in the usual high band, while recall quietly drops. Always check the model card and apply the same convention at index time and query time.
code
python · 10 linesfrom sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/e5-base-v2")
query = model.encode("query: how do I rotate an API key?", normalize_embeddings=True)
doc = model.encode(
"passage: To rotate a key, open Settings, choose Credentials, then Rotate.",
normalize_embeddings=True,
)
print(float(query @ doc))go deeper
Know that some embedding models expect a marker such as E5's query: and passage:, that it comes from the model card, and that it must be applied consistently.
Explain why the space is asymmetric — models are trained on question/answer pairs, so a query is mapped near the passages that answer it — and describe the two implementations: separate encoders or one encoder plus a role prefix.
Show how you would catch the silent failure in production: labelled retrieval regression tests on deploy, a single shared encoding function with the role as a required argument, and prefix strings pinned alongside the model version.
Frame the prefix as part of the model contract, so a model swap is a coordinated change across ingestion and serving, with a reindex plan and an evaluation gate rather than a config edit.
## The asymmetry is a property of the trained space Semantic retrieval asks a strange thing of an embedding space: it must place a short question next to a long answer even though the two texts barely overlap in surface form. "How do I rotate an API key?" and a paragraph of key-rotation instructions share few words and have completely different shapes. A space trained only on symmetric similarity — sentences that mean the same thing — will not do this well, because it will judge the question as most similar to *other questions*. So retrieval models are trained on pairs where the two sides play different roles, and the geometry that results is **asymmetric**. Query vectors and document vectors do not live in interchangeable regions; the model learned a mapping in which a query is placed near the documents that *answer* it. Feeding a document through the query path, or vice versa, asks the model for the wrong mapping. ## Two ways of expressing the asymmetry **Two towers.** Some architectures use physically separate encoders, one for questions and one for passages, trained jointly. Dense Passage Retrieval is the classic example. Here the asymmetry is unmissable: there are literally two models, and calling the wrong one is a code-level mistake. **One encoder plus a role marker.** More common now is a single shared encoder with a text prefix that tells it which role this input plays. The E5 family expects the literal strings `query: ` and `passage: ` in front of the text. The BGE English v1.5 models take an instruction prefix on the query side ("Represent this sentence for searching relevant passages: ") and recommend no instruction for the passages themselves. Other instruction-tuned embedding models generalize this into a task instruction supplied per call. The details differ per model family and per release, so the model card is the authority. ## Why the failure is silent This is the part interviewers probe. If you drop the prefixes and embed queries and documents identically, nothing throws. The prefix is just text; the encoder happily produces a vector. The similarity scores come back in the same high band they always occupy, because the space is anisotropic and everything scores high anyway. Nothing in the response distinguishes a well-formed retrieval from a degraded one. What actually happened is that both sides were mapped through the same role, so the model is now doing symmetric similarity. Queries drift toward other question-shaped text; retrieval starts favouring documents that *sound like* the question — FAQ headers, other tickets, chunks that happen to be phrased interrogatively — over the passages that answer it. Recall on the answers you care about drops, often by a large margin, and the only way to see it is an evaluation on labelled query/document pairs. The symmetric variant of the bug is just as common: the ingestion job applies the passage prefix, then a later query service is written by a different team, or a model is swapped without updating the prefix constant, and now index-time and query-time conventions disagree. The corpus is fine, the queries are fine, and results are subtly worse forever. ## Practical discipline Put the prefixing in one shared function, with the role as a required argument, so that no call site can forget it. Pin the prefix strings next to the model identifier, because they are part of the model contract; changing the model without changing the prefixes is a breaking change. Add a regression test built from a handful of labelled query/document pairs that asserts the correct passage is ranked first — this catches the whole class of silent mismatch, including a partially reindexed corpus. When you swap models, re-check whether the new family uses prefixes at all. Some do not, and prepending `query:` to a model that never saw it in training is just noise injected into every embedding. ## Related consequences Because query and document vectors sit in different regions, some cross-side statistics stop being meaningful. Clustering a mixed pool of query and document embeddings will tend to separate them by role rather than by topic. Similarly, a similarity threshold calibrated on document-to-document pairs does not transfer to query-to-document scoring, because the two distributions are not the same. Keep the role in mind whenever you compare two embeddings: the number only means something when both sides were produced the way the model expects.
- How would you detect that a running system had lost its query prefix?Score-level monitoring will not show it, so evaluate on labelled pairs: keep a small set of queries with their known-correct passages and assert the target is retrieved in the top-k. A sudden drop in that hit rate on a deploy points at the prefix. A softer signal is the character of the results — if answers are being displaced by other question-shaped chunks, both sides are probably being encoded in the same role.
- Does this asymmetry change how you would cluster a mixed pool of embeddings?Yes. If you embed queries and documents with their respective roles and then cluster the pool together, the strongest split you find will often be role rather than topic, because the two groups occupy different regions of the space. Cluster each side separately, or embed everything in one role if the task is genuinely symmetric — but then accept that it is no longer the retrieval mapping the model was trained for.
- Is a threshold calibrated on document-to-document pairs valid for query-to-document scoring?No. The two score distributions are produced by different mappings, so their baselines and their spreads differ. A cutoff derived from deduplicating documents against each other will be systematically wrong when applied to query-to-document retrieval. Calibrate each comparison type separately on labelled examples of exactly that comparison.
saying these in an interview costs you the question
- Assuming any text can be embedded the same way for retrieval
- Calling the prefixes cosmetic or purely documentation
- Expecting an error when the prefix is omitted
- Prefixing at query time but not during indexing
- Copying prefix strings across unrelated model families