In Cohere's Embed API, why must input_type differ for documents and queries?
answer
- the model asks what the text is for
- queries and passages are not alike
- a task marker is prepended before encoding
- wrong value, no error, worse recall
- baked into the vector, so re-embed
basics
~20 sCohere's Embed models are asymmetric: input_type tells the model what the text is for, and it conditions the vector accordingly. Corpus chunks go in as search_document, user queries as search_query, so a short question lands near the long passage that answers it.
solid answer
~50 sCohere's current Embed models are trained asymmetrically. `input_type` is a required parameter that tells the model the *role* of the text, and the model conditions the resulting vector on it — internally it is prepended as a task marker before encoding. The retrieval pair is `search_document` for the passages you index and `search_query` for the user's question; the two are trained so that a short interrogative query lands near the long declarative passage that answers it, which plain symmetric encoding does badly. The other values serve non-retrieval work: `classification` for vectors fed to a downstream classifier, `clustering` for topic grouping, and `image` for image inputs. Getting it wrong is dangerous precisely because it is not an error — embedding queries as `search_document` returns perfectly valid vectors and quietly costs you recall. And because the value is baked into the stored vector, changing it later means re-embedding the corpus.
go deeper
Remember the retrieval pair: corpus chunks are embedded with search_document, the user's question with search_query, and the parameter is required rather than optional.
Explain asymmetry — short interrogative queries versus long declarative passages — and that the value is prepended as a task marker so the same string yields different vectors.
Show you know the failure is silent: no error, just lost recall, detectable only on an eval set. Talk about wrapping the two calls so the value is never a caller's choice.
Own it as schema: model id, embedding type and input type together define an index generation, and changing any of them is a re-embed and rebuild you must plan and budget for.
## The parameter `input_type` is a required field on Cohere's `/v2/embed` request for the current Embed generations. It takes one of a small set of values: - `search_document` — text that will be **stored** in an index and later retrieved. - `search_query` — text that a user is **searching with**. - `classification` — text whose vector will be a feature for a classifier. - `clustering` — text whose vector will be grouped by similarity. - `image` — image input on the multimodal models. It is not documentation and not telemetry. It changes the numbers that come back. ## Why an embedding model would care Naively you would expect one text to have one embedding. That is the *symmetric* assumption, and it fits symmetric problems: comparing two paraphrases, deduplicating headlines, clustering support tickets. Retrieval is not symmetric. A query is short, interrogative, keyword-ish and often ungrammatical ("reset pw admin"). A passage is long, declarative and full of surrounding context. Those two texts are stylistically nothing alike even when one perfectly answers the other, and a single symmetric encoder must trade off matching *like with like* against matching *question with answer*. Asymmetric embedding models resolve that by conditioning on the role. Cohere implements this behind `input_type`: the value is turned into a task marker prepended to the text before encoding, so the same string yields a different vector under `search_query` than under `search_document`. The training objective pulls query-role vectors toward the document-role vectors of the passages that answer them. Other model families expose the same idea through instruction prefixes you write yourself; Cohere makes it an API parameter so you cannot forget the exact prefix string. ## The failure mode: silent, not loud This is the part interviewers push on. If you embed everything as `search_document` — including the queries — the API returns valid vectors, similarity search runs, and results come back. They are simply worse. Typical symptoms are a retrieval system that "kind of works": exact keyword overlaps still rank fine, paraphrased questions miss, and top-1 accuracy on an eval set sits several points below what the model is capable of. There is no exception, no warning field, and no way to detect it from the response — only from an evaluation set or a careful code review of your ingest and query paths. The reason it happens is structural: ingest and query usually live in different code paths, often different services, written at different times. The ingest job passes `search_document`; someone later writes the query path by copy-pasting the ingest call. A good defence is to wrap the two calls in two named functions — `embed_for_index()` and `embed_for_query()` — so the value is never a caller's choice, and to assert on it in tests. ## Non-retrieval values `classification` and `clustering` exist because those tasks want a different geometry from retrieval. For classification you want vectors that separate along the label boundaries you care about; for clustering you want vectors whose neighbourhood structure reflects topical grouping rather than question–answer relevance. Use them for their stated purpose, and use them **consistently**: a classifier trained on `classification` vectors must be served `classification` vectors, exactly as a retrieval index built with `search_document` must be queried with `search_query`. ## Consequences for the index lifecycle Because the value is fused into the stored vector, `input_type` behaves like part of the index schema, alongside the model id and the embedding type: - **Changing it is a full re-embed.** You cannot mix vectors built under different input types in one index and expect coherent distances. - **Record it.** Store the model id, embedding type and input type alongside the collection so a future engineer can tell what the vectors mean. - **Keep the raw text.** Re-embedding is the standard migration path for any of these changes, and it is only possible if the source text is still around. ## What a strong answer sounds like Name the parameter, name the retrieval pair, explain asymmetry in one sentence about queries versus passages, then land on the operational point: the failure is silent and the fix is a re-embed. Candidates who only recite the enum values are reciting docs; candidates who mention the silent recall loss have run this in production.
- You inherit an index and cannot tell which input_type built it. How do you find out?You cannot read it back from the vectors, so treat it empirically: build a small labelled query/answer eval set, then measure retrieval quality with queries embedded as `search_query` versus `search_document` against the existing index. Whichever scores materially better tells you how the corpus was embedded. Then record the answer as index metadata, and if the corpus turns out to be wrong, re-embed from the stored source text.
- Does input_type matter when you are only deduplicating near-identical documents?Much less — deduplication is a symmetric comparison of like with like, so both sides should simply use the same value. `clustering` is the natural choice since you are grouping by similarity rather than answering questions. The rule that survives is consistency: both sides of any comparison must be embedded under the same input type, whatever it is.
- Can you mix search_document and search_query vectors in one collection?No — distances between them are not meaningful in the way the model intended, and nearest-neighbour results become a mix of two geometries. Index one role only, typically `search_document`, and embed queries at request time as `search_query`. If you genuinely need query-role vectors stored, for example to cluster past searches, put them in a separate collection.
saying these in an interview costs you the question
- Treating input_type as optional metadata or telemetry
- Using search_document for both indexing and querying
- Expecting an error when the input type is wrong
- Believing one text always has exactly one embedding
- Thinking you can change input_type without re-embedding