What does a Cohere v2/rerank request contain, and what does the API return?
answer
- one query, a list of candidate strings
- top_n trims the response only
- results come back already ranked
- each result points back by index
- v2 documents are plain strings
basics
~20 sA Cohere v2/rerank call sends a model name, one query string, and a documents list of plain strings, plus an optional top_n. It returns a results array ordered by relevance, each entry carrying the document's original index and a relevance_score.
solid answer
~40 sYou POST to `https://api.cohere.com/v2/rerank` (or call `co.rerank(...)` on a `ClientV2`) with four things that matter: `model`, the `query` string, the `documents` list, and optionally `top_n`. In the v2 API `documents` is a list of **plain strings** — if your candidates are structured records you stringify the relevance-bearing fields yourself before sending. The response is a `results` array sorted by descending `relevance_score`, and each entry gives you `index` — the position that document occupied in the request list — so you join back to your own objects. `return_documents` defaults to false, so by default you get indexes and scores, not the text back. The response also carries `meta.billed_units.search_units`. It is one extra network round trip placed after vector recall and before you build the prompt.
code
python · 17 linesimport cohere
co = cohere.ClientV2("<COHERE_API_KEY>")
response = co.rerank(
model="rerank-v3.5",
query="How do I reset my password?",
documents=[
"Password resets are available from the account settings page.",
"Our office is open from 9am to 5pm on weekdays.",
"Contact support to unlock a disabled account.",
],
top_n=2,
)
for result in response.results:
print(result.index, result.relevance_score)go deeper
Be able to name the four things you send — model, query, documents, optional top_n — and say that what comes back is a ranked list of indexes and relevance scores, not rewritten text.
Explain why the response returns indexes rather than documents, what return_documents changes, and how you flatten structured records into the plain strings the v2 API expects.
Show where the call sits relative to vector recall, how candidate count and payload size drive its latency and cost, and how you keep the model id in configuration so a model retirement is a config change.
Frame Rerank as a bought capability with a contract you depend on: a single-vendor hop inside your search path. Be ready to discuss the abstraction you put around it so scoring can be swapped without rewriting the pipeline.
## What the endpoint is for Cohere's Rerank API takes a single query and a list of candidate documents and returns those candidates re-ordered by how well each one answers the query. It does not generate text and it does not retrieve anything: you must already have candidates, typically from a vector or keyword search. In a retrieval pipeline it sits between recall and generation. ## The request A v2 request (`POST https://api.cohere.com/v2/rerank`, or `co.rerank(...)` in the Python SDK) has these fields: - **`model`** (required) — the rerank model id, for example `rerank-v3.5` in the mid-2026 line-up. Model ids change; keep it in config, not hard-coded across the codebase. - **`query`** (required) — a single string. There is exactly one query per call; you cannot batch multiple queries into one request. - **`documents`** (required) — the candidate list. In the **v2** API these are plain strings. The v1 API accepted objects together with a `rank_fields` parameter that told the model which JSON keys to read; v2 dropped that, so structured records must be flattened by you (a small YAML- or JSON-shaped string per record works well). - **`top_n`** (optional) — how many ranked results to return. Omit it and you get every document back, ranked. - **`max_tokens_per_doc`** (optional) — how much of each document is considered; the default is 4096 tokens. Longer documents are truncated. - **`return_documents`** (optional, default false) — echo the document text back in each result. ## The response ``` { "id": "...", "results": [ { "index": 4, "relevance_score": 0.93 }, ... ], "meta": { "billed_units": { "search_units": 1 } } } ``` Three things to internalise: 1. **`results` is already sorted** by descending `relevance_score`. You do not sort it again. 2. **`index` is the position in the request's `documents` list**, not a rank and not an id from your database. It is the join key back to your own candidate objects. 3. **`relevance_score` is a normalised 0-1 number** expressing how well that document answers *this* query. It is a ranking signal; treat comparisons across different queries with care. ## How it composes with the rest of the pipeline The usual shape is: embed the query, pull 50-200 candidates from the vector store, send those candidates to Rerank, keep the top handful, and put only those into the model's context. The recall stage is optimised for not missing anything; the rerank stage is optimised for ordering what survived. Because the call is a network hop with the full candidate text in the request body, payload size, not just candidate count, drives its latency. ## Common first-time mistakes Sending embeddings instead of text — Rerank reads the raw text, not vectors. Expecting the document text back without setting `return_documents`. Expecting `top_n` to make the call cheaper (it only trims the response). Assuming a document longer than the per-document token limit is fully read. And, most often, iterating the results by loop position instead of using each result's `index` to map back to the original candidate.
- What comes back if you leave top_n out entirely?Every document you sent is returned, ranked by descending relevance_score. That is useful when you want the full ordering — for example to apply your own downstream filter or to log scores for evaluation — but it makes the response as large as the candidate list, so production paths normally set top_n to the number of passages they will actually put in the prompt.
- Your candidates are database rows, not strings. How do you rerank them?Flatten each row into one string containing only the fields that carry relevance signal — title, body, maybe a category — and send that list, keeping a parallel array of the original rows. A compact YAML- or JSON-shaped string per record works well. Then map results back through each result's index. Padding the string with ids, timestamps and boilerplate wastes the per-document token budget.
- Can you rerank several queries in one call?No. Each rerank request carries exactly one query string. Multiple queries mean multiple requests, which you can fire concurrently. This also matters for cost, since billing is counted per request against the number of documents it carried.
Search gives you a stack of resumes that broadly match; Rerank is the hiring manager who reads each one against the actual job description and hands the stack back in order.
saying these in an interview costs you the question
- Thinks Rerank retrieves documents or generates an answer
- Sends embedding vectors instead of raw document text
- Assumes the response echoes document text by default
- Believes top_n limits how many documents get scored
- Iterates results by loop position instead of the index field