skip to content

Rerank API

A cross-encoder as a service: send a query plus candidate documents and get relevance-scored ordering back. It is the cheapest quality win in most retrieval pipelines, paid for with one extra network hop.

on this pageshow

questions

6

What does a Cohere v2/rerank request contain, and what does the API return?

level: juniorimportance: must knowfreq 68%

answer

  1. one query, a list of candidate strings
  2. top_n trims the response only
  3. results come back already ranked
  4. each result points back by index
  5. v2 documents are plain strings

basics

~20 s

A Cohere v2/rerank call sends a model name, one query string, and a documents list of plain strings, plus an optional top_n. It returns a results array ordered by relevance, each entry carrying the document's original index and a relevance_score.

solid answer

~40 s

You POST to `https://api.cohere.com/v2/rerank` (or call `co.rerank(...)` on a `ClientV2`) with four things that matter: `model`, the `query` string, the `documents` list, and optionally `top_n`. In the v2 API `documents` is a list of **plain strings** — if your candidates are structured records you stringify the relevance-bearing fields yourself before sending. The response is a `results` array sorted by descending `relevance_score`, and each entry gives you `index` — the position that document occupied in the request list — so you join back to your own objects. `return_documents` defaults to false, so by default you get indexes and scores, not the text back. The response also carries `meta.billed_units.search_units`. It is one extra network round trip placed after vector recall and before you build the prompt.

code

python · 17 lines
python
import cohere

co = cohere.ClientV2("<COHERE_API_KEY>")

response = co.rerank(
    model="rerank-v3.5",
    query="How do I reset my password?",
    documents=[
        "Password resets are available from the account settings page.",
        "Our office is open from 9am to 5pm on weekdays.",
        "Contact support to unlock a disabled account.",
    ],
    top_n=2,
)

for result in response.results:
    print(result.index, result.relevance_score)

go deeper

for a junior

Be able to name the four things you send — model, query, documents, optional top_n — and say that what comes back is a ranked list of indexes and relevance scores, not rewritten text.

for a middle

Explain why the response returns indexes rather than documents, what return_documents changes, and how you flatten structured records into the plain strings the v2 API expects.

for a senior

Show where the call sits relative to vector recall, how candidate count and payload size drive its latency and cost, and how you keep the model id in configuration so a model retirement is a config change.

for a principal

Frame Rerank as a bought capability with a contract you depend on: a single-vendor hop inside your search path. Be ready to discuss the abstraction you put around it so scoring can be swapped without rewriting the pipeline.

## What the endpoint is for Cohere's Rerank API takes a single query and a list of candidate documents and returns those candidates re-ordered by how well each one answers the query. It does not generate text and it does not retrieve anything: you must already have candidates, typically from a vector or keyword search. In a retrieval pipeline it sits between recall and generation. ## The request A v2 request (`POST https://api.cohere.com/v2/rerank`, or `co.rerank(...)` in the Python SDK) has these fields: - **`model`** (required) — the rerank model id, for example `rerank-v3.5` in the mid-2026 line-up. Model ids change; keep it in config, not hard-coded across the codebase. - **`query`** (required) — a single string. There is exactly one query per call; you cannot batch multiple queries into one request. - **`documents`** (required) — the candidate list. In the **v2** API these are plain strings. The v1 API accepted objects together with a `rank_fields` parameter that told the model which JSON keys to read; v2 dropped that, so structured records must be flattened by you (a small YAML- or JSON-shaped string per record works well). - **`top_n`** (optional) — how many ranked results to return. Omit it and you get every document back, ranked. - **`max_tokens_per_doc`** (optional) — how much of each document is considered; the default is 4096 tokens. Longer documents are truncated. - **`return_documents`** (optional, default false) — echo the document text back in each result. ## The response ``` { "id": "...", "results": [ { "index": 4, "relevance_score": 0.93 }, ... ], "meta": { "billed_units": { "search_units": 1 } } } ``` Three things to internalise: 1. **`results` is already sorted** by descending `relevance_score`. You do not sort it again. 2. **`index` is the position in the request's `documents` list**, not a rank and not an id from your database. It is the join key back to your own candidate objects. 3. **`relevance_score` is a normalised 0-1 number** expressing how well that document answers *this* query. It is a ranking signal; treat comparisons across different queries with care. ## How it composes with the rest of the pipeline The usual shape is: embed the query, pull 50-200 candidates from the vector store, send those candidates to Rerank, keep the top handful, and put only those into the model's context. The recall stage is optimised for not missing anything; the rerank stage is optimised for ordering what survived. Because the call is a network hop with the full candidate text in the request body, payload size, not just candidate count, drives its latency. ## Common first-time mistakes Sending embeddings instead of text — Rerank reads the raw text, not vectors. Expecting the document text back without setting `return_documents`. Expecting `top_n` to make the call cheaper (it only trims the response). Assuming a document longer than the per-document token limit is fully read. And, most often, iterating the results by loop position instead of using each result's `index` to map back to the original candidate.

  • What comes back if you leave top_n out entirely?
    Every document you sent is returned, ranked by descending relevance_score. That is useful when you want the full ordering — for example to apply your own downstream filter or to log scores for evaluation — but it makes the response as large as the candidate list, so production paths normally set top_n to the number of passages they will actually put in the prompt.
  • Your candidates are database rows, not strings. How do you rerank them?
    Flatten each row into one string containing only the fields that carry relevance signal — title, body, maybe a category — and send that list, keeping a parallel array of the original rows. A compact YAML- or JSON-shaped string per record works well. Then map results back through each result's index. Padding the string with ids, timestamps and boilerplate wastes the per-document token budget.
  • Can you rerank several queries in one call?
    No. Each rerank request carries exactly one query string. Multiple queries mean multiple requests, which you can fire concurrently. This also matters for cost, since billing is counted per request against the number of documents it carried.

Search gives you a stack of resumes that broadly match; Rerank is the hiring manager who reads each one against the actual job description and hands the stack back in order.

saying these in an interview costs you the question

  • Thinks Rerank retrieves documents or generates an answer
  • Sends embedding vectors instead of raw document text
  • Assumes the response echoes document text by default
  • Believes top_n limits how many documents get scored
  • Iterates results by loop position instead of the index field

context

open as a page

Why does each Cohere rerank result carry an index instead of the document text?

level: middleimportance: must knowfreq 55%

basics

~20 s

Rerank returns a reordered list, so each result must say which input it came from. The index is the document's position in the request's documents array, which is how you join back to your own records without paying to send the text back.

open as a page

How does Cohere's v2/rerank handle a document longer than max_tokens_per_doc?

level: middleimportance: should knowfreq 42%

basics

~20 s

It truncates the document to that token budget (4096 by default) and scores only the kept portion; the overflow is invisible to the model. The v1 API instead split long documents into chunks and kept the best chunk's score, which v2 dropped.

open as a page

Which Cohere rerank model fits a corpus mixing English, French and Japanese?

level: middleimportance: should knowfreq 32%

basics

~20 s

Use a multilingual rerank model such as rerank-v3.5, which covers 100-plus languages and scores a query in one language against documents in another. You send one mixed candidate list to one call — no per-language routing and no translation step.

open as a page

How do you keep a Cohere Rerank hop from breaking a production search request?

level: seniorimportance: should knowfreq 36%

basics

~20 s

Treat Rerank as an optional enhancement, not a dependency: give it a timeout smaller than the remaining request budget and, on timeout, 429 or 5xx, serve the vector-recall order instead of failing. Log every degradation so the silent quality loss stays visible.

open as a page

How is a Cohere Rerank call billed, and does a smaller top_n make it cheaper?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Rerank is billed in search units, where one unit covers a single query against up to 100 documents. A smaller top_n changes nothing: every document you send is still scored, so the cost lever is the candidate count, not the number of results returned.

open as a page