skip to content

Why does each Cohere rerank result carry an index instead of the document text?

level: middleimportance: must knowfreq 55%

answer

  1. the response is a permutation
  2. rank position is not identity
  3. join by result.index, not loop i
  4. you already hold the text you sent
  5. silent wrong-document bug

basics

~20 s

Rerank returns a reordered list, so each result must say which input it came from. The index is the document's position in the request's documents array, which is how you join back to your own records without paying to send the text back.

solid answer

~50 s

The response is a permutation of your input, so position in `results` no longer tells you which document you are looking at — `index` does. It is the zero-based position that document occupied in the `documents` array you sent, so the correct join is `candidates[result.index]`, never `candidates[i]` from the enumeration loop. Getting that wrong is silent: you still get plausible-looking passages, just the wrong ones, and it usually surfaces as unexplained quality loss rather than an error. Cohere also defaults `return_documents` to false precisely because you already hold the text; echoing it back would double the bytes on the wire for no gain. Set it to true only when the caller genuinely has no candidate list to join against — for instance a thin proxy service. Alongside `index` you get `relevance_score`, a normalised 0-1 measure of how well that document answers this query, and the results arrive already sorted by it in descending order.

code

python · 23 lines
python
import cohere

co = cohere.ClientV2("<COHERE_API_KEY>")

candidates = [
    {"id": "doc-a", "text": "Refunds are issued within five business days."},
    {"id": "doc-b", "text": "The warehouse closes at 6pm."},
    {"id": "doc-c", "text": "To request a refund, open the order and choose Return."},
]
texts = [c["text"] for c in candidates]

response = co.rerank(
    model="rerank-v3.5",
    query="how do I get my money back?",
    documents=texts,
    top_n=2,
)

ranked = [
    {**candidates[r.index], "score": r.relevance_score}
    for r in response.results
]
print([r["id"] for r in ranked])

go deeper

for a junior

Remember that the response tells you which document by number, not by text, and that you look up your own candidate list with result.index.

for a middle

Explain why a reordered response needs an index at all, demonstrate the correct join, and describe the silent failure mode when someone maps by loop position instead.

for a senior

Talk about keeping the sent list and the object list immutable together, logging doc id plus score plus rank for evaluation, and why return_documents stays off in a service that already holds the text.

for a principal

Own the contract at the boundary: rerank returns positional references, so any caching, batching or candidate-mutation layer you introduce must preserve request-order identity or you get quality regressions no test catches.

## The response is a permutation Rerank does not filter and does not rewrite; it reorders. The `results` array you get back is your input list, sorted by relevance and optionally truncated to `top_n`. Because it is reordered, the array position in the response is the *rank*, and the only thing that identifies *which document* a result refers to is its `index` field. ``` request documents: [A, B, C, D] response results: [{index: 2, score: .91}, {index: 0, score: .44}] ``` Here the best document is C, the second is A, and B and D were dropped by `top_n`. Nothing in the response says "C" — only the number 2, pointing at the request array. ## The bug this design invites The classic defect is this: ```python for i, result in enumerate(response.results): picked.append(candidates[i]) # WRONG ``` That reads candidates 0 and 1 in request order, ignoring the ranking entirely. It throws no exception, produces the right *number* of passages, and the answers merely get quietly worse. The correct form is: ```python picked = [candidates[result.index] for result in response.results] ``` The same mistake appears in a subtler form when the candidate list is rebuilt or re-sorted between building the request and reading the response — for example when deduplication runs in between. The parallel array you index into must be the exact list, in the exact order, that you serialised into `documents`. Build the request from that list and keep it alive until you have mapped the results. ## Why the text is not echoed by default `return_documents` defaults to false. The caller sent the text, so the caller already has it; echoing hundreds of passages back doubles response size and serialization time for zero information gain. Turning it on is justified in one situation: when whatever consumes the response does not hold the candidate list — a gateway that forwards a rerank call on behalf of another service, or an ad-hoc script exploring results in a notebook. In a normal RAG service, leaving it false and joining by index is both cheaper and clearer. ## Reading relevance_score Each result also carries `relevance_score`, a normalised value in the 0-1 range where higher means more relevant to this query. Two properties matter for using it correctly. First, the ordering it induces is the product you are paying for — the ranking is more trustworthy than any individual number. Second, it is a score against *this* query over *these* candidates; treating a raw number as a corpus-wide calibrated probability, or comparing scores across different queries as though they were on one absolute scale, is where teams get into trouble. Whatever you do with the numbers downstream should be validated on your own labelled data rather than assumed from the range. ## Practical shape of the join A robust implementation keeps three things together: the original candidate objects (with their own ids and metadata), the parallel list of strings actually sent, and the response. Map with `index`, attach the `relevance_score` onto your object for logging and evaluation, and carry your own document id forward into whatever you build next. Logging `(query_id, doc_id, relevance_score, rank)` for a sample of traffic is what later lets you measure whether the rerank hop is actually earning its latency.

  • What breaks if you deduplicate the candidate list after building the request?
    The index values point into the list you sent, so any mutation of the parallel array between request and response silently misaligns the join and you attach scores to the wrong documents. Build the string list and the object list together, treat both as immutable until the results are mapped, and deduplicate either before constructing the request or after mapping.
  • When is turning return_documents on actually worth the extra bytes?
    When the consumer of the response does not hold the candidate list — a proxy or gateway service reranking on someone else's behalf, or interactive exploration where you want to eyeball the text next to the score. In a normal service that just sent the documents, it doubles payload size and adds nothing, so leave it false and join by index.
  • Can two results in one response share the same index?
    No. The response is a permutation of the input list, optionally truncated by top_n, so each input document appears at most once and every index is distinct. If you see duplicates in your ranked output, the duplication came from your candidate list — the same passage retrieved twice under different ids — not from the rerank call.

saying these in an interview costs you the question

  • Uses the loop counter instead of result.index to map back
  • Thinks index is the document's rank in the response
  • Assumes index is an id from the vector store
  • Re-sorts results that already arrive ranked
  • Treats relevance_score as a calibrated cross-query probability

context