skip to content

How does Cohere's v2/rerank handle a document longer than max_tokens_per_doc?

level: middleimportance: should knowfreq 42%

answer

  1. truncation, not chunking
  2. the tail is never scored
  3. v1 split, v2 cuts
  4. default is 4096 tokens per document
  5. chunk at ingestion instead

basics

~20 s

It truncates the document to that token budget (4096 by default) and scores only the kept portion; the overflow is invisible to the model. The v1 API instead split long documents into chunks and kept the best chunk's score, which v2 dropped.

solid answer

~50 s

In the v2 API, `max_tokens_per_doc` is a hard truncation limit — default 4096 tokens — and anything past it simply is not read. That is a real behaviour change from v1, where `max_chunks_per_doc` split an over-long document into chunks, scored each, and used the best chunk's score for the document. So a long PDF page whose answer lives in the last paragraph will now score as if that paragraph did not exist, and it will score low, quietly, with no error. The fix is to do the chunking yourself: split documents into passages at ingestion, rerank the passages, then map back to parent documents and deduplicate if several passages of the same document survive. That is usually what you want anyway, because the passage you rerank is the passage you feed the model. If you are migrating from v1, treating `max_chunks_per_doc` as a drop-in rename of `max_tokens_per_doc` is a silent quality regression, not a compile error.

code

python · 26 lines
python
import cohere

co = cohere.ClientV2("<COHERE_API_KEY>")

# passages carry their parent id so we can dedupe after ranking
passages = [
    {"doc_id": "policy-1", "text": "Section 1. Scope of the returns policy."},
    {"doc_id": "policy-1", "text": "Section 9. Refunds are issued to the original card."},
    {"doc_id": "hours-1", "text": "Stores open at 9am."},
]

response = co.rerank(
    model="rerank-v3.5",
    query="where is my refund paid back to?",
    documents=[p["text"] for p in passages],
    top_n=3,
    max_tokens_per_doc=512,
)

seen, kept = set(), []
for r in response.results:
    p = passages[r.index]
    if p["doc_id"] not in seen:
        seen.add(p["doc_id"])
        kept.append(p)
print(kept)

go deeper

for a junior

Know that documents longer than the per-document token limit are cut off, not rejected, and that the part beyond the cut is not scored at all.

for a middle

Explain the v1 chunking versus v2 truncation difference, name max_tokens_per_doc and its 4096 default, and say why passages should be chunked before they reach the rerank call.

for a senior

Diagnose the symptom — long documents never surfacing while short ones dominate — and describe the migration audit: token-length distribution, explicit chunking, parent-id dedupe, and a re-run evaluation set.

for a principal

Treat silent semantic changes across a vendor's API versions as a class of risk: this one returns 200 and degrades quality. Argue for retrieval evaluation in CI so a provider migration cannot ship unmeasured.

## The parameter and its default Cohere's v2 rerank request accepts `max_tokens_per_doc`, which caps how much of each supplied document the model considers. The default is 4096 tokens, which also reflects the context the rerank model works within — query plus document have to fit. Documents shorter than the cap are scored in full. Documents longer than it are **truncated**: the model scores the retained prefix and never sees the rest. ## What changed from v1 The v1 rerank API had a different knob, `max_chunks_per_doc`. Its behaviour was fundamentally different: an over-long document was split into chunks internally, each chunk was scored against the query, and the document's score was taken from its best chunk, up to the configured chunk ceiling. That meant a long document with one highly relevant paragraph deep inside could still surface. The v2 API removed that mechanism. There is no server-side chunking; there is only truncation. This is the single most consequential difference for anyone porting a v1 pipeline, and it is dangerous precisely because it is not a breaking change at the type level — you drop one parameter, add another with a similar name, everything still returns 200, and recall on long documents quietly falls off. ## Why truncation hurts and how it shows up Long documents put their conclusions in inconvenient places. A support article buries the fix under three paragraphs of symptoms; a contract puts the liability clause on page nine. Under truncation, the query-relevant span is either inside the retained window (fine) or outside it (the document scores as though it were irrelevant). The failure is asymmetric and invisible: no error, no flag, just a good document ranked below a mediocre short one whose entire text was read. Symptoms in production look like: short FAQ entries dominating the top of every ranking; a known-good long document never appearing however the query is phrased; rerank *hurting* measured quality on a corpus of long files while helping on a corpus of short ones. ## The right shape: chunk before you rerank The durable answer is that reranking operates on passages, not whole files. Split at ingestion into passages that are meaningful on their own — a section, a few hundred tokens with overlap, a table plus its caption — and index those passages. Then: 1. Recall returns passages. 2. Rerank scores passages against the query. 3. You keep the top passages, dedupe by parent document if one document contributed several, and pass those passages into the prompt. This alignment matters beyond truncation: the unit you score is the unit you feed the model, so a high relevance score actually predicts that the retained text answers the question. Reranking a whole document and then feeding the model a different slice of it breaks that link. ## Tuning the parameter itself Lowering `max_tokens_per_doc` below the default is a deliberate trade: it caps the text the model reads per candidate, which reduces the work per document, but it increases the chance of cutting the answer off. Raising it is bounded by the model's context. In practice, if you find yourself wanting a big value, that is a signal your passages are too large and the chunking, not the parameter, is what needs work. ## Migration checklist When moving a v1 rerank pipeline to v2: search the codebase for `max_chunks_per_doc` and treat every hit as a design question, not a rename; measure the token length distribution of your candidates against the 4096 default; add explicit passage chunking wherever documents exceed it; keep parent-document ids on each passage so you can dedupe after ranking; and re-run your retrieval evaluation set before and after, because this change moves quality without moving any status code.

  • You are porting a v1 rerank call that set max_chunks_per_doc to 10. What do you change?
    Do not rename it to max_tokens_per_doc and move on — the semantics are different. v1 chunked long documents server-side and kept the best chunk's score; v2 only truncates. Introduce explicit passage chunking in your ingestion or retrieval layer, keep parent ids for post-rank deduplication, and re-run your retrieval evaluation set, since this change degrades long-document recall without any error.
  • What does lowering max_tokens_per_doc below the default buy you?
    It caps how much text the model reads per candidate, trimming the work each document costs, and it enforces discipline when your candidates are inconsistently sized. The price is a higher chance of cutting off the span that actually answers the query. If a small value hurts quality, the real fix is smaller, self-contained passages rather than a bigger cap.
  • Several top-ranked passages come from the same source document. Is that a problem?
    Usually yes, if you are filling a limited prompt budget — one document monopolises the context and the answer loses breadth. Carry a parent id on every passage and apply a per-document cap after ranking, keeping the highest-scoring passage or two per source. Do that after reranking, not before, so the scoring still sees all candidates.

saying these in an interview costs you the question

  • Thinks v2 still chunks long documents server-side
  • Treats max_tokens_per_doc as a rename of max_chunks_per_doc
  • Expects an error when a document exceeds the limit
  • Assumes truncation keeps the most relevant part
  • Reranks whole files but feeds the model different slices

context