How is a Cohere Rerank call billed, and does a smaller top_n make it cheaper?
answer
- priced per search unit, not per token
- one unit covers up to 100 documents
- cost follows candidates sent
- step function at the block boundary
- meta.billed_units reports the truth
basics
~20 sRerank is billed in search units, where one unit covers a single query against up to 100 documents. A smaller top_n changes nothing: every document you send is still scored, so the cost lever is the candidate count, not the number of results returned.
solid answer
~50 sCohere prices Rerank in **search units** rather than tokens: one search unit is one query with up to 100 documents, so a call carrying 250 candidates bills as three units and a call with 40 bills as one. The response reports what you were charged under `meta.billed_units.search_units`, which is what you should log rather than inferring cost from your own arithmetic. The consequence people miss is that `top_n` is a response-shaping parameter, not a cost control — all supplied documents are scored regardless of how many results come back, so dropping `top_n` from 20 to 5 saves you nothing. The real levers are how deep your recall stage goes, whether you rerank at all for cheap or cached queries, and whether you are sending near-duplicate candidates. Note also the block boundary: going from 100 candidates to 101 doubles the bill for that call, so candidate depth is worth aligning to the block size.
go deeper
Know that Rerank is not billed per token — it is billed per search unit, roughly one query against a batch of up to 100 documents.
Do the arithmetic out loud: 250 documents is three units, and top_n does not change it because every supplied document is still scored.
Show the instrumentation and the levers: log meta.billed_units.search_units per query, tune recall depth against a quality curve, dedupe before the call, and align depth to the 100-document block boundary.
Frame rerank spend as a function of one tunable — recall depth — that also drives latency, and argue for choosing it from measured quality gains rather than intuition, with fan-out patterns budgeted explicitly.
## Search units, not tokens Most LLM endpoints bill per token in and per token out. Rerank does not: it is priced in **search units**, defined as one query evaluated against up to 100 documents. This is a coarse unit by design — it makes the cost of a rerank hop a function of how many candidates you feed it, and nothing else. The arithmetic is a ceiling division: - 1 query, 40 documents -> 1 search unit - 1 query, 100 documents -> 1 search unit - 1 query, 101 documents -> 2 search units - 1 query, 250 documents -> 3 search units And because each request carries exactly one query, ten queries are ten requests and at least ten search units. ## Read the number, don't compute it The response carries `meta.billed_units.search_units`. Log that field per request alongside your query id. It is the vendor's own accounting, it survives any future change to the block size, and it lets you build a real cost-per-query dashboard instead of an estimate. Teams that infer billing from their own candidate counts drift as soon as retry logic, fan-out or a fallback path adds calls they forgot to count. ## Why top_n is not a cost control A rerank model must score every candidate to know which ones are best; there is no way to rank the top five without evaluating all of them. `top_n` therefore applies after scoring, trimming the response. Lowering it reduces response bytes and your own downstream work, but the billed unit count is identical. The same goes for `return_documents`: it changes payload size, not price. This matters because `top_n` is the parameter most people reach for when asked to cut rerank spend, and it is the wrong one. The parameters and design decisions that actually move the bill are: - **Recall depth.** Whether the vector stage returns 50, 100 or 300 candidates directly sets the unit count. This is the primary lever. - **Whether you call at all.** Very short or navigational queries, exact-id lookups, and repeat queries served from a cache can skip the hop entirely. - **Duplicate candidates.** Sending the same passage twice, or ten near-identical boilerplate chunks, buys nothing and can push you across a block boundary. Dedupe before the call. - **Fan-out.** Multi-query expansion, where one user question becomes four searches, multiplies rerank calls too. Deduplicate the merged candidate set before reranking rather than reranking each branch separately. ## The block boundary is a real design constraint Because a unit covers *up to* 100 documents, cost is a step function, not a line. Candidate depths of 100, 200 and 300 are efficient; depths of 105 or 210 pay a full extra unit for a handful of documents that rarely change the top of the ranking. If your recall stage is configurable, snapping its depth to the block boundary is free money, and if a filter sometimes trims candidates below a boundary the saving is automatic. ## Putting cost next to value Rerank is normally one of the cheapest quality improvements in a retrieval pipeline, so the interview point is not miserliness — it is knowing the unit and being able to say where the money goes. A defensible answer: instrument `search_units` per query, hold recall depth as a tunable, evaluate quality at several depths, and pick the depth where the quality curve flattens rather than the one that felt safe. The extra latency of the hop and the extra units both scale with the same knob, which makes depth the single number worth arguing about.
- Your recall stage returns 105 candidates. What is the cheap change?Trim to 100 or raise to 200 deliberately. A search unit covers up to 100 documents, so 105 candidates bills two units — a full extra unit for five documents that almost never reach the top of the ranking. Either snap recall depth to the block boundary, or spend the unit you are already paying for by going to 200 and measuring whether the deeper pool actually improves results.
- How does query expansion into four sub-queries affect rerank cost?Naively it quadruples it, because each sub-query is its own rerank request. The cheaper shape is to run the four retrievals, merge and deduplicate the union of candidates, then issue one rerank request with the user's original query against that merged set. You pay for the merged candidate count once instead of four separate calls, and you avoid scoring the same passage several times.
- Where do you get the authoritative number for what a rerank call cost you?From the response itself: meta.billed_units.search_units. Log it per request with the query id and recall depth. Computing it from your own candidate count works until retries, fallbacks or a fan-out path add calls your accounting does not see, and it breaks silently if the vendor ever changes the block size.
saying these in an interview costs you the question
- Thinks lowering top_n reduces the bill
- Assumes Rerank is priced per token like chat endpoints
- Believes one request always equals one search unit
- Ignores that recall depth is the real cost lever
- Sends duplicate candidates and pays for them twice