How do you keep a Cohere Rerank hop from breaking a production search request?
answer
- you already have an answer without it
- degrade, never fail the request
- timeout derived from remaining budget
- the body carries every candidate's text
- count the fallbacks or never see them
basics
~20 sTreat Rerank as an optional enhancement, not a dependency: give it a timeout smaller than the remaining request budget and, on timeout, 429 or 5xx, serve the vector-recall order instead of failing. Log every degradation so the silent quality loss stays visible.
solid answer
~50 sThe pipeline already has an answer before the rerank call — the recall order — so the correct failure mode is degradation, not an error page. Set an explicit client timeout derived from the budget left after retrieval, and wrap the call so any timeout, connection error, 429 or 5xx falls back to the pre-rerank ordering. Retry at most once, with jitter, and only if the deadline still allows it; retrying into an exhausted budget converts a slow request into a failed one. Because the request body carries the full candidate text, payload discipline is also latency work: strip boilerplate, send only relevance-bearing fields, and pre-truncate passages rather than shipping whole documents the model will truncate anyway. Cache by the pair of query and candidate id set so repeated queries skip the hop. Finally, emit a metric on every fallback — degraded results look fine to users and to your error rate, so without that counter you can be serving unranked results for days without noticing.
code
python · 19 linesimport logging
import cohere
co = cohere.ClientV2("<COHERE_API_KEY>", timeout=0.4)
def rerank_or_fallback(query, candidates, top_n=5):
texts = [c["text"] for c in candidates]
try:
response = co.rerank(
model="rerank-v3.5",
query=query,
documents=texts,
top_n=top_n,
)
except Exception:
logging.warning("rerank_fallback", exc_info=True)
return candidates[:top_n]
return [candidates[r.index] for r in response.results]go deeper
Know that if the rerank call fails you can still return the search results you already have, and that the code should catch the error rather than let the request fail.
Describe the concrete wrapper: an explicit timeout, catch timeout and 5xx and 429, return the recall order, and log the fallback so it is not invisible.
Derive the timeout from the request's remaining budget, bound retries by the deadline, add a circuit breaker and a cache keyed on query plus candidate set plus model, and alert on fallback rate rather than error count.
Argue the position explicitly: rerank is a quality enhancement bought from a third party inside a latency-critical path, so it gets a kill switch, a measured quality delta and an availability budget separate from the components the product cannot run without.
## The key property: you already have an answer Rerank is unusual among the calls in a RAG pipeline because it is strictly an improvement over something you already hold. Retrieval has returned candidates in *some* order; rerank makes that order better. If the call fails, the pre-rerank order is still a serviceable result. That single fact should drive the whole error-handling design: **a rerank failure must never become a user-visible failure.** Compare that with the generation call, where a failure genuinely has no fallback. Teams that wrap every LLM-ish call in the same retry-and-propagate helper miss this distinction and turn a quality dip into an outage. ## Deadline first, retries second The request arrives with a latency budget. Retrieval consumes part of it; generation will consume most of the rest. What remains for rerank is a *derived* number, not a constant. Compute it, set the client timeout to it, and pass a deadline down rather than hard-coding a timeout at each layer. Retries then follow from the deadline: retry once, with jittered backoff, **only if** the remaining budget can absorb another attempt. A blind retry policy is actively harmful here — it spends the generation stage's budget on an optional improvement. On 429 specifically, respect any backoff signalling the API returns and prefer shedding the rerank hop for that request over queuing behind it; under a rate-limit episode the whole fleet is retrying at once, and adding load makes it worse. ## Payload is latency Unlike a chat call, a rerank request body carries the full text of every candidate. With 100 candidates of a few hundred tokens each, serialization plus transfer is a real component of the round trip, and it grows exactly when you are trying to improve recall. Practical levers: - Send only the fields that carry relevance signal. Ids, timestamps, ACL blobs and navigation chrome cost bytes and dilute the text the model reads. - Pre-truncate passages yourself to something close to what will be scored, rather than shipping whole files that get cut anyway. - Deduplicate candidates before the call: near-identical boilerplate chunks add payload, add nothing to ranking, and can push you across a billing block. - Keep the connection warm. Client reuse and connection pooling matter more than usual for a call in a latency-sensitive path. ## Cache the obvious repeats The rerank result is a pure function of the query string, the candidate set and the model. That makes it cacheable: key on a hash of the normalised query plus the ordered candidate ids plus the model id, and store the resulting ranking. Head queries in search traffic repeat heavily, and cached rankings remove both the latency and the search units. Invalidate on model change — which is another reason the model id belongs in configuration and in the cache key. ## Make degradation visible The dangerous property of a graceful fallback is that nobody notices it. Requests still succeed, latency may even improve, and error dashboards stay green while every user quietly gets pre-rerank ordering. So instrument it explicitly: - A counter for rerank fallbacks, split by cause (timeout, 429, 5xx, circuit open). - The rerank call's own latency histogram, separate from the end-to-end one. - An alert on fallback *rate*, not just on hard errors. - Ideally, a sampled quality signal — click-through or answer-acceptance on reranked versus fallback traffic — so you know what the degradation actually costs. ## Circuit breaking and staged rollout When the provider is having a bad day, a circuit breaker around the rerank client stops you from spending the deadline of every request on a call that is going to time out. Open the breaker on a sustained failure rate, serve recall order while it is open, and half-open probe periodically. Pair it with a feature flag so rerank can be turned off per tenant or globally without a deploy — which is also how you run the A/B that proves the hop is earning its latency in the first place.
- Rerank starts returning 429 across your fleet. Retry or shed?Shed for that request. A rate-limit episode means everyone is retrying at once, so adding attempts deepens it while spending the budget the generation stage needs. Respect whatever backoff the API signals, fall back to recall order immediately, and let a circuit breaker keep traffic off the endpoint until it recovers. Retrying is defensible only for an isolated request with plenty of deadline left.
- How would you cache rerank results, and what invalidates the entry?Key on the normalised query, the ordered candidate ids and the model id — the ranking is a pure function of those three. Head queries repeat heavily, so this removes both latency and search units. The model id in the key means a model change invalidates naturally; a document edit invalidates through the candidate id or content hash if your ids are content-derived.
- Your fallback works perfectly and nobody complains. Why is that a problem?Because a silent fallback is indistinguishable from success on every dashboard you have: requests succeed, error rate stays flat, latency may even improve. You can serve unranked results for days. Emit a fallback counter split by cause, alert on its rate rather than on hard errors, and sample a quality signal such as click-through on reranked versus fallback traffic so the cost of degradation is measurable.
saying these in an interview costs you the question
- Propagates a rerank failure as a request error
- Retries with a fixed policy regardless of remaining deadline
- Retries hard on 429 and deepens the rate-limit episode
- Sends full documents with metadata the model does not need
- Falls back silently with no metric or alert