When do you add a Pinecone reranking stage rather than keep tuning hybrid alpha?
answer
- ask which metric is actually broken
- one lever moves recall, one moves order
- a reordering stage sees only what arrived
- the second stage costs per query
- over-fetch first, then cut
basics
~20 sTune alpha while the right documents are already in the top-k but ordered badly by a linear blend. Add reranking when ordering needs the query and document read together — it costs latency and money per query, and it can only reorder what retrieval already returned.
solid answer
~50 sThey fix different problems. Alpha is a single global scalar blending two independent scores; it is free, applies at retrieval time, and is the right lever when your failure is "the lexical or semantic side is under-weighted". A reranker scores each query-document pair jointly with a model, so it can capture interactions no linear blend can express — negation, entity mismatch, the answer being in the third sentence — but it runs after retrieval, on a bounded candidate set. In Pinecone the pattern is: over-fetch with hybrid (say top_k=100), then call `pc.inference.rerank(...)` with a hosted model to cut to top_n. The costs are real: an extra network hop, tens to hundreds of milliseconds, and a per-request charge that scales with candidates. And the hard constraint is that reranking cannot recover a document retrieval missed — recall is owned by alpha, top_k and filtering; precision at the top is what reranking buys.
code
python · 20 linesfrom pinecone import Pinecone
pc = Pinecone(api_key="YOUR_KEY")
index = pc.Index("docs-hybrid")
res = index.query(
vector=[0.02] * 768,
sparse_vector={"indices": [7, 20514], "values": [0.3, 0.15]},
top_k=100,
include_metadata=True,
)
docs = [{"id": m["id"], "text": m["metadata"]["chunk"]} for m in res["matches"]]
ranked = pc.inference.rerank(
model="bge-reranker-v2-m3",
query="how do I rotate an api key",
documents=docs,
top_n=5,
return_documents=True,
)go deeper
Know that reranking is a second pass over results retrieval already returned, and that it cannot add documents the first stage missed.
Explain the two-stage shape — over-fetch with a hybrid query, then rerank to a smaller top_n — and why a linear alpha blend cannot express query-document interactions.
Diagnose with numbers: separate recall@k from top-of-list precision, decide which lever the failure calls for, and design the timeout and fallback for the extra dependency.
Own the tradeoff explicitly — quantify the relevance gain against added p95 latency and per-query cost at your traffic level, and set the policy for when the second stage is worth paying for and when chunking is the real fix.
## Two levers, two different failure modes Diagnose before you choose. Take a labelled query set and measure two things separately: **recall@100** (is the right chunk anywhere in the candidate set?) and **precision or NDCG@5** (is it near the top?). - Low recall@100 → the right document is not being retrieved at all. Reranking cannot help; it only reorders what it is given. Fix this with alpha, with a larger top_k, with better chunking, or by loosening a metadata filter. - High recall@100 but poor NDCG@5 → retrieval finds the document and ranks it 40th. This is exactly what a reranker fixes, and no value of alpha will, because the ordering error comes from interactions a sum of two independent similarity scores cannot represent. Stating this decomposition is the answer an interviewer is listening for. ## Why alpha has a ceiling The hybrid score is alpha·(dense similarity) + (1 - alpha)·(lexical similarity). Both terms are computed *without either model ever seeing the query and the document together*: the embedding of the chunk was fixed at ingest, the query embedding is fixed at query time, and the comparison is a dot product. That bi-encoder structure is what makes it fast enough to search millions of vectors — and it is also why it cannot tell that a document containing every query term is nonetheless about the opposite case, or that the matched entity is a different product with a similar name. One scalar cannot buy an interaction the representation never encoded. ## What Pinecone's reranking gives you Pinecone hosts reranking models behind `pc.inference.rerank(...)`, taking a `model`, the `query`, a list of `documents`, and `top_n` for how many to return, with `return_documents` controlling whether text comes back. Because it is hosted, you avoid running a GPU-backed cross-encoder yourself, which is the practical reason many teams adopt it. Indexes configured with integrated embedding can also request reranking as part of a `search` call, keeping it to one round trip. The shape is always two-stage: retrieve broadly and cheaply, then rescore narrowly and expensively. ## The costs to budget - **Latency.** A rerank call is an extra hop and a model forward pass over every candidate. Reranking 100 candidates costs meaningfully more than reranking 25, and the relationship is roughly linear in candidate count and in document length. For an interactive assistant this often lands in the same order as the LLM's own first-token latency, so it is not free perceptually. - **Money.** It is a per-request, per-candidate charge on top of the query. At high QPS this can exceed the vector search cost itself. - **Operational surface.** A second dependency that can rate-limit, time out or degrade. It needs a fallback: on failure, serve the hybrid ordering rather than erroring, since a slightly worse ranking beats no answer. ## How to size the candidate set Sweep top_k for the retrieval stage against recall@k on your labelled set and pick the smallest k where the curve flattens — typically somewhere between 50 and 150 for RAG over documentation. Then rerank down to the number of chunks the prompt can actually carry, often 3 to 10. Over-fetching beyond the recall plateau buys nothing and costs linearly. ## The decision, stated as policy 1. Always tune alpha first; it is free and it moves recall, which is the constraint reranking cannot relax. 2. Add reranking when measurements show recall is adequate and top-of-list precision is not, and when the downstream consumer is sensitive to order — an LLM given 5 chunks is very sensitive; a search page showing 20 results is less so. 3. Justify it with a number: the NDCG@5 or answer-accuracy delta against the added p95 latency and the per-query cost. If the delta is small and the traffic is high, keep the money. 4. Keep both levers under continuous evaluation, because corpus and traffic drift will move the optimum for each independently. ## A caution Reranking can also paper over a chunking problem. If the right answer is split across two chunks so that neither is individually convincing, no reranker fixes it — that is an ingestion-design issue, and reaching for a reranker will produce a small, expensive, disappointing improvement.
- Your recall@100 is 62 percent. Would a reranker help?No. Nearly four in ten queries never see the right chunk in the candidate set, and a reranker only reorders what retrieval returned — its ceiling is that same 62 percent. Spend the effort on retrieval instead: sweep alpha, raise top_k, revisit chunk size and overlap, and check whether a metadata filter is excluding valid candidates. Reconsider reranking once recall is comfortably high.
- How do you size the candidate set you send to the reranker?Sweep retrieval top_k against recall@k on a labelled set and pick the smallest k where the curve plateaus — commonly 50 to 150 for documentation RAG. Beyond the plateau you pay linearly in rerank latency and cost for no recall gain. Then rerank down to the number of chunks the prompt can actually carry, usually 3 to 10.
- What is your fallback when the reranking call times out?Serve the hybrid ordering as-is and log the degradation. Reranking is a precision improvement on top of a ranking that is already usable, so failing open costs a little relevance while failing closed costs the whole answer. Set a tight timeout, cap the candidate count so the call stays bounded, and alert on the fallback rate rather than on individual failures.
saying these in an interview costs you the question
- Reaches for reranking when recall, not ordering, is broken
- Assumes a reranker can surface documents retrieval missed
- Ignores the per-query latency and cost of the extra stage
- Reranks the full candidate set without sizing top_k
- Treats reranking as a substitute for fixing chunking