Hosted rerank API or self-hosted cross-encoder — how do you decide between them?
answer
- billed per document, not per query
- depth multiplies both cost and latency
- a network hop you cannot tune
- passages leave, not just the query
- pinning a checkpoint enables fine-tuning
basics
~20 sDecide on three axes: unit economics, since hosted reranking is billed per document scored and rerank depth multiplies every query; tail latency and control, since a network hop adds variance you cannot tune; and data exposure, because reranking sends your corpus passages out, not just the query.
solid answer
~50 sHosted rerank endpoints are the right default early: no GPU to operate, a competitive model maintained for you, and cost that stays trivial at low volume. They stop being obviously right on three fronts. **Economics** — you pay per document scored, so cost scales with queries multiplied by rerank depth, and a top-100 rerank at sustained traffic can exceed the generation bill. **Latency** — the extra network round trip plus the provider's own queueing lands in your p95, and you cannot batch, pin or prioritise it. **Data** — unlike an embedding call, reranking ships whole passages of your corpus to a third party on every query, which is a real constraint under residency or confidentiality obligations. Self-hosting buys control over all three and costs you a GPU pool, batching and autoscaling, version pinning, and an eval you must rerun on every model upgrade. The honest framing is that this is a point on a curve you revisit as volume grows, not a permanent architectural stance.
go deeper
Know that reranking can be a hosted API call or a model you run yourself, and that the hosted call sends both the query and the candidate passages over the network.
Explain the concrete tradeoffs: per-document billing that scales with rerank depth, an added network round trip inside the latency budget, and the operational work self-hosting adds.
Show you would instrument the decision — cost per thousand queries, rerank p95 and p99 separately, actual depth — and that you have a graceful fallback when the reranker is slow or unavailable.
Own it as a revisitable point on a cost-control-risk curve with explicit triggers, and connect it to model lifecycle: pinning a checkpoint is what makes fine-tuning and reproducible evaluation possible, at the price of owning upgrades.
## Why this is a genuine decision and not a default Rerankers are small by modern standards — hundreds of millions of parameters, not hundreds of billions — which makes self-hosting genuinely feasible in a way that self-hosting a frontier generator is not. At the same time, several providers offer rerank endpoints that are cheap, good, and one HTTP call away. So both options are live, and the choice turns on operating constraints rather than on capability. ## Axis one: unit economics Hosted rerank pricing is generally per *document scored*, not per query. That multiplier is the whole story. A system doing 20 queries per second with a rerank depth of 100 is scoring two thousand documents per second; the same system at depth 25 scores five hundred. Rerank depth therefore appears directly in the invoice, alongside its appearance in the latency budget — the two constraints push the same lever in the same direction, which is convenient. Self-hosting converts that variable cost into a fixed one: a GPU (or a CPU pool for small quantized rerankers at low QPS) that costs the same whether it is busy or idle. The crossover is a utilization question. Low, spiky traffic favours the hosted meter; steady high traffic favours owned capacity. Do the arithmetic with real projected depth, not with the depth you wish you needed. ## Axis two: latency and control Every hosted call adds a network round trip and the provider's own queueing to the critical path. In a pipeline where the whole interactive budget might be a few hundred milliseconds, that is not noise. More importantly, it is variance you cannot manage: you cannot increase the batch, cannot pin the model to reserved capacity, cannot deprioritise a background job in favour of an interactive one, and cannot see why a p99 spiked. You can only retry, and retrying a latency problem usually makes it worse. Self-hosting gives you those knobs and hands you their operational cost: batching policy, admission control, autoscaling that reacts fast enough to a traffic spike, and a capacity plan. It also gives you the ability to fail gracefully — dropping to first-stage order under load is a decision you own, not one you negotiate through a provider's error semantics. ## Axis three: what leaves your boundary This axis is often missed and often decisive. A first-stage embedding call sends the *query*. A rerank call sends the query **and the candidate passages** — that is, chunks of your corpus, on every request, at your rerank depth. For a healthcare, legal or internal-engineering corpus this can be the constraint that ends the discussion, regardless of the economics. Check what the provider retains, in which region processing happens, and whether a zero-retention or in-region arrangement is available. Self-hosting inside your own network sidesteps the question entirely. ## Axis four: model control and lifecycle Hosted means the provider maintains, improves and occasionally retires the model. That is mostly a benefit — you get upgrades free — but it is also a dependency: a model version can change under you, shifting the ordering your evaluation was tuned on, and endpoints do get deprecated. Self-hosting pins a checkpoint. Pinning is what makes results reproducible and what makes fine-tuning possible at all; a fine-tuned domain reranker is inherently self-hosted unless the provider supports custom checkpoints. The flip side is that you now own an ML lifecycle: track new checkpoints, re-run your labelled evaluation before adopting one, and keep a rollback path. ## How to actually decide A workable sequence: 1. **Start hosted.** Prove the reranker earns its place at all against an unreranked baseline. Most of the value of this stage is learning whether you need reranking, and at what depth. 2. **Instrument.** Track cost per thousand queries, rerank p95 and p99 separately from end-to-end, and the depth actually used. 3. **Re-decide at a trigger, not on a schedule.** Sensible triggers: rerank spend becomes a visible line item, the network hop becomes a measurable share of the tail, a compliance review flags corpus egress, or a domain evaluation says you need to fine-tune. 4. **When you move, keep the hosted path as a fallback** if data rules permit — a self-hosted pool with a hosted overflow gives you burst capacity without over-provisioning. The answer an interviewer is listening for is not a side. It is that you know which measurements would change your mind, and that you noticed reranking sends documents rather than just queries.
- Which is the stronger argument for self-hosting: cost or data residency?They dominate at different scales. Cost only becomes decisive at sustained volume, because rerank billing is per document scored and depth multiplies it; below that it is a rounding error. Data residency is binary and can rule out hosted reranking on day one, since every request ships corpus passages rather than just the query. In regulated domains residency usually decides first, and economics only arrives later.
- You want to fine-tune a reranker on your own domain data. How does that change the hosting decision?It largely settles it. A custom checkpoint has to run somewhere you control unless your provider explicitly supports hosting fine-tuned rerankers, so fine-tuning generally implies self-hosting plus the lifecycle that comes with it: version pinning, a labelled evaluation gate before each promotion, and a rollback path. Weigh that permanent operational cost against the measured domain gain before committing to the fine-tune.
- What would you keep measuring so you know when to revisit this choice?Rerank spend per thousand queries, rerank p95 and p99 isolated from end-to-end latency, the depth actually being sent, and provider error and timeout rates. Set explicit triggers rather than a review cadence: spend crossing a threshold, the network hop becoming a visible share of the tail, or a compliance review flagging corpus egress. Each trigger names a specific number that flips the decision.
saying these in an interview costs you the question
- Assumes only the query is sent, not the candidate passages
- Compares per-query pricing while ignoring that billing is per document scored
- Treats self-hosting as free once the GPU exists, ignoring batching and autoscaling
- Ignores that a hosted model version can change and shift evaluated ordering
- Argues one option is universally correct regardless of volume or domain