skip to content

Describe request hedging — sending a second copy of an in-flight request before the first one has failed. When does it improve tail latency, and what does it cost?

level: seniorimportance: nice to knowfreq 28%

answer

  1. fire second copy at p95, first answer wins, cancel loser
  2. independent slowness -> tail becomes p^2
  3. ~5% extra load if triggered at p95
  4. useless or harmful under global saturation
  5. needs idempotency, real cancellation, a hedge budget, shared deadline

basics

~20 s

If a request is still pending at roughly the p95 latency, send a duplicate to another replica, take whichever answers first, and cancel the loser. It cuts tail latency when slowness is per-request bad luck, and costs about 5% extra load plus a strict idempotency requirement.

solid answer

~60 s

Hedging attacks the tail rather than failures. Start the request; if no response by a trigger point — typically the p95 or p99 latency — issue a second attempt to a **different** replica while the first is still running. First usable response wins; cancel the other immediately. Both attempts share the original deadline. It works when tail latency comes from **per-request bad luck**: an unlucky slow node, a pause, a cold cache, a queue you happened to land behind. A second try is independent of that bad luck, so the combined tail is close to the product of the individual tail probabilities. It does not work — and actively harms — when the system is globally saturated: hedges then add load to the exact bottleneck causing the slowness, a positive feedback loop that can push a struggling system into collapse. Costs and requirements: extra load bounded by the trigger percentile; strict idempotency, since both copies may execute; cancellation that really stops the loser; a hedge budget capping hedges as a small fraction of traffic; and disabling hedges above a utilization threshold.

go deeper

for a junior

Know the shape: if it is taking unusually long, send a second copy, use whichever replies first, and cancel the other.

for a middle

Explain the trigger percentile controlling extra load, the need for idempotency, and cancelling the loser.

for a senior

Analyse the independence assumption, the p^2 tail effect, failure under saturation, hedge budgets and cut-offs, and the distinction from retries.

for a principal

Decide when hedging is the right lever at all versus reducing the tail at the source, and specify the guardrails — global budget, utilization-based disabling, live percentile triggers, idempotency guarantees — with metastable-failure risk explicitly considered.

## The problem hedging solves In a fan-out system, user-visible latency is governed by the slowest component, not the average. A service whose median is 10 ms but whose p99 is 500 ms, called on 100 shards in parallel, will show a slow shard in almost every request. Reducing the *median* does nothing for that; you must attack the tail. Tail latency usually comes from per-request accidents — a garbage-collection or compaction pause on one node, a cold cache, a rebalancing replica, an unlucky queue position — which are largely independent between replicas. ## The mechanism ```text t=0 send request to replica A t=p95 still no answer -> send the same request to replica B take the first usable response; cancel the other attempt both attempts share the ORIGINAL deadline (no extension) ``` The trigger point is the key parameter. Set it at the p95 of the operation's own latency distribution and, by construction, only about 5% of requests ever hedge, so the extra load is roughly 5%. Set it at the p99 and you pay about 1% for a smaller improvement. Setting it near the median doubles your traffic and usually makes things worse. Why it works: if p(slow) is independent across replicas, the probability that *both* copies are slow is roughly p(slow)^2 — turning a 1-in-100 tail into something near 1-in-10,000, at a few percent extra cost. Independence is the load-bearing assumption, which is why the hedge must go to a *different* replica, ideally in a different failure domain, and why the client should exclude the replica already suspected to be slow. ## When it fails **Global saturation.** If requests are slow because the whole fleet is at capacity, the slowness is *not* independent — it is a shared cause. Hedging then adds 5% (or much more, since far more than 5% of requests now exceed the trigger) to the exact resource that is the bottleneck. The trigger fires more often, which adds more load, which makes it fire even more often: a runaway loop that turns a slowdown into an outage and can keep the system down after the original trigger clears. Countermeasures: an absolute **hedge budget** (hedges may not exceed a few percent of total requests over a rolling window), automatic disabling above a utilization or error-rate threshold, and a trigger derived from a *live* latency distribution rather than a static constant that goes stale. **Non-idempotent work.** Both copies may execute to completion; cancelling the loser is best-effort and racy. Hedging is therefore safe only for reads or for operations protected by an idempotency key with server-side deduplication. **Expensive work.** Duplicating a request that costs seconds of CPU or scans a large amount of data is a poor trade even at 5% frequency. Hedging suits short, cheap, latency-sensitive operations. **Weak cancellation.** If the loser cannot really be stopped — cancellation is ignored, or the peer does not honour it — you pay full duplicate cost on every hedged request, not the fraction you budgeted. ## Hedging versus retrying A retry starts *after* a failure or timeout; a hedge starts *while the first attempt is still alive*. Retries recover from errors; hedges cut latency. They compose: within one shared deadline you may hedge once and still retry an outright error if budget remains. Both require idempotency; both must be globally budgeted; and neither may extend the deadline — the original deadline bounds all copies, which also means a hedge issued very close to the deadline is pointless and should be suppressed. ## Related tactics *Tied requests* send both copies immediately but have the replicas coordinate so that whichever starts executing cancels the other's queued copy — cheaper than true duplication because only queueing delay is duplicated. *Backup requests with cross-server cancellation* is the same idea. On the server side, the complementary fix is to make the tail smaller: shed load early on expired deadlines, avoid long stop-the-world pauses, and prevent head-of-line blocking. ## Operating it Measure the hedge rate, the win rate (how often the hedge actually answers first), and the extra load. A hedge rate far above the trigger percentile means the distribution has shifted and you are amplifying an overload. A win rate near zero means the slowness is not independent — turn hedging off, because you are paying for nothing.

  • How do you pick the hedge trigger point?
    Use the operation's own live latency distribution and trigger at a high percentile, typically p95 or p99, because that directly bounds the fraction of requests that hedge and therefore the extra load. Compute it from a rolling window rather than hard-coding a constant, since a static trigger becomes an amplifier when the distribution shifts under load. Also suppress hedges when the remaining deadline is too small for a second attempt to plausibly win.
  • Why can hedging make an overloaded system worse?
    Because it assumes slowness is independent per request. Under saturation the cause is shared, so far more than the intended fraction of requests cross the trigger, each adds a duplicate to the bottleneck, latency rises further, and even more requests cross the trigger. That positive feedback can sustain an outage after the original trigger has passed, which is why hedging needs a global budget and an automatic cut-off tied to utilization or error rate.

It is like queueing at two checkout lanes when your first one stalls: brilliant when one cashier is unlucky, useless when the whole store is packed — and then you are also the person making both lines longer.

saying these in an interview costs you the question

  • Hedging non-idempotent operations because "the loser gets cancelled" — cancellation is best-effort and races with execution.
  • Setting the trigger near the median, roughly doubling traffic.
  • Giving the hedge its own fresh deadline instead of sharing the original one.
  • Leaving hedging enabled during overload, where it amplifies the bottleneck it is reacting to.
  • Confusing hedging with retrying — a hedge runs while the first attempt is still in flight.
  • Sending the hedge to the same replica that is already slow.

context