An Envoy route is configured with `timeout: 3s` and a `retry_policy` of `num_retries: 3` with `per_try_timeout: 3s`, but slow upstreams are never retried in production. Why not, and how should the two timeout values relate?
answer
- overall timeout covers all attempts
- per-try must fit inside the budget
- upstream_rq_retry stuck at zero
- retry_on must match the real failure
- backoff also spends the budget
basics
~20 sEnvoy's route timeout covers the whole request including every retry, so a per_try_timeout equal to it means the first attempt consumes the entire budget and the request fails before a retry can start. Per-try must be a fraction of the overall timeout.
solid answer
~50 sIn Envoy the route-level `timeout` is an end-to-end budget for the request *including* all retry attempts, while `per_try_timeout` bounds each individual attempt. Setting both to 3s means attempt one runs out the overall budget at the same instant it hits its own per-try limit, so Envoy returns 504 with a `UT` flag and never gets to attempt two — `num_retries` is effectively dead configuration. The fix is to size per-try as a fraction of the total: with a 3s budget and one retry, roughly 1.4s per try leaves room for both attempts plus backoff, which starts at a 25ms base interval and grows exponentially. Also confirm `retry_on` actually covers the failure — a per-try timeout is retried under `5xx`, whereas `retry_on: connect-failure` alone would not cover it. And remember retries consume the cluster's `max_retries` circuit breaker or its `retry_budget`, so an exhausted budget silently skips retries too.
code
yaml · 17 linesroutes:
- match: { prefix: "/payments" }
route:
cluster: payments
timeout: 3s
retry_policy:
retry_on: "5xx,reset,connect-failure"
num_retries: 1
per_try_timeout: 1.2s
retry_back_off:
base_interval: 0.025s
max_interval: 0.25s
retry_host_predicate:
- name: envoy.retry_host_predicates.previous_hosts
typed_config:
"@type": type.googleapis.com/envoy.extensions.retry.host.previous_hosts.v3.PreviousHostsPredicate
host_selection_retry_max_attempts: 3go deeper
Know that Envoy has both an overall request timeout on the route and a per-attempt timeout inside the retry policy, and that the per-attempt one must be smaller for retries to be possible.
Explain the arithmetic: the overall budget must hold every attempt plus the backoff between them. Be able to name the retry_on conditions and say which one covers a per-try timeout.
Diagnose it from telemetry — a retry counter pinned at zero, an overflow counter climbing — and pick concrete values for a real budget. Discuss retrying onto a different host and when retrying a slow upstream makes the incident worse.
Own the policy across services: default attempt counts, whether retry budgets replace fixed limits, which layer of the call chain is allowed to retry at all, and how header-based overrides are governed at trust boundaries.
## Two different clocks Envoy exposes two timeouts on the same route, and they measure different things. - **`timeout`** (route-level, default 15 seconds) is the overall request timeout. It bounds the entire request as seen by the client, *including every retry attempt and the backoff between them*. Setting it to `0s` disables it. - **`per_try_timeout`** (inside `retry_policy`) bounds a **single attempt**. If it is not set, each attempt is bounded only by the global route timeout. Because the overall timeout is a hard ceiling on the whole sequence, `per_try_timeout >= timeout` makes the retry policy unreachable: the first attempt uses the full budget, the overall timeout fires, Envoy sends a 504 with response flag `UT`, and no further attempt is made. The configuration looks like it retries; the stat `cluster.<name>.upstream_rq_retry` stays at zero, which is how you confirm it. ```yaml route: cluster: payments timeout: 3s # covers ALL attempts retry_policy: retry_on: "5xx,reset,connect-failure" num_retries: 1 per_try_timeout: 1.2s # must be a fraction of 3s ``` ## Sizing the two numbers A workable rule: pick the overall budget from what the caller can actually wait for, decide how many attempts that budget can afford, then divide, leaving headroom for backoff. With a 3s budget and two total attempts, about 1.2–1.4s per try is right. More attempts inside the same budget means each attempt gets less time, which makes the retry more likely to be cut short than to succeed — that trade is the real design decision, not the value itself. Envoy's default retry backoff uses an exponential schedule with a base interval of 25ms and a maximum of 250ms; `retry_back_off` lets you set `base_interval` and `max_interval` explicitly. The backoff time is spent inside the overall budget too. ## retry_on must actually match the failure Retries only happen for conditions listed in `retry_on`. The common values are `5xx`, `gateway-error` (502/503/504), `reset`, `connect-failure`, `retriable-4xx`, `refused-stream`, `retriable-status-codes` and `envoy-ratelimited`. A per-try timeout is treated as a 5xx-class failure, so `retry_on: 5xx` covers it while `connect-failure` alone does not. A second common cause of "retries never happen" is simply that the failure mode in production is not in the `retry_on` list. ## The other silent skips Even with the timeouts sized correctly, Envoy may decline to retry: - **Concurrency limits.** Retries are gated by the cluster's `circuit_breakers` `max_retries` threshold (default 3 concurrent retries) or, if configured, `retry_budget` with `budget_percent` and `min_retry_concurrency`. Overflow increments `cluster.<name>.upstream_rq_retry_overflow` and the original failure is returned as-is. - **Body buffering.** A request whose body exceeds the buffer limit cannot be replayed, so it will not be retried. - **Response already started.** Once upstream response headers have been forwarded downstream, the attempt is committed. ## Per-request overrides Envoy honours request headers that override route settings when the route allows it: `x-envoy-retry-on`, `x-envoy-max-retries`, `x-envoy-upstream-rq-timeout-ms` and `x-envoy-upstream-rq-per-try-timeout-ms`. These are useful for a specific slow client path, and they are a hazard at the edge — a proxy facing untrusted clients should strip them rather than let callers dial up their own retry count against your backends. ## Where retries make things worse Retrying a *slow* upstream, as opposed to a failed one, adds load to something already struggling. Two guards matter. First, cap total attempts low — one retry is usually the right answer for an interactive path. Second, prefer a different host on the retry: `retry_host_predicate` with `envoy.retry_host_predicates.previous_hosts` plus `host_selection_retry_max_attempts` makes Envoy avoid the endpoint that just failed, which turns a retry into a real second chance rather than a second hit on the same sick host. ## Confirming the fix After changing values, watch three counters per cluster: `upstream_rq_retry` (attempts made), `upstream_rq_retry_success` (attempts that rescued a request) and `upstream_rq_retry_overflow` (attempts refused by the limit). A retry policy where success stays near zero is buying latency and load for nothing and should be turned down or off.
- How would you confirm from telemetry that retries are firing at all?Watch `cluster.<name>.upstream_rq_retry`, `upstream_rq_retry_success` and `upstream_rq_retry_overflow` on the admin stats endpoint. Zero attempts means the policy is unreachable — usually a per-try timeout at or above the route timeout, or a `retry_on` list that does not match the actual failure. Attempts with near-zero successes means retries are costing load without rescuing requests.
- What stops the retry from hitting the same failing host again?Configure `retry_host_predicate` with `envoy.retry_host_predicates.previous_hosts` and set `host_selection_retry_max_attempts` so Envoy re-runs host selection a few times to avoid endpoints already tried for this request. Without it, load-balancer selection can legitimately return the same sick host and the retry buys nothing.
- Why should an edge Envoy strip x-envoy-max-retries from inbound requests?Because those headers let a caller override the route's retry policy. An untrusted client could ask for a large retry count and multiply its own load against your backends, turning one request into many. Header-based overrides are appropriate between trusted internal hops; at the edge, strip them and let route configuration decide.
saying these in an interview costs you the question
- Thinking the route timeout applies per attempt
- Setting per_try_timeout equal to the overall timeout
- Raising num_retries when the budget cannot fit another try
- Ignoring that backoff also consumes the request budget
- Assuming any failure is retried regardless of retry_on