An Istio VirtualService for a slow service sets `timeout: 10s` together with `retries: {attempts: 3, perTryTimeout: 5s}`. Explain how those two settings interact, and why a chain of services each configured this way can turn a slowdown into an outage.
answer
- two clocks, one budget
- attempts times per-try versus the total
- retries you never configured
- multiplicative down the chain
- budget shrinks as you go inward
basics
~20 sThe VirtualService timeout is the total budget for the whole request including every retry, while perTryTimeout caps one attempt. Three attempts of five seconds cannot all run inside a ten-second budget, and retries at every hop multiply, so a slow dependency becomes many times its normal load.
solid answer
~50 s`timeout` bounds the **entire** request as seen by the caller — all attempts, plus the gaps between them — while `perTryTimeout` bounds a single attempt. With `attempts: 3` and `perTryTimeout: 5s` the arithmetic wants up to fifteen seconds, but the ten-second overall timeout wins and cuts the request mid-retry, so the third attempt often never runs and you have paid for load you got no benefit from. The dangerous part is that retries **multiply along a call chain**: if three hops each retry twice on failure, one user request can become many requests at the deepest service, and that amplification arrives exactly when the deepest service is already struggling. Istio also applies default retries when you write no `retries` block at all, so services retry without anyone deciding to. The discipline is a decreasing timeout budget going inward, `retryOn` restricted to conditions that are actually safe to repeat, and `attempts: 0` on the paths that must not repeat.
code
yaml · 25 linesapiVersion: networking.istio.io/v1beta1
kind: VirtualService
metadata:
name: payments
spec:
hosts:
- payments
http:
- match:
- method:
exact: POST
uri:
prefix: /charge
route:
- destination: {host: payments, subset: v1}
timeout: 8s
retries:
attempts: 0
- route:
- destination: {host: payments, subset: v1}
timeout: 10s
retries:
attempts: 3
perTryTimeout: 3s
retryOn: connect-failure,refused-streamgo deeper
Know that timeout covers the whole request while perTryTimeout covers one attempt, and that both are set on an http route entry in the VirtualService.
Do the arithmetic out loud — attempts times perTryTimeout against the overall timeout — and explain why the overall budget truncates the retry sequence rather than extending it.
Show the amplification across a call chain, connect it to the fact that retries fire exactly when the dependency is already struggling, and prescribe a decreasing timeout budget plus a restricted retryOn list.
Set the policy: which endpoints may be retried at all, who owns the per-hop budget so it decreases inward, and why connection-pool limits in the DestinationRule are the required companion to any retry policy.
## Two different clocks The VirtualService exposes both settings on the same HTTP route entry, and they measure different things: ```yaml - route: - destination: {host: payments, subset: v1} timeout: 10s retries: attempts: 3 perTryTimeout: 5s retryOn: connect-failure,refused-stream,unavailable ``` - **`timeout`** is the deadline for the request as the caller experiences it: the first attempt, every retry, and everything in between. When it fires, the caller gets a failure regardless of what any attempt was doing. - **`perTryTimeout`** is the deadline for one attempt. When it fires, that attempt is abandoned and — if attempts remain and the outcome is retryable — another is started. The overall timeout always wins. `3 × 5s` wants fifteen seconds; the ten-second budget truncates it. The practical consequence is that the configuration promises three attempts and delivers roughly two, while the upstream has been made to do work for a request whose answer will be discarded. Whenever you set both, do the arithmetic: `attempts × perTryTimeout` should fit inside `timeout` with room to spare, or you should accept explicitly that the last attempt is decorative. ## Retries you did not ask for Istio applies a **default retry policy** to HTTP routes even when the VirtualService contains no `retries` block — a small number of attempts on connection-level failures and a limited set of retryable conditions. This surprises people: a team reasons that they never enabled retries, so a doubled request count cannot be theirs. It can be. Make the policy explicit on any route where the behaviour matters, and use `attempts: 0` to switch it off rather than relying on omission. ```yaml retries: attempts: 0 # explicitly no retries for a non-idempotent path ``` ## Why chains amplify Retry amplification is multiplicative, not additive. Take three hops, each retrying twice after the initial attempt — three total attempts per hop: ``` edge → A : up to 3 requests to A A → B : each of those, up to 3 → 9 requests to B B → C : each of those, up to 3 → 27 requests to C ``` One user request can become twenty-seven at the deepest service. And the amplification is **triggered by** the deepest service being slow or erroring, so the extra load arrives precisely when it is least survivable. This is the classic retry storm: a brief blip at the bottom of the stack is converted by the middle tiers into sustained overload, and the system does not recover on its own even after the original cause clears, because the retries themselves now supply the load. A layered timeout budget is the structural defence. Each hop's `timeout` should be meaningfully **shorter** than its caller's, so an inner call is abandoned before the outer deadline expires. When the outer timeout is shorter than the inner one, the caller gives up while the inner attempt is still running, the work continues to consume capacity, and the retry it triggers piles on top. ## Choosing retryOn honestly `retryOn` takes a comma-separated list of conditions — `connect-failure`, `refused-stream`, `unavailable`, `gateway-error`, `retriable-status-codes`, `5xx` among them. The safe end of that list is the conditions where the request demonstrably **never reached** the application: a connection that failed to establish, or a stream the server refused. Retrying those risks nothing. `5xx` is at the other end: the request may well have been processed and had its side effect before the error was produced, so retrying can duplicate work. Retry-safety is a property of the endpoint, so the setting belongs per route, not mesh-wide. ## Where the neighbouring controls live The `retries` block has no budget or ceiling of its own. What limits concurrent pressure on an upstream is the **DestinationRule's** `trafficPolicy.connectionPool` — `tcp.maxConnections`, `http.http1MaxPendingRequests`, `http.http2MaxRequests`, `http.maxRequestsPerConnection` — and its `outlierDetection`, which takes a persistently failing endpoint out of rotation. A complete answer places the retry policy in the VirtualService and the pressure limits in the DestinationRule, and notes that retries without those limits are a load multiplier with no governor. ## What to say in an interview Start with the two clocks and the arithmetic, name the default retries as something you have to switch off deliberately, then move straight to amplification along the chain and the decreasing-budget rule. Finish on retry safety: attempts are only free when the request never landed.
- How do you make sure a non-idempotent endpoint is never retried by the mesh?Give it its own route entry matched on method and path, and set `retries: {attempts: 0}` there. Omitting the block is not enough, because a default retry policy still applies. Keeping it as a separate, earlier entry in the http list also documents the intent where the next person will see it.
- Which resource limits how much concurrent pressure the proxy will put on an upstream, given that retries have no budget field?The DestinationRule's `trafficPolicy.connectionPool` — `tcp.maxConnections`, `http.http1MaxPendingRequests`, `http.http2MaxRequests`, `http.maxRequestsPerConnection` — with `outlierDetection` removing persistently failing endpoints. Retries live in the VirtualService, the governor lives in the DestinationRule, and configuring one without the other is how storms get through.
- Why must each hop's timeout be shorter than its caller's, rather than the same?If the caller's deadline expires first, it abandons the request while the inner attempt keeps running and consuming capacity, then usually retries on top of that. A strictly decreasing budget going inward means the inner call fails and reports before the outer deadline, so the outer hop can act on a real answer instead of guessing.
- With attempts: 3 and perTryTimeout: 5s under a 10s timeout, how many attempts actually run?Usually two. The first attempt burns up to five seconds, the second up to five more, and the overall ten-second budget fires before a third can complete — so the third attempt is largely decorative while still being possible to start and waste upstream work. Size perTryTimeout so attempts × perTryTimeout fits inside timeout.
saying these in an interview costs you the question
- Thinks timeout applies to each attempt separately
- Assumes no retries happen unless a retries block is written
- Treats retry amplification as additive across hops
- Retries on 5xx for a request that may have had side effects
- Sets the same timeout at every hop in the chain