skip to content

A Linkerd ServiceProfile lets you mark a route `isRetryable` and configure a `retryBudget` rather than a fixed number of attempts. What does a retry budget actually enforce, and why is it preferred to "retry three times"?

level: middleimportance: nice to knowfreq 40%

answer

  1. per-request count versus aggregate load
  2. a ratio of ordinary traffic, not a multiplier
  3. a floor so quiet services can still retry
  4. unmatched routes fall into the default bucket
  5. the mesh cannot know what is idempotent

basics

~20 s

A retry budget caps retries as a proportion of ordinary requests — Linkerd's default allows roughly 20% extra plus a small floor — so a struggling dependency cannot be hit with a multiple of its normal load the way a fixed per-request retry count allows.

solid answer

~40 s

A fixed count is per-request and therefore unbounded in aggregate: with three attempts, a dependency failing everything receives three times its normal traffic exactly when it is least able to serve it. Linkerd's `retryBudget` is a rate limit on retries instead, expressed as `retryRatio` (extra retries as a fraction of ordinary requests, default 0.2), `minRetriesPerSecond` (a floor so low-traffic services can still retry, default 10) and `ttl` (the sliding window over which the ratio is computed, default 10s). Retries only happen on routes you have explicitly marked `isRetryable: true` in a ServiceProfile, and a request that matches no defined route falls into the default bucket and is never retried. That opt-in default is deliberate: Linkerd will not replay a request unless you have asserted it is safe to replay.

code

yaml · 22 lines
yaml
apiVersion: linkerd.io/v1alpha2
kind: ServiceProfile
metadata:
  name: payments.prod.svc.cluster.local
  namespace: prod
spec:
  routes:
    - name: GET /v1/quote
      condition:
        method: GET
        pathRegex: /v1/quote
      isRetryable: true
      timeout: 300ms
    - name: POST /v1/charge
      condition:
        method: POST
        pathRegex: /v1/charge
      timeout: 2s
  retryBudget:
    retryRatio: 0.2
    minRetriesPerSecond: 10
    ttl: 10s

go deeper

for a junior

Know that Linkerd can retry failed requests for you, that it does so only on routes you explicitly mark retryable, and that retrying a non-idempotent operation can duplicate its effect.

for a middle

Explain the three budget fields and what each prevents: the ratio caps amplification, the floor keeps low-traffic services able to retry, and the ttl sets the window. Say why unmatched traffic is never retried.

for a senior

Show the failure mode you are avoiding: fixed counts multiplying load during an outage and stacking across tiers. Tie retries to the timeout budget down the call chain and to what the caller does when the budget is spent.

for a principal

Own where retry policy should live at all — mesh, client library, or the caller's explicit fallback — and how you keep retry amplification bounded across a multi-tier system rather than service by service.

## The problem with a retry count Retries are per-request; overload is aggregate. That mismatch is the whole story. Suppose a service normally handles 1,000 RPS from a client configured with "up to 3 attempts". Under healthy conditions almost nothing retries and the extra load is negligible. Now the dependency starts failing — a slow database, a bad deploy, a saturated thread pool. Every request fails, every request is retried twice more, and the dependency now receives 3,000 RPS at exactly the moment it had trouble with 1,000. The retry policy, which existed to improve availability, has become the mechanism that prevents recovery. Add a second retrying tier upstream and the multiplication compounds. ## What a budget enforces instead A retry budget converts "how many times may this request be retried" into "how much *total* retry traffic may this client generate". Linkerd's ServiceProfile expresses it with three fields: ```yaml apiVersion: linkerd.io/v1alpha2 kind: ServiceProfile metadata: name: payments.prod.svc.cluster.local namespace: prod spec: routes: - name: GET /v1/quote condition: method: GET pathRegex: /v1/quote isRetryable: true timeout: 300ms retryBudget: retryRatio: 0.2 minRetriesPerSecond: 10 ttl: 10s ``` - **`retryRatio`** — retries permitted as a fraction of ordinary (non-retry) requests. The default 0.2 means at most 20% additional traffic, so a total failure adds a fifth of the load rather than multiplying it. - **`minRetriesPerSecond`** — a floor, default 10. Without it a service receiving one request per second could never retry at all, since 20% of one request rounds to nothing. - **`ttl`** — the sliding window over which the ratio is computed, default 10 seconds. It decides how quickly the budget forgets a burst. When the budget is exhausted, the failure is simply returned to the caller rather than retried. That is the intended behaviour: under sustained failure the mesh stops amplifying and lets the error surface, where a circuit breaker or the caller's own degradation logic can act on it. ## Why retries are opt-in per route Linkerd does not retry anything by default. Two conditions must hold: 1. A ServiceProfile exists for the target service, named for its fully-qualified service name (`<service>.<namespace>.svc.cluster.local`), and defines the route via `condition` (method and `pathRegex`). 2. That route sets `isRetryable: true`. Requests that match no defined route fall into a default bucket — visible as `[DEFAULT]` in `linkerd viz routes` — and are never retried. This catches people out: they configure a budget, see no retries, and have not realised their traffic never matched a route condition. The opt-in design is a safety property, not an inconvenience. Replaying a request is only safe if the operation is idempotent, and the mesh cannot know that. Marking `POST /v1/charge` retryable when the handler is not idempotent produces duplicate charges, and no budget saves you from that — the budget limits volume, not correctness. ## The mechanical constraint To retry a request the proxy must be able to replay it, which means holding the request body. Linkerd will not retry requests whose bodies it could not buffer, so large-payload or streaming requests are effectively non-retryable regardless of configuration. Expect a follow-up on this: the interviewer is checking whether you understand that a retry is a *re-send*, not a magic re-run. ## Timeouts belong in the same conversation The same route carries a `timeout`. Retries and timeouts interact: a per-route timeout that is longer than the caller's own deadline means the retries never happen within the caller's patience, and the caller times out while the mesh is still working. Budget your timeouts down the call chain — each hop must allow less time than the hop above it, with room for the retries you have authorised. ## Version note ServiceProfile has been the long-standing mechanism for per-route retries, timeouts and per-route metrics. More recent Linkerd releases also support expressing retry and timeout policy via route resources with Linkerd-specific annotations; if you are configuring this today, check which mechanism your release documents as current, and do not mix two sources of truth for the same route.

  • You added a retryBudget but see no retries at all in `linkerd viz routes`. What do you check first?
    Whether the traffic actually matches a defined route. A ServiceProfile must be named for the fully-qualified service name, its route `condition` must match the real method and path, and the route must set `isRetryable: true`. Traffic matching nothing lands in the `[DEFAULT]` bucket, which is never retried no matter what budget you configured.
  • Would you mark a POST route retryable?
    Only if the handler is genuinely idempotent — typically because it accepts an idempotency key and deduplicates. The mesh has no way to know whether replaying a POST creates a second charge or order, so `isRetryable` on a write route is an assertion you are making about the application, and a wrong one produces duplicates that the budget cannot prevent.
  • What happens to a request when the retry budget is already exhausted?
    It is not retried; the failure is returned to the caller immediately. That is the intended behaviour under sustained failure — the mesh stops amplifying load onto a struggling dependency and lets the error surface where the caller's own fallback or circuit-breaking logic can respond.

saying these in an interview costs you the question

  • Treats a retry budget as just another maximum attempt count
  • Assumes Linkerd retries failed requests by default
  • Marks write routes retryable without idempotency
  • Ignores that retries multiply load precisely during an outage
  • Thinks a budget makes non-idempotent retries safe

context