skip to content

An Envoy cluster's `circuit_breakers` block sets thresholds for max_connections, max_pending_requests, max_requests and max_retries. What does each one limit, and what does a client see when one is exceeded?

level: middleimportance: should knowfreq 50%

answer

  1. concurrency caps, no half-open state
  2. pending overflow is the fast 503
  3. HTTP/2 binds on max_requests
  4. retries capped concurrently, not per request
  5. per proxy, multiplied by fleet size

basics

~20 s

They are concurrency caps, not a tripping breaker: max_connections limits upstream connections, max_pending_requests limits requests queued waiting for one, max_requests limits in-flight requests, max_retries limits concurrent retries. Exceeding a limit yields an immediate 503 flagged UO.

solid answer

~50 s

Envoy's `circuit_breakers` are per-cluster, per-priority concurrency limits rather than the open/half-open state machine the name suggests. `max_connections` caps how many upstream connections the pool may open, `max_pending_requests` caps how many requests may sit waiting for a connection to become available, `max_requests` caps requests in flight (for HTTP/2 that is concurrent streams on the pool), and `max_retries` caps how many retries can be running concurrently across the cluster. Defaults are 1024 for the first three and 3 for retries. When a request would exceed pending or in-flight limits, Envoy fails it immediately with 503 and response flag `UO`, incrementing `upstream_rq_pending_overflow` or `upstream_rq_overflow`; an exceeded retry limit increments `upstream_rq_retry_overflow` and the original error is returned. Crucially the limits are per Envoy process, so a hundred sidecars means a hundred times the configured concurrency arrives at the upstream.

code

yaml · 15 lines
yaml
clusters:
  - name: checkout
    connect_timeout: 0.25s
    type: EDS
    circuit_breakers:
      thresholds:
        - priority: DEFAULT
          max_connections: 200
          max_pending_requests: 50
          max_requests: 400
          max_retries: 5
          track_remaining: true
          retry_budget:
            budget_percent: { value: 20.0 }
            min_retry_concurrency: 3

go deeper

for a junior

Know that Envoy's circuit_breakers set limits on how many connections and requests a cluster may have outstanding, and that exceeding one causes an immediate 503 rather than waiting.

for a middle

Be able to walk through all four thresholds, say which binds for HTTP/1.1 versus HTTP/2, and name the overflow counter and response flag you would look for. Correct the interviewer's assumption that this is a state machine.

for a senior

Size the numbers from measured concurrency and explain why fast shedding beats deep queueing. Connect thresholds to retry policy and timeouts as one shared concurrency budget, and use the remaining-headroom gauges rather than only overflow counters.

for a principal

Decide where concurrency limiting belongs across the platform — per-client sidecar limits versus enforcement at the shared dependency — and set defaults teams inherit, given that per-proxy limits multiply by fleet size.

## Not the pattern you are picturing The name misleads. Envoy's `circuit_breakers` do not trip open, sit in a half-open probing state, and close again. They are **hard concurrency limits on the upstream connection pool**, evaluated per cluster and per routing priority (`DEFAULT` and `HIGH` are configured separately). The behaviour that most resembles a classic breaker in Envoy is `outlier_detection`, which ejects individual failing hosts. Being able to say that plainly is half the value of this question. ## The four thresholds ```yaml clusters: - name: checkout circuit_breakers: thresholds: - priority: DEFAULT max_connections: 200 max_pending_requests: 100 max_requests: 400 max_retries: 5 track_remaining: true ``` - **`max_connections`** — the maximum number of connections Envoy's pool will open to all hosts in the cluster. It is the binding limit for HTTP/1.1, where one connection carries one request at a time. Hitting it does not fail the request directly; the request waits in the pending queue and `upstream_cx_overflow` increments. - **`max_pending_requests`** — how many requests may be queued waiting for a connection. This is the one that turns saturation into fast failure: exceed it and Envoy immediately returns 503 with flag `UO`, incrementing `upstream_rq_pending_overflow`. For HTTP/1.1 clusters this is effectively your queue-depth control. - **`max_requests`** — the maximum number of requests in flight to the cluster at once. For HTTP/2 and gRPC clusters, where one connection multiplexes many streams, this rather than `max_connections` is the meaningful limit. Overflow yields 503 `UO` and increments `upstream_rq_overflow`. - **`max_retries`** — how many retries may be *in flight simultaneously* across the whole cluster; it is not a per-request attempt count (that is `num_retries` in the route's retry policy). Default 3, deliberately small, so a broad failure cannot double the load on an already-failing dependency. Overflow increments `upstream_rq_retry_overflow` and the original response is returned unretried. Modern configurations often replace it with `retry_budget`, expressing the cap as `budget_percent` of active requests with a `min_retry_concurrency` floor. Defaults are 1024 for connections, pending requests and requests, and 3 for retries — high enough that many deployments never notice the limits exist until a stall makes the pending queue overflow. ## What the client experiences Overflow is **fast failure, not queueing**. That is the design intent: shedding a request in microseconds keeps the proxy's own memory and file descriptors bounded and pushes back on the caller rather than growing an unbounded queue whose entries will time out anyway. From the caller's perspective a 503 arrives faster than any healthy response ever could, which is a useful signature: sub-millisecond 503s with flag `UO` mean the proxy shed load, not that the upstream failed. ## Observing them Setting `track_remaining: true` publishes gauges such as `cluster.<name>.circuit_breakers.default.remaining_rq` and `remaining_pending`, so a dashboard can show headroom instead of only recording the moment it ran out. The overflow counters (`upstream_cx_overflow`, `upstream_rq_pending_overflow`, `upstream_rq_overflow`, `upstream_rq_retry_overflow`) are the alerting signal. On the admin interface, `/clusters` shows per-cluster and per-host connection and request state alongside them. ## Sizing them The useful frame is Little's law: concurrency equals throughput times latency. A cluster serving 500 requests per second at 40ms average latency needs roughly 20 concurrent requests; a limit of 400 is generous headroom, while 10 would shed traffic under normal conditions. Set the limit above healthy peak concurrency but far below the point where the upstream collapses, so the proxy sheds before the backend does. Pending-request depth should be small — a deep queue only converts a fast failure into a slow one. ## The distributed-limit trap The thresholds are enforced **per Envoy process**. In a sidecar mesh, each of a hundred proxies independently permits `max_requests` concurrent requests, so the upstream can see a hundred times the number you wrote down. Limits configured as if they were global protection for a shared backend are a common and expensive mistake; the number has to be reasoned about as per-client concurrency multiplied by client count, or the protection has to live at the receiving end instead. ## Interaction with retries and timeouts Retries occupy both a request slot and a retry slot, so an aggressive route retry policy can push a cluster into overflow during a partial failure — the exact moment you least want extra load. Similarly, a long route timeout keeps slots occupied, so a slow upstream drains capacity into the pending queue until it overflows. Timeouts, retries and circuit-breaker thresholds are three views of the same concurrency budget and should be chosen together.

  • Why is max_requests, not max_connections, the meaningful limit for a gRPC cluster?
    Because HTTP/2 multiplexes many concurrent streams onto one connection. A gRPC cluster may serve hundreds of in-flight requests over a handful of connections, so a connection cap never binds while the request cap does. Size max_requests from expected concurrency — throughput times latency — and treat max_connections as a resource guard rather than a load control.
  • How does max_retries differ from num_retries in a route's retry policy?
    `num_retries` is how many extra attempts one request may make. `max_retries` is how many retries may be in flight across the entire cluster at once, defaulting to 3. It exists so a widespread failure cannot multiply load on a struggling dependency; when it is full, retries are skipped and `upstream_rq_retry_overflow` increments.
  • Your sidecars each allow 400 concurrent requests to a shared backend. What is wrong?
    The threshold is enforced per proxy, so with a hundred sidecars the backend can receive up to 40,000 concurrent requests. Per-client limits protect the client's own resources, not a shared server. Real protection for the dependency has to be enforced where it is shared — at the receiving side or a fronting gateway — or the per-client number must be divided by the expected client count.

saying these in an interview costs you the question

  • Expecting an open/half-open state machine like a library breaker
  • Believing overflow queues the request rather than failing it
  • Assuming the thresholds protect the upstream globally
  • Confusing max_retries with per-request num_retries
  • Setting a huge pending queue to avoid seeing 503s

context