skip to content

After enabling `outlier_detection` on an Envoy cluster, a short burst of upstream errors ejects most of the pool and latency gets worse instead of better. Which fields drive that behaviour, and what does Envoy do once too many hosts are ejected?

level: seniorimportance: should knowfreq 42%

answer

  1. passive health check, not active probes
  2. cluster-wide burst hits every host at once
  3. cap how much of the pool can leave
  4. below half healthy, Envoy stops trusting health
  5. app 500s are not host failures

basics

~20 s

Aggressive consecutive_5xx with a large max_ejection_percent ejects healthy hosts during an error burst, concentrating load on survivors. Once healthy hosts fall below healthy_panic_threshold (default 50%), Envoy enters panic mode and load balances across all hosts, ignoring health.

solid answer

~50 s

Envoy's outlier detection is passive health checking: every `interval` it reviews per-host results and ejects any host that hit `consecutive_5xx` (default 5) for `base_ejection_time` (default 30s), with each successive ejection lasting proportionally longer up to `max_ejection_time`. Two settings decide whether a burst becomes an outage. `max_ejection_percent` caps how much of the pool can be ejected — leave it at a high value and a burst that hits every host ejects nearly all of them, so the remaining hosts absorb the whole load and slow down. Then `common_lb_config.healthy_panic_threshold`, default 50%, kicks in: when fewer than half the hosts are healthy Envoy assumes its health data is wrong and balances across every host regardless, which you see as the `lb_healthy_panic` counter rising. The other trap is that application 500s count toward `consecutive_5xx`; use `consecutive_gateway_failure` when you only want connection-level failures to eject.

code

yaml · 17 lines
yaml
clusters:
  - name: orders
    connect_timeout: 0.25s
    type: EDS
    outlier_detection:
      interval: 10s
      base_ejection_time: 30s
      max_ejection_time: 300s
      max_ejection_percent: 20
      consecutive_5xx: 5
      enforcing_consecutive_5xx: 0
      consecutive_gateway_failure: 3
      enforcing_consecutive_gateway_failure: 100
      split_external_local_origin_errors: true
    common_lb_config:
      healthy_panic_threshold:
        value: 50.0

go deeper

for a junior

Know that Envoy can remove a misbehaving backend from rotation based on the errors real requests see, and that this is configured on the cluster rather than as a separate probe.

for a middle

Explain the detectors and their timers — consecutive errors, the sweep interval, ejection duration and the ejection cap — and say where you would see an ejected host on the admin interface.

for a senior

Diagnose the cluster-wide burst case: why per-host detection fires everywhere at once, why the cap matters, what panic mode does, and which counters prove it. Distinguish application errors from host failures when picking a detector.

for a principal

Set the platform default for passive ejection and decide where it belongs at all — versus active health checking or capacity work — knowing that ejection helps only when the rest of the pool can absorb the load.

## What outlier detection actually is Outlier detection is Envoy's **passive** health check: rather than probing endpoints on a schedule, it watches the results of real traffic and temporarily removes hosts that look broken relative to the rest. It runs on a sweep every `interval` (default 10 seconds) and marks ejected hosts with the health flag `/failed_outlier_check`, visible per host on the admin `/clusters` page. ```yaml clusters: - name: orders outlier_detection: consecutive_5xx: 5 interval: 10s base_ejection_time: 30s max_ejection_percent: 20 enforcing_consecutive_5xx: 100 consecutive_gateway_failure: 3 enforcing_consecutive_gateway_failure: 100 max_ejection_time: 300s common_lb_config: healthy_panic_threshold: value: 50.0 ``` ## The detectors - **`consecutive_5xx`** (default 5) — N consecutive 5xx-class results from one host trigger ejection. Enforced with probability `enforcing_consecutive_5xx`, default 100. - **`consecutive_gateway_failure`** — the narrower variant, counting only 502/503/504 and connection-level failures. Its enforcing percentage defaults to **0**, meaning it is measured but not acted on until you turn it on. - **Success-rate detection** — compares each host against the cluster mean and ejects statistical outliers, gated by `success_rate_minimum_hosts` (default 5) and `success_rate_request_volume` (default 100) so it stays quiet on small or idle clusters. - **Failure-percentage detection** — an absolute-threshold alternative to the statistical one. - **`split_external_local_origin_errors`** — separates errors Envoy itself produced (connect failures, local resets) from real upstream responses, so you can eject on transport failures without counting application errors. ## Why a burst becomes an outage The failure mode is mechanical. A dependency of your dependency hiccups, so *every* host in the cluster returns 5xx for a few seconds. Consecutive-error detection is per host, but the condition is cluster-wide, so all hosts qualify at once. If `max_ejection_percent` is generous, most of the pool disappears from load balancing simultaneously. The survivors now receive all the traffic, saturate, and start failing or slowing — and each ejection lasts `base_ejection_time` multiplied by the number of times that host has been ejected, so recovery is not immediate. A transient upstream blip has been converted into a self-inflicted capacity incident. The defence is `max_ejection_percent`. It exists precisely because ejection is only meaningful when the *rest* of the pool is healthy; ejecting a majority is a statement that the whole cluster is broken, which is not something removing hosts can fix. Values in the 10–20% range keep ejection as a scalpel. ## Panic mode Envoy has a second guard. `common_lb_config.healthy_panic_threshold` (default 50%) says: if fewer than this share of hosts are healthy, distrust the health data and load balance across **all** hosts, healthy or not. The reasoning is that a cluster reporting 80% unhealthy is more often a broken health signal than 80% genuinely dead machines, and sending traffic to possibly-bad hosts beats sending it all to a handful of survivors. Operationally this matters because it explains a confusing symptom: you can see hosts ejected in `/clusters` *and* still see traffic going to them. The counter `cluster.<name>.lb_healthy_panic` tells you panic mode is active. It is a signal to investigate, and one of the most useful alerts on a mesh. ## Application errors are not host failures The subtler trap: by default `consecutive_5xx` counts any 5xx, including a genuine application 500 from a buggy endpoint. If one endpoint reliably throws, every host serving it accumulates consecutive errors and gets ejected in turn — even though every host is perfectly healthy and the problem is in the code. Outlier detection is meant to find hosts that are *different from their peers*, so for the common case prefer `consecutive_gateway_failure` (turning its enforcing percentage up from the default 0), or enable `split_external_local_origin_errors`, or rely on success-rate detection which compares hosts against each other rather than against an absolute rule. ## Interaction with retries Ejection and retries reinforce each other when configured together: a retry with `envoy.retry_host_predicates.previous_hosts` avoids the host that just failed, so one bad endpoint costs a single slow request rather than a failed one. Without ejection, retries keep rediscovering the same bad host; without retries, ejection still lets the first few requests fail before the host is removed. ## Observing it Watch `cluster.<name>.outlier_detection.ejections_enforced_total` (ejections actually acted on), `ejections_active` (currently ejected), `ejections_overflow` (ejections refused because `max_ejection_percent` was reached) and `lb_healthy_panic`. `ejections_overflow` climbing is the direct evidence that your detector is trying to eject more of the pool than the cap allows — meaning the problem is cluster-wide, not host-specific, and ejection is the wrong tool for it.

  • You see hosts marked ejected in /clusters yet traffic still reaches them. Why?
    Panic mode. When the healthy share of the cluster drops below `common_lb_config.healthy_panic_threshold` — 50% by default — Envoy stops honouring health state and balances across all hosts, on the theory that the health signal is more likely wrong than the fleet is dead. The `lb_healthy_panic` counter confirms it, and it is a signal to look at why so much of the pool was ejected.
  • Why does ejecting on consecutive_5xx misfire on an application with one buggy endpoint?
    Because that detector counts any 5xx, including honest application 500s. A route that reliably throws makes every host accumulate consecutive errors and be ejected in turn, though all of them are healthy. Prefer `consecutive_gateway_failure` with its enforcing percentage raised from the default 0, or success-rate detection, which compares hosts with their peers instead of an absolute rule.
  • How long does an ejected host stay out of rotation?
    `base_ejection_time` (default 30s) multiplied by the number of times that host has been ejected, so repeat offenders are removed for progressively longer, bounded by `max_ejection_time`. After the period elapses the host returns to the load-balancing set and is judged again on live traffic, since outlier detection has no separate probe to confirm recovery.

saying these in an interview costs you the question

  • Leaving max_ejection_percent high so most of the pool can be ejected
  • Treating outlier detection as active health checking with probes
  • Assuming an ejected host is never sent traffic again
  • Counting application 500s as evidence a host is broken
  • Ignoring panic mode when diagnosing traffic to unhealthy hosts

context