In an Envoy access log the %RESPONSE_FLAGS% field shows short codes such as UH, UF, UO, UT and URX. What does each one tell you about where the request actually died, and how do you use them to triage a burst of 503s?
answer
- proxy-side cause, not the status code
- empty flag means the app answered
- UH versus UF: attempted or not
- UO is Envoy refusing, not upstream
- flag, then counter, then /clusters
basics
~20 sEnvoy's response flags name the proxy-side cause of an abnormal stream: UH no healthy upstream host, UF upstream connection failure, UO circuit-breaker overflow, UT upstream request timeout, URX retry limit reached. They distinguish upstream faults from Envoy's own decisions.
solid answer
~50 s`%RESPONSE_FLAGS%` is the field in Envoy's access log that says *why* a stream ended badly, which the status code alone cannot. A 503 with `UH` means the cluster had no healthy hosts, so nothing was even attempted; `UF` means Envoy tried to connect and the TCP or TLS handshake failed; `UO` means Envoy refused the request itself because a `circuit_breakers` threshold was full; `URX` means the retry limit or maximum connect attempts was reached; `UT` means the upstream did not answer inside the route timeout. Others you meet are `UR`/`UC` (upstream reset or connection termination), `NR` (no route matched), `DC` (the downstream client hung up) and `RL` (rate limited). Triage flows flag → the matching cluster stat → the `/clusters` admin page: the flag tells you which knob or which side to look at, so you stop guessing whether the backend or the proxy produced the 503.
code
json · 12 lines{
"start_time": "%START_TIME%",
"method": "%REQ(:METHOD)%",
"path": "%REQ(:PATH)%",
"response_code": "%RESPONSE_CODE%",
"response_flags": "%RESPONSE_FLAGS%",
"response_code_details": "%RESPONSE_CODE_DETAILS%",
"upstream_host": "%UPSTREAM_HOST%",
"upstream_transport_failure_reason": "%UPSTREAM_TRANSPORT_FAILURE_REASON%",
"duration_ms": "%DURATION%",
"request_id": "%REQ(X-REQUEST-ID)%"
}go deeper
Know that Envoy writes an access-log line per request and that the response-flags field is a short code explaining why a request failed. Be able to say that an empty flag means the upstream itself produced the response.
Be ready to decode the common flags out loud — UH, UF, UC, UR, UT, UO, URX, NR, DC — and to say which of them mean Envoy never reached the upstream at all. Explain how a flag maps to a cluster stat you would check next.
Show the triage loop: flag, then the matching counter, then per-host state on the admin interface, then a config change. Demonstrate judgment about which flags are user-caused noise and which indicate the proxy is shedding load.
Own the logging contract across the fleet: which fields are mandatory, how flags are indexed and split in dashboards and alerts, and how log volume and sampling are budgeted so the diagnostic field survives cost pressure.
## What the field is Envoy's HTTP connection manager writes one access-log entry per request, and the entry is rendered from a format string or a JSON field map you control. `%RESPONSE_FLAGS%` renders one or more short alphabetic codes describing the proxy-side reason the stream terminated abnormally. It appears in Envoy's default access-log format, so most deployments already have it even if nobody configured logging deliberately. The field matters because Envoy is the thing that *manufactures* many of the error responses your clients see. When a client gets a 503, the status code alone does not say whether an upstream returned 503, whether Envoy could not find a host, or whether Envoy rejected the request before it ever left the proxy. The response flag disambiguates all three. ## The codes you meet most - `UH` — **no healthy upstream**. The cluster had zero hosts passing health checks (or all were ejected). No connection was attempted. Look at endpoint discovery, health checking, or outlier ejection. - `UF` — **upstream connection failure**. Envoy tried and the TCP connect or TLS handshake failed. Pair it with `%UPSTREAM_TRANSPORT_FAILURE_REASON%`, which carries the TLS error text. - `UC` — upstream connection termination: the upstream closed a connection mid-stream. - `UR` — upstream remote reset (for HTTP/2, the upstream sent RST_STREAM). - `UT` — **upstream request timeout**. The route timeout expired while waiting for the upstream response. - `UO` — **upstream overflow**: a `circuit_breakers` threshold on the cluster was full, so Envoy returned 503 immediately without using a connection. - `URX` — the request was rejected because the **retry limit** (HTTP) or maximum connect attempts (TCP) was reached. - `NR` — **no route configured**: the request matched no virtual host or route. Usually a routing-config bug, and it produces 404, not 503. - `NC` — no cluster found for the matched route. - `DC` — **downstream connection termination**: the client went away. Very common and usually *not* your bug. - `SI` — stream idle timeout; `DPE`/`UPE` — downstream/upstream protocol error; `RL` — rejected by the rate-limit service; `UAEX` — rejected by external authorization. Multiple flags can appear joined together on one line, because more than one condition can apply to a single stream. ## Flags disambiguate identical status codes This is the whole point. A wave of 503s can be any of: ``` "GET /api" 503 UH -> cluster empty or fully ejected "GET /api" 503 UF -> cannot connect: bad port, mTLS, firewall "GET /api" 503 UO -> Envoy's own circuit_breakers thresholds full "GET /api" 503 URX -> retries exhausted "GET /api" 503 - -> the upstream application genuinely returned 503 ``` The last line is the important one: an empty flag (`-`) means Envoy proxied the response through unchanged, so the error came from your application. That single distinction saves hours. ## Pair flags with RESPONSE_CODE_DETAILS `%RESPONSE_CODE_DETAILS%` complements the flag with a longer machine-readable reason such as `via_upstream` (the upstream produced this response) or a named local-reply reason. Flag plus details plus `%UPSTREAM_HOST%` is the minimum useful triple: what went wrong, who produced the response, and which endpoint was involved. Adding `%REQ(X-REQUEST-ID)%` lets you join the proxy line to the application's own logs. ## The triage loop 1. **Read the flag** to pick a hypothesis. 2. **Confirm with a counter** on the admin interface, because logs may be sampled but stats are not: `UH` ↔ `cluster.<name>.membership_healthy`; `UO` ↔ `cluster.<name>.upstream_rq_pending_overflow` and `upstream_cx_overflow`; `UT` ↔ `cluster.<name>.upstream_rq_timeout`; `URX` ↔ `upstream_rq_retry` and `upstream_rq_retry_overflow`. 3. **Look at `/clusters`** for per-host state, including health flags such as `/failed_outlier_check` on an ejected host. 4. Only then change configuration. ## Pitfalls Teams frequently alert on 5xx rate without splitting by flag, so a spike of `DC` (clients navigating away or a mobile network dropping) looks identical to a real backend outage. Splitting the log or metric by response flag is usually a one-line change that removes an entire class of false pages. Equally, a JSON access log that omits `response_flags` throws away the most diagnostic field in the record — always include it, and include `upstream_host` next to it.
- An access-log line shows a 503 with an empty response-flags field. What does that tell you?That Envoy did not generate the response — it proxied a real 503 from the upstream, which `%RESPONSE_CODE_DETAILS%` will typically report as `via_upstream`. The investigation belongs in the application, not in proxy configuration. It is the cleanest way to separate "the backend is failing" from "the proxy refused to reach the backend".
- Your dashboards show a large, steady share of requests flagged DC. Is that an incident?Usually not by itself. `DC` means the downstream client terminated the connection — users navigating away, mobile networks dropping, or a client-side timeout shorter than yours. It becomes interesting when the rate jumps in step with rising `%DURATION%`, which suggests clients are giving up because you got slow. Alert on the combination, not on `DC` alone.
- Which extra access-log fields would you add so the flags are actually actionable?`%RESPONSE_CODE_DETAILS%` for the longer reason, `%UPSTREAM_HOST%` to name the endpoint that served or failed, `%UPSTREAM_TRANSPORT_FAILURE_REASON%` to get TLS handshake text behind a `UF`, `%DURATION%` plus the upstream timing fields, and `%REQ(X-REQUEST-ID)%` to join proxy lines with application logs for the same request.
saying these in an interview costs you the question
- Treating every 503 as an upstream application error
- Assuming UH and UF mean the same thing
- Believing UO comes from the backend rather than Envoy
- Alerting on 5xx rate without splitting by response flag
- Removing response_flags when switching to JSON access logs