skip to content

Envoy exposes an admin interface, commonly on port 9901. What do its /stats, /clusters and /server_info endpoints give you, and why must that listener never be reachable from untrusted networks?

level: juniorimportance: should knowfreq 52%

answer

  1. built-in introspection listener
  2. counters, gauges, histograms in one dump
  3. per-host health flags live here
  4. no auth, no TLS, mutating endpoints
  5. bind loopback or a unix socket

basics

~20 s

Envoy's admin interface reports every counter, gauge and histogram at /stats, per-cluster and per-host state at /clusters, and build and configuration details at /server_info. It has no authentication and offers mutating endpoints, so bind it to localhost only.

solid answer

~50 s

The admin interface is Envoy's built-in introspection surface, declared with an `admin` block in the bootstrap config. `/stats` dumps every counter, gauge and histogram the process maintains, with `/stats/prometheus` rendering the same data for scraping and `?filter=` narrowing it. `/clusters` shows each cluster's endpoints with their health flags — including `/failed_outlier_check` on an ejected host — plus per-host connection and request counts, which is where you look when traffic is going somewhere unexpected. `/server_info` reports the Envoy version, build, uptime and current state such as `LIVE` or `DRAINING`. Others in daily use are `/ready`, `/listeners`, `/logging` for runtime log levels, and `/config_dump`. It ships with **no authentication or authorization**, and includes destructive endpoints like `/quitquitquit`, `/drain_listeners` and `/healthcheck/fail`, so the documented practice is to bind it to 127.0.0.1 or a Unix domain socket and reach it through a controlled path.

code

bash · 5 lines
bash
curl -s localhost:9901/server_info | head -20
curl -s localhost:9901/clusters | grep -E 'health_flags|::rq_error'
curl -s 'localhost:9901/stats?filter=orders.*(retry|timeout|overflow)'
curl -s localhost:9901/stats/prometheus | head -5
curl -s -XPOST 'localhost:9901/logging?upstream=debug'

go deeper

for a junior

Know that Envoy has an admin listener and be able to name what /stats, /clusters and /server_info return. Say plainly that it has no authentication and belongs on localhost.

for a middle

Walk a debugging order across the endpoints — server state, listeners, clusters, filtered stats — and explain the difference between the text, JSON and Prometheus renderings of the stats output.

for a senior

Discuss the exposure risk concretely: which endpoints disclose topology and which can terminate or drain the proxy, and how you would expose read-only stats to scrapers without opening the whole interface.

for a principal

Set the fleet-wide contract: where the admin listener binds, what automation may call, how scrape access is granted, and how that policy survives a platform where every workload runs a sidecar with this port present.

## Declaring it The admin interface is configured in the Envoy bootstrap, not through xDS: ```yaml admin: address: socket_address: address: 127.0.0.1 port_value: 9901 ``` The bind address is the security decision, and it should almost always be loopback or a Unix domain socket. ## What each endpoint answers **`/stats`** renders every metric the process holds: counters (monotonic, e.g. `cluster.orders.upstream_rq_retry`), gauges (point-in-time, e.g. `cluster.orders.upstream_cx_active`) and histograms (latency distributions). It is text by default; `/stats?format=json` gives structured output and `/stats/prometheus` (equivalently `/stats?format=prometheus`) gives Prometheus exposition format. With thousands of stats per proxy, `?filter=` with a regex is how you actually use it interactively — for example filtering to one cluster name while debugging. **`/clusters`** is the routing-truth endpoint. For every cluster it lists the endpoints Envoy currently believes exist, their weight and priority, per-host counters for connections and requests, success-rate figures, and a `health_flags` field. Seeing `/failed_outlier_check` there tells you a host has been ejected by outlier detection; seeing an endpoint list that does not match reality tells you the problem is in service discovery rather than in routing. `/clusters?format=json` is easier to parse in a script. **`/server_info`** reports the Envoy version and build, command-line options, uptime, hot-restart epoch, and the server state — `LIVE`, `DRAINING`, `PRE_INITIALIZING` or `INITIALIZING`. It answers two everyday questions immediately: which build is running, and is this process still initializing or already draining. Other endpoints you meet constantly: **`/ready`** returns 200 only when the server is live, so it makes a natural container readiness target; **`/listeners`** shows the listeners and their bound addresses; **`/config_dump`** returns the effective configuration; **`/logging`** reads and changes per-component log levels at runtime, which is how you get debug logging out of a misbehaving proxy without restarting it; **`/certs`** lists loaded certificates with their expiry. ## Why exposure is dangerous There is no authentication, no authorization, and no TLS on the admin listener. Two distinct risks follow. **Disclosure.** `/config_dump` and `/certs` reveal your entire routing topology, cluster names, upstream addresses and certificate metadata. `/stats` reveals traffic volumes and error patterns. This is a detailed internal map handed to whoever can reach the port. **Control.** Several endpoints are mutating: `/quitquitquit` terminates the process, `/drain_listeners` starts draining connections, `/healthcheck/fail` makes the proxy report itself unhealthy so a load balancer removes it, `/reset_counters` wipes your metrics, and `/logging` can be turned to trace level and flood the host with log output. Any of these is a denial of service available to an unauthenticated caller. Mutating endpoints require POST, which is a guardrail against a stray browser GET, not a security control. The real control is the bind address. In a Kubernetes sidecar the admin port is on the pod's loopback interface and reachable from inside the pod; that is already broad enough to be worth thinking about, and it is why platforms often front the admin port with a filtered path that exposes only the read-only stats endpoint. ## How it is normally consumed Humans use it interactively during an incident. Machines use `/stats/prometheus` for scraping and `/ready` for readiness. Automation should prefer specific endpoints over screen-scraping the text `/stats`, whose format is convenient but not a stability contract in the way the Prometheus and JSON renderings are. ## A practical debugging order When a request is not behaving, the sequence is usually: `/server_info` to confirm which build and state you are in, `/listeners` to confirm the traffic can arrive, `/clusters` to confirm the destination endpoints exist and are healthy, then `/stats?filter=` on the cluster name to see the counters for connections, retries, timeouts and overflow. That four-step walk resolves a large share of "the proxy is broken" reports without touching configuration at all.

  • Which admin endpoints can change the proxy's behaviour rather than just report on it?
    `/quitquitquit` terminates the process, `/drain_listeners` begins draining connections, `/healthcheck/fail` and `/healthcheck/ok` flip the proxy's advertised health, `/reset_counters` clears stats, and `/logging` changes log levels live. They require POST, which prevents accidental GETs but is not access control — the bind address is.
  • You suspect a routing problem. What do you look at on /clusters?
    The endpoint list for the cluster in question: whether the addresses you expect are present at all, their priority and weight, and the `health_flags` field — `/failed_outlier_check` means outlier detection ejected the host. Per-host connection and request counters then show whether traffic is spread as expected or concentrated on one endpoint.
  • Why is /ready a better container readiness target than /stats?
    `/ready` returns 200 only once the server is live and has finished initialization, and it is cheap. `/stats` renders thousands of metrics on every call, so using it as a frequent probe wastes CPU and returns 200 even while the proxy is still initializing or already draining, which is exactly when you want the probe to fail.

saying these in an interview costs you the question

  • Exposing the admin port on the pod or node's external address
  • Assuming the admin interface requires authentication
  • Thinking POST-only endpoints are an access control
  • Using /stats as a readiness probe target
  • Believing /clusters shows configured hosts rather than live state

context