You operate several hundred Envoy proxies. How would you decide between scraping each proxy's /stats/prometheus endpoint and configuring stats_sinks, and how do you keep the metric volume from each proxy manageable?
answer
- pull gives liveness, push needs no open port
- flush interval versus scrape interval
- clusters multiply per proxy
- never create the stat you will not query
- histogram buckets are series too
basics
~20 sScraping /stats/prometheus is pull-based and needs discovery of every proxy; stats_sinks push to statsd, dog_statsd or a gRPC metrics service on a flush interval. Either way, cardinality is controlled with stats_matcher inclusion or exclusion lists and tuned histogram buckets.
solid answer
~50 sThe choice is pull versus push. Scraping `/stats/prometheus` on each admin interface keeps the proxy stateless about telemetry, gives you scrape-time freshness and per-target up/down signals, but requires service discovery for every proxy and access to a port you otherwise want closed. `stats_sinks` — `envoy.stat_sinks.statsd`, `dog_statsd` or `metrics_service` over gRPC — push on `stats_flush_interval` (default 5s), which suits short-lived or unreachable proxies but adds a delivery dependency and loses the natural liveness signal. The bigger problem at fleet scale is not transport but **cardinality**: each proxy emits a stat set per cluster and per listener, so a few hundred proxies times a few hundred clusters is millions of series. Control it in the proxy with `stats_config.stats_matcher` inclusion or exclusion lists so unwanted families are never created, keep `histogram_bucket_settings` narrow, and treat per-host statistics as opt-in for specific clusters rather than a default.
code
yaml · 12 linesstats_flush_interval: 15s
stats_config:
use_all_default_tags: true
stats_matcher:
inclusion_list:
patterns:
- prefix: "cluster_manager."
- safe_regex: { regex: "^cluster\\.[^.]+\\.(upstream_rq_(total|timeout|retry.*|pending_overflow)|membership_.*|outlier_detection\\..*)$" }
- safe_regex: { regex: "^listener\\..*\\.downstream_cx_(total|active)$" }
histogram_bucket_settings:
- match: { prefix: "cluster." }
buckets: [5, 10, 25, 50, 100, 250, 500, 1000, 2500, 5000]go deeper
Know that Envoy can either be scraped for metrics on its admin interface or push them to a collector, and that each proxy emits a large number of metrics.
Name the sink types and the flush interval, explain the Prometheus rendering of the stats endpoint, and describe why per-cluster stat families multiply so quickly in a mesh.
Argue the pull-versus-push tradeoff for a specific environment, and apply the concrete cardinality levers — matcher lists, histogram buckets, per-host opt-in — with a defensible core set of metrics to keep.
Own the telemetry budget as policy: what every proxy emits by default, who may widen it, how the cost scales with the service catalogue rather than traffic, and how debuggability is preserved when most families are off.
## The two transports **Pull.** Envoy's admin interface renders every stat in Prometheus exposition format at `/stats/prometheus`. A scraper discovers each proxy and pulls on its own schedule. This is the common choice in Kubernetes because pod discovery already exists, and it comes with a useful property: a failed scrape is itself a signal that a proxy is gone, so you get liveness for free without the proxy needing to know anything about your monitoring stack. **Push.** `stats_sinks` in the bootstrap send metrics outward on each flush, governed by `stats_flush_interval` (default 5 seconds). The built-in sinks include `envoy.stat_sinks.statsd`, `envoy.stat_sinks.dog_statsd` (statsd with tag support) and `envoy.stat_sinks.metrics_service`, a gRPC sink speaking Envoy's own metrics service protocol. Push suits proxies that are short-lived, sit behind NAT, or live on a network the scraper cannot reach — and it avoids exposing the admin port at all. ```yaml stats_flush_interval: 15s stats_sinks: - name: envoy.stat_sinks.metrics_service typed_config: "@type": type.googleapis.com/envoy.config.metrics.v3.MetricsServiceConfig grpc_service: envoy_grpc: { cluster_name: metrics_cluster } ``` The decision is rarely about which is technically better. It is about which one your platform already operates well, whether the admin port may be reached at all, and whether losing scrape-based liveness costs you an alerting signal you rely on. ## Cardinality is the real constraint Envoy's stat names are structured — `cluster.<name>.upstream_rq_timeout`, `listener.<address>.downstream_cx_total` — and the per-cluster families are the multiplier. A sidecar in a mesh often has a cluster for **every other service it might call**, so the stat count per proxy scales with the size of the service catalogue rather than with what that workload actually talks to. Multiply by the fleet and by a scrape every fifteen seconds and the storage bill, not the proxy, becomes the binding limit. Three levers, in order of effectiveness: 1. **Do not create the stat.** `stats_config.stats_matcher` takes an `inclusion_list` or `exclusion_list` of name matchers. Excluded stats are never allocated, saving proxy memory as well as pipeline cost. An inclusion list — naming the families you actually alert and dashboard on — is far stronger than an exclusion list that must be maintained against every new stat Envoy adds. 2. **Bound histograms.** `histogram_bucket_settings` sets explicit bucket boundaries per histogram. Default bucket sets are generous; each bucket is a series in Prometheus, so trimming them to the latency range you care about cuts a large share of the volume. 3. **Keep per-host stats opt-in.** Per-endpoint statistics multiply by the number of endpoints and turn a stable series count into one that grows with autoscaling. Enable them for a handful of clusters you are actively debugging, not fleet-wide. Tagging matters too: `use_all_default_tags` extracts embedded name parts (cluster name, response code class) into tags rather than leaving them in the metric name, which makes downstream queries workable. Custom `stats_tags` regexes can extract your own naming conventions, but each extracted dimension is a dimension you now pay for. ## Choosing what to keep A defensible core per cluster: request counts by response class, `upstream_rq_timeout`, retry counters, the circuit-breaker overflow counters, `membership_healthy` versus `membership_total`, outlier ejection counters, and upstream latency histograms. On the listener side, connection counts and downstream protocol errors. That set answers every question in this leaf's triage flows and is a small fraction of the full stat surface. The test to apply per family is: does an alert or a dashboard read it, or would an engineer reach for it during an incident? If neither, exclude it. The stats you removed are still reconstructible for a single proxy during an incident by reading its admin interface directly, which is a genuinely different budget from storing them for every proxy for months. ## Sampling the logs, not the metrics Access logs and stats scale differently and should be budgeted separately. Logs are per request and can be sampled or restricted to error conditions; stats are aggregate and cheap per data point but explode through cardinality. A platform that samples logs but leaves every stat family enabled has usually optimised the wrong one. ## Flush interval and freshness `stats_flush_interval` and scrape interval buy resolution. Fifteen seconds is a common compromise; five seconds doubles the data for detail that most alerts, evaluated over minutes, never use. Set it deliberately rather than leaving the default, and be aware that counters are monotonic, so a longer interval loses resolution but not totals.
- Why is a stats_matcher inclusion list preferable to an exclusion list at fleet scale?An exclusion list must be maintained against everything Envoy might emit, including families added by new versions or new filters, so it drifts open over time. An inclusion list names the families your alerts and dashboards actually read, so the cost is bounded by design and new stats default to off rather than to on.
- What do you lose by switching from scraping to a push sink?The scrape's implicit liveness signal — a failed scrape tells you a proxy is unreachable, whereas silence from a push sink is ambiguous between a dead proxy and a broken pipeline. You also inherit a delivery dependency: if the sink's destination is down, that window of metrics is simply gone, where a pull model would recover on the next successful scrape.
- How would you preserve incident debuggability after excluding most stat families?By keeping the admin interface reachable operationally. Excluded families are not stored centrally, but for a single proxy under investigation you can read the full stat set directly from its admin endpoint, or temporarily widen the inclusion list for one cluster. The tradeoff is deliberate: broad detail on demand, narrow detail retained for everyone.
saying these in an interview costs you the question
- Treating stat volume as a storage problem rather than a config one
- Enabling per-endpoint statistics fleet-wide by default
- Assuming histogram buckets are free once the metric exists
- Leaving the default flush interval without considering resolution
- Exposing the admin port broadly just to enable scraping