skip to content

Envoy

Envoy is the L4/L7 proxy that sits under most modern service meshes and API gateways, and it is the first proxy designed to be configured entirely over an API rather than a text file. Interviewers ask about it because understanding listeners, filter chains, clusters and xDS is what separates someone who can debug a mesh from someone who only edits its YAML.

on this pageshow

explore

questions

18

Envoy can be configured from a static local file or dynamically through xDS. What is xDS, and what configuration does an xDS-driven Envoy still have to read from a local file?

level: juniorimportance: must knowfreq 60%

answer

  1. configuration as a subscription, not a file
  2. a management server pushes, Envoy applies
  3. no reload signal, no restart
  4. the control plane's own address cannot be discovered
  5. bootstrap file: node, admin, dynamic_resources

basics

~20 s

xDS is Envoy's family of discovery APIs: a management server streams listeners, routes, clusters and endpoints into a running proxy, applied with no restart. Only the bootstrap — node identity and how to reach that server — stays in a local file.

solid answer

~50 s

xDS stands for "x Discovery Service" — a family of APIs (LDS, RDS, CDS, EDS, SDS and others) that a control plane implements and Envoy subscribes to, over a gRPC stream or REST polling. Instead of editing a config file and reloading, a management server sends Envoy the resources it should be running, and Envoy swaps them in on the live process: no restart, no dropped connections, no reload signal. What stays local is the **bootstrap** file: the `node` identity Envoy presents, an `admin` block if you want one, `dynamic_resources` saying which discovery services to use, and a `static_resources` section that must contain the cluster pointing at the management server itself — Envoy cannot discover the thing it discovers config from. Anything in the bootstrap, and the binary itself, still needs a process restart (or a hot restart) to change.

code

yaml · 37 lines
yaml
node:
  id: sidecar-1
  cluster: checkout

dynamic_resources:
  ads_config:
    api_type: GRPC
    transport_api_version: V3
    grpc_services:
      - envoy_grpc:
          cluster_name: xds_cluster
  lds_config:
    resource_api_version: V3
    ads: {}
  cds_config:
    resource_api_version: V3
    ads: {}

static_resources:
  clusters:
    - name: xds_cluster
      type: STRICT_DNS
      connect_timeout: 1s
      typed_extension_protocol_options:
        envoy.extensions.upstreams.http.v3.HttpProtocolOptions:
          "@type": type.googleapis.com/envoy.extensions.upstreams.http.v3.HttpProtocolOptions
          explicit_http_config:
            http2_protocol_options: {}
      load_assignment:
        cluster_name: xds_cluster
        endpoints:
          - lb_endpoints:
              - endpoint:
                  address:
                    socket_address:
                      address: control-plane
                      port_value: 18000

go deeper

for a junior

Be able to say plainly that xDS is a push API: a management server sends Envoy its listeners, routes, clusters and endpoints while it runs, and Envoy needs no restart. Name the bootstrap file as the one static piece.

for a middle

Explain the transport underneath — a streaming gRPC subscription per resource type or one aggregated stream — and why the cluster pointing at the management server has to be in static_resources.

for a senior

Show you know what is still not dynamic: bootstrap fields and the binary. Describe how you would roll those out with hot restart or a rolling replacement without dropping in-flight connections.

for a principal

Own the boundary between control plane and data plane: what belongs in a bootstrap you version with the image, what belongs in a push, and what it costs the organisation when proxy behaviour becomes owned by a central service rather than a file in a repo.

## What problem xDS solves A classic reverse proxy is configured by a file. You edit it, you signal the process, it re-reads the file. That model assumes configuration changes rarely and that a human wrote it. In a dynamic environment — instances appearing and disappearing, routes shifting during a deploy — a file plus a reload is the wrong shape: the file has to be regenerated by something, written to every host, and every proxy has to be told to re-read it, with no feedback about whether it worked. Envoy inverts this. It exposes a set of gRPC/REST APIs, collectively called **xDS**, that it *consumes* as a client. Some other process — the control plane, or management server — implements the server side and pushes configuration to each connected proxy. Envoy applies each update in place on the running process. ## The resource types Each discovery service carries one kind of resource: - **LDS** — Listener Discovery Service: the listeners (sockets and their filter chains). - **RDS** — Route Discovery Service: the route configurations an HTTP listener refers to by name. - **CDS** — Cluster Discovery Service: the upstream clusters. - **EDS** — Endpoint Discovery Service: the actual endpoint addresses backing a cluster. - **SDS** — Secret Discovery Service: TLS certificates and keys, so private keys never have to sit in a config file on disk. There are more (runtime layers, extension configs, scoped routes), but these five are what interviewers mean by xDS. Each is delivered as a stream of typed resources, and each is optional — you can run any mix of static and dynamic configuration. ## The bootstrap stays static The one thing that cannot come over xDS is the description of how to reach the management server. Envoy reads a bootstrap file at startup (`-c bootstrap.yaml`) containing: ```yaml node: id: sidecar-1 cluster: checkout dynamic_resources: lds_config: { ads: {}, resource_api_version: V3 } cds_config: { ads: {}, resource_api_version: V3 } ads_config: api_type: GRPC transport_api_version: V3 grpc_services: - envoy_grpc: { cluster_name: xds_cluster } static_resources: clusters: - name: xds_cluster # ... address of the control plane ``` The cluster named by an xDS `ConfigSource` must be defined in `static_resources`; it is bootstrapped, not discovered. The `node` block matters too — the control plane usually decides *what* to send based on the node id and metadata the proxy presents on connect. ## Transport options `api_type` selects how updates arrive: `GRPC` (a streaming bidirectional connection, the normal choice), `DELTA_GRPC` (the incremental variant), or `REST` (Envoy polls a URL on a `refresh_delay`). Streaming gRPC is what makes push semantics possible — the server sends when something changes rather than waiting to be asked. `transport_api_version: V3` pins the API version; v2 was removed from Envoy some years ago. A `ConfigSource` can also be `path` — a file Envoy watches. That is still "dynamic" configuration in Envoy's sense (it goes through the same update machinery, acknowledgements and warming) even though no control plane is involved, and it is a common way to run xDS-style config without writing a gRPC server. ## What is still not dynamic Three things need a process restart: 1. **Bootstrap changes** — node id, admin port, the xDS cluster's address, most `layered_runtime` wiring. 2. **A new binary or a new set of compiled-in extensions.** 3. **Anything below Envoy** — the container image, the mounted filesystem. For these Envoy offers *hot restart*: a new process starts with an incremented `--restart-epoch`, takes over the listening sockets from the old process over a unix domain socket, inherits its stats, and the old process drains and exits. That is a separate mechanism from xDS and is used far less often once xDS is in place, because the things people used to restart for are now pushed. ## The mental model to carry into an interview Static file configuration answers "what should this proxy do?" once, at start. xDS answers it continuously, and makes the proxy a *data plane* — a thing whose behaviour is owned by a control plane rather than by whoever last edited a file on the box. Every other xDS question (ordering, acceptance and rejection, warming) falls out of that: once configuration is a stream of updates into a live process, you need a way to sequence them, a way to say "I rejected that", and a way to avoid switching to a new version before it is ready.

  • If Envoy gets everything over xDS, why does the bootstrap still define a cluster in static_resources?
    Chicken-and-egg: the cluster that points at the management server cannot itself arrive over xDS, so Envoy requires the cluster named by an xDS `ConfigSource` to be statically defined in the bootstrap. That static cluster carries its own address, connect timeout and TLS context, and it is the one piece of upstream configuration you still deploy as a file.
  • What changes do still require restarting an Envoy process, and how do you do that without dropping connections?
    Bootstrap fields (node id, admin address, the xDS cluster) and a new binary or extension set. Envoy handles those with hot restart: the new process starts with an incremented `--restart-epoch`, takes the listening sockets and stats over from the old one through a domain socket, and the old process drains for `--drain-time-s` before exiting, so established connections finish instead of being reset.
  • Can Envoy take dynamic configuration without running a gRPC control plane at all?
    Yes. A `ConfigSource` can specify `path` instead of `api_config_source`, and Envoy watches that file for changes and applies them through the same dynamic-update path — including warming and validation. It is a common way to get xDS semantics from a file-writing agent, and it is also how people experiment with xDS resource formats before building a real server.

saying these in an interview costs you the question

  • Says Envoy needs SIGHUP or a reload to apply xDS updates
  • Thinks xDS replaces the bootstrap file entirely
  • Confuses xDS with a specific control plane — xDS is the API
  • Claims a config file change and an xDS push are the same mechanism
  • Assumes the management server address itself arrives over xDS

context

open as a page

In an Envoy http_connection_manager, http_filters is an ordered list. Why must envoy.filters.http.router be the last entry, and in what order do the filters actually see a request versus its response?

level: middleimportance: must knowfreq 52%

basics

~20 s

The router is Envoy's terminal HTTP filter: it sends the request upstream and ends the chain, so anything listed after it can never run. Envoy rejects such a config. Filters see requests in listed order and responses in reverse order.

open as a page

Trace a single HTTP request through Envoy's configuration objects, from the TCP port it arrives on to the backend instance that finally serves it. Which object handles each step, and what does each one own?

level: middleimportance: must knowfreq 78%

basics

~20 s

A listener binds the port and selects a filter chain; the http_connection_manager network filter parses HTTP and runs the HTTP filters; its route_config picks a virtual host and a route that names a cluster; the cluster load-balances across its endpoints.

open as a page

In an Envoy access log the %RESPONSE_FLAGS% field shows short codes such as UH, UF, UO, UT and URX. What does each one tell you about where the request actually died, and how do you use them to triage a burst of 503s?

level: middleimportance: must knowfreq 60%

basics

~20 s

Envoy's response flags name the proxy-side cause of an abnormal stream: UH no healthy upstream host, UF upstream connection failure, UO circuit-breaker overflow, UT upstream request timeout, URX retry limit reached. They distinguish upstream faults from Envoy's own decisions.

open as a page

Envoy's xDS APIs are split into LDS, RDS, CDS and EDS (plus SDS). What does each one deliver, and how do they depend on one another?

level: middleimportance: must knowfreq 65%

basics

~20 s

LDS delivers listeners, RDS the route tables they name, CDS the upstream clusters routes point at, EDS the endpoint addresses inside those clusters, and SDS the TLS secrets. The dependency runs listener to route to cluster to endpoint, which is why updates are ordered CDS, EDS, LDS, RDS.

open as a page

An Envoy route is configured with `timeout: 3s` and a `retry_policy` of `num_retries: 3` with `per_try_timeout: 3s`, but slow upstreams are never retried in production. Why not, and how should the two timeout values relate?

level: seniorimportance: must knowfreq 50%

basics

~20 s

Envoy's route timeout covers the whole request including every retry, so a per_try_timeout equal to it means the first attempt consumes the entire budget and the request fails before a retry can start. Per-try must be a fraction of the overall timeout.

open as a page

In Envoy's terminology and configuration, what do the words 'downstream' and 'upstream' refer to, and which Envoy objects sit on each side?

level: juniorimportance: should knowfreq 55%

basics

~20 s

In Envoy, downstream is the client that connects to Envoy and upstream is the service Envoy connects out to. Listeners and their filter chains are the downstream side; clusters and their endpoints are the upstream side.

open as a page

Envoy exposes an admin interface, commonly on port 9901. What do its /stats, /clusters and /server_info endpoints give you, and why must that listener never be reachable from untrusted networks?

level: juniorimportance: should knowfreq 52%

basics

~20 s

Envoy's admin interface reports every counter, gauge and histogram at /stats, per-cluster and per-host state at /clusters, and build and configuration details at /server_info. It has no authentication and offers mutating endpoints, so bind it to localhost only.

open as a page

Inside an Envoy route_config, how does Envoy choose a virtual host and then a route for a request, and why can a routes entry with match: { prefix: "/" } placed first make every later entry unreachable?

level: middleimportance: should knowfreq 48%

basics

~20 s

Envoy first selects a virtual host by matching the request authority against its domains, most specific first, then scans that host's routes strictly in order and takes the first match. A leading prefix of "/" matches everything, so nothing after it is ever reached.

open as a page

An Envoy cluster's `circuit_breakers` block sets thresholds for max_connections, max_pending_requests, max_requests and max_retries. What does each one limit, and what does a client see when one is exceeded?

level: middleimportance: should knowfreq 50%

basics

~20 s

They are concurrency caps, not a tripping breaker: max_connections limits upstream connections, max_pending_requests limits requests queued waiting for one, max_requests limits in-flight requests, max_retries limits concurrent retries. Exceeding a limit yields an immediate 503 flagged UO.

open as a page

Envoy can subscribe to each xDS resource type on its own gRPC stream, or aggregate them onto a single ADS stream. What does ADS guarantee that separate streams do not, and why does that matter during a configuration change?

level: middleimportance: should knowfreq 45%

basics

~20 s

ADS puts every resource type on one gRPC stream to one management server, so updates arrive in a guaranteed sequence. Separate streams give no relative ordering, so a route can arrive before the cluster it names and requests fail with 503 until the gap closes.

open as a page

An Envoy cluster with type: LOGICAL_DNS points at a DNS name that resolves to many backend addresses, yet almost all traffic lands on one backend. What does LOGICAL_DNS do that STRICT_DNS does not, and which would you use here?

level: seniorimportance: should knowfreq 35%

basics

~20 s

LOGICAL_DNS keeps a single logical host built from one resolved address, so the load balancer has exactly one endpoint to choose. STRICT_DNS keeps every address returned by DNS as a separate host and balances across all of them, which is what a multi-backend name needs.

open as a page

After enabling `outlier_detection` on an Envoy cluster, a short burst of upstream errors ejects most of the pool and latency gets worse instead of better. Which fields drive that behaviour, and what does Envoy do once too many hosts are ejected?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Aggressive consecutive_5xx with a large max_ejection_percent ejects healthy hosts during an error burst, concentrating load on survivors. Once healthy hosts fall below healthy_panic_threshold (default 50%), Envoy enters panic mode and load balances across all hosts, ignoring health.

open as a page

You push a new configuration to a running Envoy and the proxy keeps serving the old one. Explain how Envoy's xDS ACK/NACK works — version_info, the nonce and error_detail — and how you would confirm the update was rejected.

level: seniorimportance: should knowfreq 50%

basics

~20 s

Every xDS response carries a version_info and a nonce, and Envoy answers with a request echoing that nonce. An ACK sets version_info to the new version; a NACK keeps the previously accepted version and adds error_detail. A rejected update is discarded whole — Envoy keeps running the last good one.

open as a page

You push a CDS update to a running Envoy and traffic does not move to the new cluster for several seconds. What is cluster warming, what blocks on it, and what happens if warming never completes?

level: seniorimportance: should knowfreq 35%

basics

~20 s

A cluster added or changed via CDS is not usable until it warms: Envoy waits for its first endpoint data — an initial EDS response or initial DNS resolution — before it can take traffic. During warming the previous version keeps serving, so an update looks delayed rather than broken.

open as a page

A single Envoy listener on port 443 carries several filter_chains with different filter_chain_match blocks. How does Envoy decide which chain handles an incoming connection, and what has to happen before it can match on the TLS server name?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Envoy picks the single most specific matching filter chain, comparing criteria in a fixed priority order rather than taking the first that matches. Matching on SNI requires a listener filter to read the TLS ClientHello before selection happens.

open as a page

You operate several hundred Envoy proxies. How would you decide between scraping each proxy's /stats/prometheus endpoint and configuring stats_sinks, and how do you keep the metric volume from each proxy manageable?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Scraping /stats/prometheus is pull-based and needs discovery of every proxy; stats_sinks push to statsd, dog_statsd or a gRPC metrics service on a flush interval. Either way, cardinality is controlled with stats_matcher inclusion or exclusion lists and tuned histogram buckets.

open as a page

Envoy's xDS comes in a State-of-the-World form and a delta (incremental) form. For a large fleet of proxies, how would you decide between them, and what does each cost the control plane and the proxy?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

State-of-the-World resends the full resource set for a type on every change; delta sends only what changed plus removed resource names. Delta cuts push volume sharply when churn is high and resource sets are large, at the cost of per-client state tracking in the control plane.

open as a page