Envoy can subscribe to each xDS resource type on its own gRPC stream, or aggregate them onto a single ADS stream. What does ADS guarantee that separate streams do not, and why does that matter during a configuration change?
answer
- one stream, one server, one ordering
- cross-type references need a sequencer
- make before break, break after make
- the 503 nobody can reproduce, at deploy time
- ads: {} per type, ads_config once
basics
~20 sADS puts every resource type on one gRPC stream to one management server, so updates arrive in a guaranteed sequence. Separate streams give no relative ordering, so a route can arrive before the cluster it names and requests fail with 503 until the gap closes.
solid answer
~50 sWithout aggregation, each resource type is its own subscription — potentially its own connection, potentially even a different server — and nothing sequences them against each other. That breaks the one rule Envoy's config graph has: a resource must exist before something references it. Push a route table naming `cluster-v2` while CDS is still in flight, and every matching request 503s until the cluster lands. ADS fixes this by multiplexing all types onto a single bidirectional stream (`StreamAggregatedResources`) to a single management server, which can then order its sends: clusters, endpoints, listeners, routes on the way up, and the reverse on the way down. In the bootstrap you write `ads: {}` as the `config_source` for each type and configure `ads_config` once. It is the default choice for anything non-trivial; separate streams are mainly a simple-setup or legacy arrangement.
code
yaml · 13 linesdynamic_resources:
ads_config:
api_type: GRPC
transport_api_version: V3
grpc_services:
- envoy_grpc:
cluster_name: xds_cluster
cds_config:
resource_api_version: V3
ads: {}
lds_config:
resource_api_version: V3
ads: {}go deeper
Know that ADS means all resource types share one connection to one management server, and that this is what lets updates be applied in a safe order.
Explain the cross-type reference problem concretely — a route arriving before its cluster — and how ordering on a single stream avoids it. Be able to point at ads: {} in the bootstrap.
Bring the production symptom: a burst of 503s at deploy time that the application never logged. Say what you would check on the control-plane side, since ADS only makes correct ordering possible.
Own the trade: one stream is one ordering domain but also one failure domain and one head-of-line queue. Reason about whether a single management server per proxy is the right consistency boundary at your fleet size.
## What "aggregated" actually means By default each xDS subscription is independent: Envoy opens a stream per resource type against the `ConfigSource` you gave that type. Nothing in the protocol relates one stream to another. They can be handled by different processes, they can reconnect at different times, and there is no way for a management server to say "apply this cluster before that route". ADS — the Aggregated Discovery Service — is not a new resource type. It is one gRPC method (`StreamAggregatedResources`) that carries `DiscoveryRequest`/`DiscoveryResponse` messages for *all* types over one stream, discriminated by `type_url`. In the bootstrap, each resource type's `config_source` becomes `ads: {}`, and a single `ads_config` under `dynamic_resources` says where that stream goes. ```yaml dynamic_resources: ads_config: api_type: GRPC transport_api_version: V3 grpc_services: - envoy_grpc: { cluster_name: xds_cluster } cds_config: { ads: {}, resource_api_version: V3 } lds_config: { ads: {}, resource_api_version: V3 } ``` ## The guarantee One stream means one ordering. That gives the management server the ability to implement the **make-before-break** sequence: send clusters, then their endpoints, then listeners, then the routes that point at those clusters. Every reference is resolvable at the moment the referencing resource is applied. Removal runs in reverse — drop the route first, then the cluster — so nothing is torn out from under a live reference. One stream also means one server. That sounds like a limitation and it is a deliberate one: consistency across resource types is only definable if a single authority is deciding what the complete configuration is. A route referencing a cluster is a cross-type invariant, and no two independent servers can maintain it. ## What goes wrong without it The failure is not subtle but it is brief, which is what makes it hard to catch. A deploy adds `payments-v2` and shifts 10% of traffic to it. The route table and the cluster set travel on separate streams. If the RDS push wins the race, every request that falls into the 10% matches a route whose cluster does not exist yet, and Envoy answers 503 immediately — it has nowhere to send them. A second later CDS lands and the errors stop. You are left with a spike of 503s at deploy time that nobody can reproduce, correlated with nothing in the application logs, because the application never saw the requests. The same shape appears on removal: the cluster is deleted before the route referencing it, so for a moment the route points at nothing. ## What ADS does not give you Three limits are worth stating, because candidates routinely over-claim here. 1. **ADS does not make an update atomic.** Resources are still applied as they are received; the guarantee is *ordering*, not a transaction. There is a window in which some of the new configuration is live and some is not — it is just a window in which every reference resolves. 2. **ADS does not remove eventual consistency.** The proxy converges toward what the server has; during convergence it is running a mixture. Traffic-shifting designs have to tolerate that. 3. **ADS does not sequence anything for you if the server does not sequence its sends.** The single stream makes correct ordering *possible*; a management server that pushes routes before clusters on that stream produces exactly the same 503s. This is a property of the pair, not of Envoy alone. ## Operational consequences One stream is also one failure domain and one reconnect. When it drops, Envoy keeps serving the configuration it last accepted — this is important and worth saying explicitly in an interview: **a dead control plane does not take the data plane down**, it freezes it. Envoy reconnects with backoff and, on reconnect, re-sends its subscriptions with the versions it currently holds so the server can decide what it needs. The flip side is that all resource types now share a head-of-line: a very large listener or route push occupies the stream that endpoint updates also need. At large scale this is one of the arguments for the delta variant of ADS, which moves less data per update. ## Delta ADS The aggregated stream has an incremental counterpart, `DeltaAggregatedResources`, selected with `api_type: DELTA_GRPC`. It keeps the single-stream ordering property while sending only changed resources. Aggregation and incrementality are independent choices: you can have either, both, or neither.
- Does using ADS make a configuration update atomic on the proxy?No. Resources are still applied one at a time as they arrive; ADS only guarantees the order they arrive in. There is always a window where part of the new configuration is live and part is not — the point is that within that window every reference still resolves, so you get a consistent intermediate state rather than dangling pointers.
- If the ADS stream to the management server drops, what happens to traffic through the proxy?Nothing immediately. Envoy keeps serving the last configuration it accepted and reconnects with backoff, re-declaring its subscriptions and current versions so the server can resend what changed. The data plane is frozen, not broken — which is why a control-plane outage is usually a change-freeze incident rather than a traffic incident, until something needs an endpoint update to stay correct.
- Your control plane uses ADS and you still see 503s for a new cluster during pushes. Where would you look?At the send order on the server side. ADS makes correct sequencing possible but does not enforce it: if the management server writes the new route configuration to the stream before the cluster it names, the proxy applies them in that order and the reference dangles exactly as it would on separate streams. Check that the server implements clusters, endpoints, listeners, routes.
saying these in an interview costs you the question
- Says ADS makes configuration updates atomic or transactional
- Thinks ADS is a separate resource type rather than a transport arrangement
- Believes losing the xDS stream immediately stops traffic
- Claims separate streams are fine because updates 'arrive fast anyway'
- Assumes ADS orders resources automatically regardless of server behaviour