Envoy's xDS comes in a State-of-the-World form and a delta (incremental) form. For a large fleet of proxies, how would you decide between them, and what does each cost the control plane and the proxy?
answer
- full state versus what changed
- who keeps the bookkeeping
- removed_resources exists because omission stops meaning deletion
- reconnect is free in one, careful in the other
- scoping beats encoding
basics
~20 sState-of-the-World resends the full resource set for a type on every change; delta sends only what changed plus removed resource names. Delta cuts push volume sharply when churn is high and resource sets are large, at the cost of per-client state tracking in the control plane.
solid answer
~50 sIn SotW, a `DiscoveryResponse` for a type carries the complete set the proxy should hold, so a single endpoint change reserializes and resends everything — fine at a hundred clusters, brutal at ten thousand across a thousand proxies. Delta xDS (`DeltaDiscoveryRequest`/`DeltaDiscoveryResponse`, `api_type: DELTA_GRPC`) sends only changed resources, each with its own version, plus `removed_resources` for deletions, and lets a proxy adjust its subscription with `resource_names_subscribe`/`resource_names_unsubscribe`. The trade is where the bookkeeping lives: SotW keeps the server nearly stateless, since it just sends current truth, while delta requires it to track what every connected proxy already holds and compute a diff per client — more memory, more code, harder recovery. I would reach for delta when push volume or control-plane CPU is measurably the constraint, and I would first check whether *scoping* — sending each proxy only the resources it actually needs — solves the same problem more cheaply, because it usually does.
go deeper
Know the distinction in one line: State-of-the-World responses carry the complete set of resources, delta responses carry only what changed plus what was removed.
Explain why deletion needs an explicit removed_resources list once responses stop being complete, and where each protocol's cost shows up on the wire.
Reason about the operational difference — reconnect is self-correcting under SotW, and delta narrows the batch that a single invalid resource can block. Say what you would measure before switching.
Own the decision with numbers: name the constraint (control-plane CPU, egress, convergence time), argue scoping before protocol, and weigh the quiet correctness risk of per-client diff state against the bandwidth it saves.
## The two protocols **State of the World.** The name is literal: a response tells the proxy what the world looks like now for that resource type. Envoy replaces its held set with what arrived. For wildcard types (listeners, clusters) omission means deletion — if a cluster is missing from the response, it is gone. That is the property that forces completeness: the server cannot send a subset without implicitly deleting everything else. **Delta (incremental).** A response carries only resources that changed, each with its own per-resource version, plus `removed_resources` naming what to drop. Subscriptions are managed explicitly: `resource_names_subscribe` and `resource_names_unsubscribe` on the request. On reconnect the client can declare `initial_resource_versions` — what it already holds — so the server sends only the difference rather than the world. Both exist in aggregated form. `StreamAggregatedResources` is SotW-over-ADS; `DeltaAggregatedResources` is delta-over-ADS, selected with `api_type: DELTA_GRPC`. Aggregation and incrementality are orthogonal choices. ## Where the cost lands ### SotW - **Server:** nearly stateless per client. It holds current truth and serializes it. It does not need to know what any proxy has. - **Wire:** every change costs a full set. Cost scales as (resources per proxy) × (proxies) × (change rate) — and the change rate of endpoints in a busy platform is not small. - **Proxy:** must ingest and diff the full set to work out what actually changed, so CPU per update tracks total configuration size, not change size. ### Delta - **Server:** must track per-client state — which resources each proxy is subscribed to and at what version — and compute a diff per connection. Memory grows with connected proxies × resources; the code is materially harder, and getting reconnect and unsubscribe semantics wrong yields the worst possible bug class, a proxy that silently holds a resource the server thinks it removed. - **Wire:** proportional to actual change. One endpoint moving is one small message. - **Proxy:** applies only what changed. There is a subtlety worth knowing on the SotW side: for wildcard types the complete set is mandatory, but for by-name types like endpoints and route configurations a SotW response does not have to include every subscribed resource, and omission there does not mean deletion. So endpoint churn under SotW is not always a full-world resend — but it does resend the entire load assignment for each affected cluster, whatever the size of the change within it. ## How I would actually decide **Start by measuring, not choosing.** The question is what is saturating: control-plane CPU, egress bandwidth, or convergence latency during large events. Delta helps the first two and helps the third only insofar as smaller messages arrive sooner. **Check scoping first.** The largest wins in this space usually come from sending each proxy *less*, not from encoding the same amount more efficiently. A sidecar that only ever calls three services does not need every cluster in the platform, and restricting its subscription cuts both its memory and everyone's push cost — and it keeps working on either protocol. A control plane that pushes the full catalogue to every proxy has a scoping problem that delta will merely make cheaper to keep having. **Then weigh the event that actually hurts.** The pathological case for SotW is high-churn endpoints across a large resource set: draining a node moves hundreds of endpoints and every affected proxy gets full resends. If that is your incident shape, delta is the right tool. **Account for the control plane you have.** If you did not write it, delta support and its maturity are a given, not a decision. If you did, be honest that per-client diff state is a significant, correctness-sensitive addition — and that its failure modes are quiet. ## Failure and recovery Recovery is the underrated axis. Under SotW, reconnect is trivially correct: the server sends the world, the proxy replaces its state, and any prior divergence is erased by construction. Under delta, reconnect depends on `initial_resource_versions` being honest on both sides; a bug there produces a proxy quietly running a resource nobody believes is live, and no amount of pushing fixes it because the server thinks the proxy is already current. When a fleet-wide reconnect storm follows a control-plane restart, SotW is self-healing and delta is only as correct as its bookkeeping. Acceptance semantics are unchanged in shape — version, nonce, `error_detail` — but the blast radius differs, and this is a genuine operational argument for delta: rejection is still per-response, and a delta response contains only the changed resources, so one invalid resource blocks a small batch rather than an entire resource type. ## What I would tell a team Default to SotW with ADS and aggressive scoping. Move to delta when you have a measurement showing push volume or control-plane CPU is the binding constraint, and when the control plane's delta implementation is one you trust to get reconnect right. Do not adopt it because it sounds more efficient; adopt it because you can name the number it improves.
- In delta xDS, why is there an explicit removed_resources list when SotW does not need one?Because omission stops carrying meaning. Under SotW a resource missing from a wildcard response is deleted by definition, since the response is the complete world. A delta response contains only what changed, so absence just means unchanged — deletion therefore needs its own explicit channel, which is what `removed_resources` provides.
- What makes reconnect harder under delta than under State-of-the-World?Under SotW the server resends everything and the proxy replaces its state, so any prior divergence is erased by construction. Under delta the proxy declares what it already holds via `initial_resource_versions` and the server sends only the difference — which is correct only if that bookkeeping is right on both sides. A mismatch leaves a proxy quietly running a resource the server believes is gone, and further pushes will not fix it.
- Before adopting delta xDS to cut push volume, what cheaper change would you evaluate first?Scoping the subscriptions. Most fleets push far more to each proxy than it needs — the whole service catalogue to a sidecar that talks to three services. Cutting what each proxy is sent reduces bandwidth, proxy memory and control-plane serialization at once, works identically under either protocol, and often removes the pressure that made delta look necessary.
- Does moving to delta xDS change how a rejected update behaves?The mechanism is identical — the proxy answers with the previous version plus error_detail — but the blast radius shrinks. Rejection is per-response, and a delta response carries only changed resources, so one invalid resource blocks a small batch instead of discarding an entire resource type's worth of state. That is a real operational argument for delta beyond raw bandwidth.
saying these in an interview costs you the question
- Calls delta strictly better without naming what it costs
- Thinks delta makes the control plane simpler
- Believes delta fundamentally changes acknowledgement semantics
- Assumes SotW always resends every resource type on any change
- Reaches for delta before checking how much each proxy is even sent