You push a new configuration to a running Envoy and the proxy keeps serving the old one. Explain how Envoy's xDS ACK/NACK works — version_info, the nonce and error_detail — and how you would confirm the update was rejected.
answer
- every push is acknowledged
- the nonce says which send, the version says which state
- a version that goes backwards is a rejection
- error_detail carries the validation message
- last good config keeps serving
basics
~20 sEvery xDS response carries a version_info and a nonce, and Envoy answers with a request echoing that nonce. An ACK sets version_info to the new version; a NACK keeps the previously accepted version and adds error_detail. A rejected update is discarded whole — Envoy keeps running the last good one.
solid answer
~50 sxDS is an acknowledged protocol. The server's `DiscoveryResponse` carries `version_info` (its opaque label for this state), a `nonce`, a `type_url` and the resources. Envoy validates the whole set, then sends a `DiscoveryRequest` on the same stream with `response_nonce` echoing what it got. The distinguishing field is `version_info`: on success Envoy sets it to the version it just accepted — an ACK; on failure it sets it to the *previously* accepted version and populates `error_detail` with the validation error — a NACK. So a NACK looks like a version that went backwards. Crucially, rejection is all-or-nothing for that response: one bad cluster loses the whole push, and Envoy carries on with its last good configuration, which is exactly the symptom you described. To confirm, look for the rejection in Envoy's logs and its xDS `update_rejected` counter, and read `error_detail` on the control-plane side, where most servers surface which proxy rejected which version and why.
go deeper
Know that xDS updates are acknowledged, and that Envoy can refuse a configuration it considers invalid while continuing to serve the previous one.
Explain the three fields — version_info, nonce, error_detail — and state the wire signature of a NACK: the version does not advance and an error is attached.
Drive the diagnosis: check the proxy's rejection logs and update_rejected counter, read error_detail on the control plane, and know that rejection discards the entire response so an unrelated resource can block your change.
Own the operational consequence: all-or-nothing acceptance means one bad resource can freeze configuration across a fleet with zero traffic impact and zero alarm. Decide where validation belongs before the push, and what alerts on proxies stuck behind the current version.
## The request/response loop An xDS subscription is a long-lived bidirectional gRPC stream, and it is chattier than "server pushes, client obeys". Every push is acknowledged. Envoy opens with a `DiscoveryRequest` naming the `type_url` it wants, the `resource_names` (for by-name types), an empty `version_info` and an empty `response_nonce` — meaning "I have nothing, send me everything". The server replies with a `DiscoveryResponse` containing the resources, an opaque `version_info` string of its own choosing, and a `nonce` unique to this response. Envoy then validates and applies, and sends another `DiscoveryRequest` on the same stream. That request is both the acknowledgement of the previous response and the standing subscription. Two fields make it an ACK or a NACK: - `response_nonce` — always the nonce of the response being answered. This is how the server matches an acknowledgement to a specific send, which matters because it may have pushed several times before hearing back. - `version_info` — the version Envoy is *now* running. **ACK:** `version_info` = the version from the response just received, `error_detail` unset. **NACK:** `version_info` = the last version Envoy successfully applied (or empty, if it never applied one), plus `error_detail` — a `google.rpc.Status` whose message is the validation failure. So the wire signature of a rejection is a version that did not advance. There is no "reject" message type; the version regression *is* the rejection. ## Rejection is per-response, not per-resource This is the part candidates get wrong and the part that bites in production. In State-of-the-World xDS, the unit of acceptance is the entire response for that type. If a push contains 400 clusters and one has an invalid field, Envoy rejects all 400 and keeps the ones it had. Nothing partially applies. The consequence is that a small, obviously-safe change can be blocked by an unrelated defect elsewhere in the same resource type — a config generated for a different service, sharing your CDS push, refuses to validate, and your change silently never lands. It never lands for *every* proxy subscribed to that push, which is why one bad resource can freeze configuration fleet-wide while every proxy continues serving traffic perfectly. (Delta xDS narrows this: rejection is still per-response, but responses carry only changed resources, so a bad resource blocks a much smaller batch.) ## What the proxy does meanwhile Nothing dramatic, and that is by design. A NACKing Envoy keeps its last accepted configuration and keeps serving traffic on it. The data plane is fail-static: bad input does not take it down, it just stops it moving. That property is what makes the symptom "my change didn't take effect" rather than "the site is down", and it is also why the failure can sit unnoticed for hours. The stream itself stays up. Envoy does not disconnect on a NACK, and a well-behaved management server does not either — it takes the `error_detail`, ideally surfaces it, and waits for someone to push a version that validates. ## Confirming it Work the two ends: **On the proxy.** Envoy logs the rejection with the validation message — usually the fastest read, because the message names the offending field. Its xDS subscription stats include an `update_rejected` counter alongside `update_success` and `update_attempt`; a rising `update_rejected` with a flat `update_success` is the signature, and `control_plane.connected_state` tells you separately whether the stream is even up. Dumping the proxy's effective configuration tells you what it is actually running, which settles "did it apply and not work" versus "it never applied". **On the control plane.** The server received `error_detail` and knows which node sent it. Most implementations expose per-proxy sync status; that is where you learn that 3 of 900 proxies are stuck two versions behind, and why. ## Why versions are opaque `version_info` is a string the management server invents. It is not a counter Envoy interprets and not something you can compare for ordering — the only thing Envoy does with it is echo it back. The pairing with `nonce` exists to disambiguate in-flight work: versions identify *state*, nonces identify *a particular send of that state*. If a server pushes v7, then v8 before the v7 acknowledgement arrives, the nonces let it tell which one the incoming acknowledgement refers to; without them, a stale ACK for v7 could be misread as acceptance of v8. ## The mistake to avoid making out loud Do not say Envoy "rolls back" on a NACK. There is no rollback: the new configuration was never applied, so there is nothing to undo. Envoy validated it, refused it, and continued with what it already had. The distinction matters because rollback implies a window in which the bad config was live, and there is none.
- A single malformed cluster is in a CDS push containing hundreds. What actually gets applied?Nothing from that push. In State-of-the-World xDS the whole response for a type is accepted or rejected as one unit, so one invalid resource discards the entire set and Envoy keeps its previous clusters. That is why an unrelated team's bad resource can silently freeze your change — and why delta xDS, where a response carries only what changed, narrows the blast radius.
- Why does the protocol need a nonce when it already has version_info?They answer different questions. `version_info` names the configuration state; the `nonce` names one particular transmission of it. A server may push again before the previous acknowledgement arrives, so without a nonce it cannot tell whether an incoming ACK refers to the push it just made or the one before. Envoy echoes the nonce in `response_nonce` purely so the server can match them up.
- Is it accurate to say Envoy rolls back to the previous configuration when it NACKs?No, and the distinction is worth making. A rejected update is never applied, so there is nothing to roll back — Envoy validated the response, refused it as a unit, and continued running the configuration it already had. Saying rollback implies the bad configuration was briefly live and affecting traffic, which it never was.
saying these in an interview costs you the question
- Says Envoy applies what it can and skips the invalid resource
- Describes a NACK as a rollback of live configuration
- Thinks a rejected update tears down the stream or the listeners
- Treats version_info as a number Envoy compares for ordering
- Assumes a silent no-op means the control plane never sent anything