In a service mesh built from per-pod Envoy sidecars plus a central control plane like Istio's istiod, what does each half actually do, and why is the split needed instead of configuring every sidecar by hand?
answer
- data plane = many sidecars enforcing
- control plane = one brain computing config
- xDS streams config to Envoy
- istiod = Pilot+Citadel+Galley merged
- sidecars keep last-known config if control plane is down
basics
~20 sThe sidecars (data plane) are the many small proxies that actually move traffic; the control plane is one central brain that tells all of them what rules to follow, so you configure once instead of touching every proxy separately.
solid answer
~40 sA service mesh splits responsibilities: the data plane is the fleet of per-pod sidecar proxies (typically Envoy) that actually intercept and forward every request, applying mTLS, retries, load balancing, and collecting telemetry in real time. The control plane (e.g., istiod) is a separate, centralized component that watches cluster state (services, endpoints, and mesh policy CRDs) and computes/pushes the corresponding configuration to every sidecar over a config-distribution protocol (xDS). This separation exists because manually editing thousands of proxy configs to change one routing rule doesn't scale operationally; the control plane gives one place to declare policy and have it consistently propagate mesh-wide, while the data plane stays a thin, fast, per-hop enforcement layer.
go deeper
Should be able to say sidecars are the many small proxies doing the actual traffic work and the control plane is the one component telling them what to do.
Should know sidecars keep running independently once configured, and name at least one concrete thing the control plane manages (like certificates or routing rules).
Should explain xDS-style config push, the caching/graceful-degradation behavior when the control plane is unavailable, and propagation-lag as a real symptom.
Should reason about control-plane blast radius, why it's designed to be off the hot path, and evaluate the operational cost of running a mesh control plane vs simpler alternatives at a given fleet size.
## What a mesh actually is A service mesh is the pattern of using the sidecar (specifically the **ambassador role** for outbound and the **inbound-interception role** for incoming traffic) at fleet scale, plus a second component that manages all those sidecars centrally. The two halves have deliberately different jobs and deliberately different performance profiles, and understanding why they're split is the core of understanding what a mesh actually is. ## The data plane The data plane is the collection of **sidecar proxies** - one per pod, most commonly Envoy in Istio, or a purpose-built lightweight proxy in Linkerd - that sit on the actual request path. Every request the application sends or receives passes through its local proxy first. This is where the real work happens on every single request: - TLS handshake and certificate validation for **mTLS** - picking a healthy backend instance via **load balancing** - applying **retry and timeout policy** - enforcing **authorization rules** (is this caller allowed to call this endpoint) - recording latency/error metrics and trace spans Because this code runs on the hot path of every request, it is written to be fast and simple to reason about - Envoy itself does not decide policy, it only enforces whatever configuration it was last told to enforce. ## The control plane The control plane is a separate, much smaller set of processes (in Istio, this is `istiod`) that does NOT sit on the request path at all. Its job is to watch the state of the cluster - what services exist, what pods are healthy and where, and what mesh policies the operator has declared (as Kubernetes custom resources like `VirtualService`, `DestinationRule`, `PeerAuthentication`) - and translate that into the low-level configuration each individual sidecar needs, then push it out. The mechanism for that push in Istio is **xDS** (Envoy's own discovery protocol family: Listener, Route, Cluster, and Endpoint Discovery Services), a gRPC streaming protocol where every sidecar maintains a persistent connection to istiod and receives incremental config updates as cluster state changes - a new pod appearing, a policy being edited, a certificate needing rotation. ## Why the split exists The reason for this split rather than hand-configuring (or scripting configuration for) every sidecar individually is pure **operational scale**. A mesh with 2,000 pods has 2,000 Envoy processes; a routing change (say, shifting 10% of traffic to a canary version) or a security change (rotating mTLS root certificates fleet-wide) needs to land consistently and near-simultaneously across all of them. Without a control plane, you'd need custom tooling to push config to every proxy and reconcile drift when it inevitably diverges; with one, an operator edits a single declarative resource and the control plane computes the diff and streams it to every affected sidecar automatically, with the mesh converging within seconds. The control plane is also the natural place to run the **certificate authority** that issues and rotates the mTLS identities every sidecar uses, since it already has a trusted channel to every proxy. ## Failure modes The trade-off of this design is a new critical-path dependency and a new class of failure mode entirely separate from application bugs. 1. **The control plane goes down.** Sidecars typically keep enforcing their last-known-good configuration (Envoy caches what it was told), so existing traffic usually keeps flowing - but any change that needed a fresh push (a new pod coming up needing initial config, a cert nearing expiry needing rotation) stalls, and a long enough outage eventually causes cascading failures as caches go stale or certs expire. 2. **Another real failure mode is config-push lag.** A pod that just scaled up may not be in every other sidecar's known-endpoints list yet, so it silently receives no traffic for a few seconds after starting - this shows up as 'why isn't my new pod getting any requests' tickets, and is usually a propagation-delay issue, not a bug in the app. ## A concrete example A concrete, named example: Istio's `istiod` process (which, in Istio 1.5+, consolidated what used to be three separate components - **Pilot** for xDS config generation, **Citadel** for certificate issuance, and **Galley** for config validation - into a single binary) streams xDS updates to every Envoy sidecar in the mesh; when an operator applies a `VirtualService` that splits traffic 90/10 between two versions of a service, istiod recomputes the routing table and pushes updated Route Discovery Service data to every sidecar that could call that service, typically converging mesh-wide within a few seconds without touching a single pod directly.
- What actually happens to in-flight traffic if the control plane (istiod) becomes completely unavailable for ten minutes?Existing sidecars keep enforcing whatever configuration they last received, since Envoy caches its config locally and doesn't need a live connection to keep routing traffic - so most established traffic keeps flowing normally. What breaks is anything requiring a fresh update: new pods won't get initial config, mTLS certificates nearing their rotation window won't be renewed, and policy changes won't propagate until the control plane comes back.
- Why does Istio use a streaming protocol (xDS over gRPC) instead of having sidecars poll a REST endpoint for config?Streaming lets the control plane push incremental updates the instant cluster state changes, so propagation is near-real-time rather than bounded by a polling interval; at fleet scale, a persistent stream per sidecar is also far cheaper than thousands of proxies repeatedly polling and re-fetching full config.
- A newly scaled-up pod receives no traffic for several seconds after its container reports ready. What's the likely explanation given how the mesh's control and data planes interact?This is typically endpoint-propagation lag: the control plane hasn't yet pushed the new pod's address to every other sidecar's Endpoint Discovery Service data, so callers' proxies don't know it exists yet. It usually self-resolves within seconds as the control plane's watch-and-push loop catches up, and readiness/health-check tuning can reduce the window.
The data plane/control plane split is like air-traffic control versus the planes themselves - the many planes (sidecars) are what actually move and interact with the world in real time, while the control tower (control plane) doesn't fly anywhere but tells every plane what route to follow, and planes keep flying their last instruction even if radio contact briefly drops.
saying these in an interview costs you the question
- Thinks the control plane sits on the request path and adds latency to every call
- Cannot say what happens to existing traffic if the control plane goes down
- Confuses the control plane with the API gateway
- Doesn't know sidecars enforce policy while the control plane computes/distributes it
- Unaware that config changes take a nonzero propagation time across the mesh