Compare the two service-mesh data-plane models — a proxy injected into every pod versus a shared proxy running on each node — in terms of enrolment, resource cost, blast radius and L7 features.
answer
- where the proxy actually runs
- restart to enrol, or not
- cost per pod versus per node
- shared component, shared blast radius
- L7 may cost an extra hop
basics
~20 sA per-pod sidecar isolates each workload and handles L7 itself, but enrolling or upgrading means restarting the pod and paying memory for every pod. A node-shared proxy avoids the restart and scales with nodes, at the cost of a shared blast radius.
solid answer
~50 sThe sidecar model gives every workload its own proxy in its own network namespace. That means strong per-workload isolation, per-workload L7 policy with no extra hop, and a failure domain of exactly one pod — but enrolling a workload, or upgrading the data plane, requires recreating the pod, and you pay proxy memory and CPU once per pod. The node-level model moves a shared proxy onto each node, handling identity and encrypted transport for every pod there. Enrolment and upgrades no longer need a pod restart, cost tracks node count rather than pod count, and the sidecar startup race disappears. In exchange, one component now serves many tenants on a node — a crash, a bug or an upgrade affects all of them — and richer L7 policy typically requires steering traffic through an additional dedicated proxy, which is an extra hop. The node model is also considerably newer, so tooling and operational experience are thinner.
go deeper
Know the basic difference: one proxy inside each pod, versus one shared proxy per node serving all the pods on it.
Explain the mechanical consequences — restart required to enrol or upgrade a sidecar, cost scaling with pods versus nodes, and why request-level features may need an additional hop.
Weigh blast radius, multi-tenancy and maturity against the operational relief of restart-free enrolment, and be specific about which of your services actually need L7 policy.
Own the migration question: whether the estate's pod-to-node ratio, tenancy model and patch-response targets justify moving data-plane models, and how you would run both during the transition.
## The two shapes Every mesh needs its policy applied on the data path. Where that proxy runs is the architectural fork. **Per-pod sidecar.** A proxy container is injected into each workload's pod and shares its network namespace. Redirection rules in that namespace push the pod's inbound and outbound traffic through it. One proxy, one workload, one identity, one configuration set. **Node-level shared proxy.** A single agent runs per node — usually as a daemon — and handles traffic for every pod on that node, providing workload identity and encrypted, authorised transport at the connection level. Application pods have no proxy of their own. Where request-level features are needed, traffic is steered through a separate dedicated proxy for that service or namespace, which the platform runs alongside rather than inside the workloads. ## Enrolment and upgrade This is the difference that changes how a platform team's year looks. With sidecars, the proxy is part of the pod spec, so **enrolling a workload requires recreating it**, and so does upgrading the proxy. That makes a mesh-wide upgrade a fleet-wide rolling restart that the platform team mostly cannot perform unilaterally — application owners have deploy windows, stateful workloads have failover costs, and some pods nobody wants to touch. Security response is bounded by the same constraint: a proxy CVE takes as long to remediate as it takes to restart everything. With a node-level proxy, enrolment is a namespace-level toggle and existing pods are covered without restarting; upgrading the data plane is an upgrade of a daemon on each node, which is ordinary node maintenance the platform team already owns. ## Resource cost Sidecar cost is *per pod* and is dominated by the configuration each proxy holds, so it scales roughly as pods times mesh size. Ten thousand pods means ten thousand proxies, each with its own memory floor, its own CPU request, and its own share of configuration pushes from the control plane. Node-level cost is *per node*. On a cluster with a high pod-to-node ratio — many small pods, sidecars, cron jobs — that is a large saving. On a cluster of a few very large pods it may be a wash. ## Blast radius and isolation Here the sidecar wins. A sidecar's failure domain is one pod: if it crashes, leaks memory or is throttled, exactly one workload is affected, and its resource limits are the workload's own. Its identity material relates to one workload. A node-level proxy is shared infrastructure on the data path of every pod scheduled there. A crash, a bad configuration push or an upgrade affects all of them; one noisy workload can consume the shared proxy's CPU and degrade its neighbours; and the component handles identities on behalf of many workloads at once, which is a real question in hard multi-tenant environments. Node-level designs answer this with a much smaller, simpler per-node component than a full L7 proxy — the smaller the surface, the smaller the risk — but the sharing itself does not go away. ## L7 features and the extra hop A sidecar sits in the request path already, so header-based routing, traffic splitting by weight, request-level authorisation, per-route retries and rich HTTP telemetry cost no extra hop. In the node-level model the shared component is deliberately kept at the connection level. To get request-level behaviour you route traffic through an additional proxy dedicated to the destination service or namespace — which means an extra network hop and another component to size and operate for the services that need it. The pay-off is that services which need *only* identity and encrypted transport — often the majority — never pay for L7 processing at all. That is the model's real thesis: charge each workload only for the layer it actually uses. ## The startup coupling One underrated difference: the sidecar model couples the proxy's lifecycle to the pod's, which is the source of the classic startup race and the Job-never-completes problem. A node-level proxy is already running before your pod is scheduled and keeps running after it exits, so neither failure mode exists. ## Choosing As of 2025 the sidecar model is the mature default with the deepest tooling, the most operational literature and full L7 coverage everywhere. Node-level data planes are newer, and support, feature parity and even terminology differ by mesh and version — check what your version actually implements rather than what a conference talk claimed. Lean node-level when your motivation is mostly encrypted, authorised service-to-service transport at fleet scale, when the pod-to-node ratio makes per-pod cost painful, or when restart-free enrolment is the blocker. Lean sidecar when you need rich per-workload L7 policy broadly, when hard isolation between tenants on a node matters, or when you value the maturity. Mixed deployments are normal: node-level for the long tail, sidecars or dedicated proxies for the services that genuinely need L7.
- Why does the node-level model make a mesh-wide security patch so much faster to apply?Because the proxy is not part of the pod. Patching it is a daemon upgrade the platform team performs as ordinary node maintenance, with no application restarts and no negotiation with service owners. In the sidecar model the fixed proxy only reaches a workload when that workload's pods are recreated, so remediation time is bounded by the slowest team's deploy schedule.
- What isolation property do you give up by sharing one proxy across every pod on a node?The failure and resource domain stops being the pod. A crash, memory leak, bad configuration push or upgrade of the shared component affects every workload on that node, and a noisy workload can starve its neighbours of the proxy's CPU. It also handles identity material for many workloads at once, which matters in hard multi-tenant clusters.
- If most services only need encrypted, authorised service-to-service traffic, why is that an argument for the node-level model?Because those services never need request-level processing, and the sidecar model makes them pay for a full L7 proxy anyway — memory, CPU and a lifecycle coupling. A node-level component provides identity and secure transport at the connection level, and only the minority of services needing header routing or traffic splitting take the extra dedicated proxy hop.
saying these in an interview costs you the question
- Says the node-level model is strictly better with no trade
- Claims sidecars can be upgraded without recreating pods
- Ignores that a shared node proxy is a multi-tenant blast radius
- Assumes full L7 policy comes free in the node-level model
- Treats a newly available mode as equally mature and supported