In a large Istio mesh, each injected istio-proxy's memory footprint and configuration-push time grow as unrelated services are added elsewhere in the cluster. Why does that happen, and what does Istio's `Sidecar` resource do about it?
answer
- every proxy sees the whole mesh
- cost scales with cluster, not dependencies
- pushes triggered by unrelated namespaces
- declare who you actually call
- undeclared dependency stops working
basics
~20 sBy default istiod sends every sidecar the configuration for every service in the mesh, so per-proxy memory and push cost scale with mesh size rather than with a workload's actual dependencies. A Sidecar resource narrows each proxy's visible scope to the namespaces and hosts it really calls.
solid answer
~50 sIstio's default is to assume any workload might call any other, so every proxy receives configuration describing the whole mesh. That makes per-pod memory grow with the *cluster's* service count, not with your service's dependency list, and it means a deployment in an unrelated namespace triggers a configuration push to proxies that will never route to it — at large scale, a control plane that is permanently busy and proxies that are needlessly fat. The `Sidecar` resource fixes the visible scope: a `workloadSelector` picks which workloads it applies to, and `egress.hosts` declares which namespaces and service hosts those workloads may see, for example `./*` plus `istio-system/*` plus the two namespaces they genuinely call. Proxies then hold a fraction of the config and stop being woken by unrelated changes. The complementary control is `discoverySelectors`, which limits which namespaces istiod watches at all.
code
yaml · 12 linesapiVersion: networking.istio.io/v1
kind: Sidecar
metadata:
name: default
namespace: payments
spec:
egress:
- hosts:
- "./*"
- "istio-system/*"
- "identity/*"
- "platform/kafka.platform.svc.cluster.local"go deeper
Know that the proxy holds configuration pushed to it by the control plane, and that by default that configuration describes the whole mesh rather than just your dependencies.
Explain why default scope makes per-proxy memory and push volume scale with cluster size, and describe what a Sidecar resource's workloadSelector and egress.hosts declare.
Show you have watched the symptoms — per-proxy memory distribution, push latency during other teams' deploys — and can sequence a scoping rollout without breaking undeclared dependencies.
Own the trade: scoping converts implicit reachability into a maintained dependency declaration. Decide whether that artefact is an asset your organisation will keep current, and pair it with discoverySelectors at the cluster level.
## Where the growth comes from Istio's default posture is maximal connectivity: without further information, any workload might talk to any service, so `istiod` computes and pushes configuration covering the whole mesh to every proxy. That is an excellent default for a demo and a poor one at scale, because it couples two things that should be independent — the size of a pod's proxy and the size of the cluster it happens to live in. The consequences show up in three places: - **Per-pod memory.** Each `istio-proxy` holds the full picture. Multiply a growing per-proxy footprint by every pod in the mesh and the mesh's overhead grows quadratically in feel, if not strictly in maths. - **Push volume and latency.** A change in any namespace can require recomputing and pushing configuration to proxies that have no interest in it. Under heavy deployment churn the control plane spends its time on pushes nobody needed, and the pushes that *do* matter queue behind them — the practical symptom being routing changes that take noticeably longer to take effect than they used to. - **Startup cost.** A new pod must receive that whole picture before it is useful, which lengthens exactly the startup window that the ordering problem already made delicate. ## What the `Sidecar` resource declares A `Sidecar` resource is a scope declaration, not a routing rule. It says which other parts of the mesh a set of workloads is allowed to *see*: ```yaml apiVersion: networking.istio.io/v1 kind: Sidecar metadata: name: default namespace: payments spec: egress: - hosts: - "./*" # everything in this namespace - "istio-system/*" # the control plane's own services - "platform/kafka.platform.svc.cluster.local" - "identity/*" ``` With no `workloadSelector`, a `Sidecar` named per convention applies to every workload in its namespace; with one, it applies only to matching workloads. `egress.hosts` entries are `namespace/host` pairs, where `.` means the current namespace and `*` matches all hosts. There is also an `outboundTrafficPolicy` field controlling whether calls to hosts outside the registry are allowed to pass through or are refused. The effect is that proxies in `payments` receive configuration for four namespaces instead of two hundred, and stop being pushed to when an unrelated team deploys. ## The operational trade This is a real trade, not free savings. Scoping turns an implicit "everything is reachable" into an explicit dependency declaration, and **an undeclared dependency stops working**. A team that adds a call to a new service without updating their `Sidecar` gets a failure that looks nothing like a missing route — the destination simply is not in that proxy's world. This is the single reason organisations delay adopting scoping: it converts a silent architectural fact into a maintained artefact. Managed well, that is the point. The `Sidecar` resource becomes a reviewed statement of what a service is permitted to call, which is valuable independently of memory. Managed badly — one enormous namespace-wide `Sidecar` copied everywhere with `*/*` in it — you get the maintenance burden and none of the savings. ## The complementary control `Sidecar` narrows what proxies see. `discoverySelectors` in the mesh configuration narrows what `istiod` watches in the first place, by selecting the namespaces the control plane tracks at all. In a cluster where large namespaces have nothing to do with the mesh — batch infrastructure, third-party operators — that removes work at the source rather than filtering it downstream. The two are usually deployed together: `discoverySelectors` for whole regions of the cluster that are simply not meshed, `Sidecar` for shaping visibility among the namespaces that are. ## How to know you need this Watch per-proxy memory as a distribution across the fleet, not as a single number, and watch how long a configuration change takes to be reflected at the proxies. When the first climbs with cluster growth rather than with your own service count, and the second degrades during other teams' deploy windows, the default scope has outgrown you. Below that — a mesh of a few dozen services — introducing `Sidecar` resources buys little and costs a maintenance obligation, which is why it is fairly graded as a scaling technique rather than a default practice.
- What breaks the first time a team adds a Sidecar resource?Any dependency they forgot to declare. The proxy no longer holds configuration for that destination, so the call fails in a way that does not resemble a routing bug — the service simply is not in that proxy's world. Roll it out by observing real traffic first, deriving the host list from what the workload actually calls, and staging it in a non-production namespace.
- How does discoverySelectors differ from a Sidecar resource?discoverySelectors is mesh configuration that limits which namespaces istiod watches at all, removing work at the source. A Sidecar resource filters what a given set of workloads sees out of whatever istiod does watch. Use the first for parts of the cluster that are not meshed, the second to shape visibility among the parts that are.
- At what size does this stop being premature optimisation?When per-proxy memory starts tracking cluster growth rather than your own dependency count, and when configuration changes visibly take longer to reach proxies during other teams' deploys. Below a few dozen services the default scope is fine and a Sidecar resource mostly buys you a maintenance obligation.
saying these in an interview costs you the question
- Thinks a Sidecar resource routes traffic rather than scoping visibility
- Assumes each proxy only receives config for its own dependencies
- Expects scoping to be free of any maintenance burden
- Confuses it with the injected sidecar container itself
- Believes proxy memory depends only on request volume