In a sidecar-based service mesh, how many extra proxies does one service-to-service request cross, and where do the added latency and the extra memory actually come from?
answer
- count both ends of the call
- chains multiply the cost
- tail moves more than the median
- the proxy holds configuration, not traffic
- per pod, times the whole mesh
basics
~20 sTwo: the caller's sidecar and the callee's sidecar, so a chain of N calls crosses 2N proxies. Latency is paid per hop and shows up worst in the tail; sidecar memory tracks how much mesh configuration the proxy holds, not how much traffic it carries.
solid answer
~50 sWith a sidecar per workload, a request leaves the application, is redirected into the local proxy, crosses the network to the callee's proxy, and only then reaches the callee — two extra userspace proxies per call, four extra trips through the local TCP stack. A four-service call chain therefore crosses six sidecars, so per-hop cost multiplies with chain depth rather than being paid once. Latency comes from parsing and forwarding at L7 plus telemetry generation, and it hurts the tail more than the median because the proxy queues behind its own CPU limits. Memory is the surprising one: a proxy's footprint is dominated by the configuration it holds — routes, clusters and endpoints for every destination it has been told about — so it scales with the size of the mesh times the number of pods, not with the pod's request rate. Restricting what each proxy is told about is the main lever.
go deeper
Remember that both ends of a call have a proxy, so mesh traffic passes through two extra processes rather than one.
Explain the three cost drivers — L7 parsing, telemetry generation and configuration held in memory — and why the last one scales with the size of the mesh rather than with traffic.
Show that you measure the tail rather than the median, size proxy CPU so it is not throttled at peak, and scope configuration visibility before the fleet grows.
Own the budget: decide which traffic classes are worth meshing at all, what per-hop latency and per-pod memory the platform will accept, and how those numbers are tracked as the estate grows.
## Counting the hops In a sidecar model, every workload runs with its own proxy sharing its network namespace, and traffic is redirected into it transparently. So a single call from A to B is not one network hop; it is: 1. A's application connects to B's service address; redirection sends the connection to A's sidecar over loopback. 2. A's sidecar picks an endpoint, applies policy, and opens (or reuses) a connection to B's host. 3. That connection lands on B's node and is redirected into B's sidecar. 4. B's sidecar forwards to B's application over loopback. Two proxies, four loopback traversals, one wire hop. The arithmetic that catches people out is the chain: for a request path A → B → C → D there are **three** calls, so **six** sidecars are on the path. Mesh cost is proportional to call-chain depth, and microservice estates with deep chains pay it repeatedly for one user-facing request. ## Where the latency actually goes Three contributors, in roughly this order: - **L7 processing.** Terminating the connection, parsing headers, matching a route, choosing an endpoint and re-serialising costs real CPU per request. A proxy operating at L4 (forwarding bytes without parsing) is markedly cheaper than one doing header-based routing. - **Telemetry.** Every request produces metrics, and often an access-log line and trace spans. Emitting and aggregating those is a per-request cost paid twice, once at each sidecar. - **Scheduling and queueing.** The proxy is a normal container with its own CPU request and limit. Under load, or when the proxy is throttled by its limit, requests wait — which is why the honest measurement is the change in p99 and p999, not the change in the median. Published benchmark figures are workload-specific; treat them as a hint and measure your own path with the mesh enabled and disabled. ## Where the memory actually goes A sidecar's resident memory is dominated by the configuration the control plane has pushed to it: the listeners, routes, upstream clusters and the endpoint lists behind them. By default many meshes tell every proxy about every destination in the mesh, because the control plane cannot know which ones the workload will call. That has two consequences: - **Footprint grows with the size of the mesh, not with the pod's traffic.** A tiny cron pod that calls one service can hold the same endpoint table as your busiest gateway. - **You pay it once per pod.** In a 2,000-pod mesh, an extra 20 MB per proxy is 40 GB of cluster memory that produces no business work, and it grows as *services × pods* rather than linearly. The standard lever is to scope visibility: most meshes provide a way to declare which destinations a given workload actually talks to, so its proxy is sent a small slice of the configuration instead of all of it. That usually cuts both memory and the CPU spent processing configuration pushes. Endpoint churn matters too — a large mesh with high pod churn means constant configuration updates flowing to every proxy that was told about the changed service. ## The cost nobody puts in the spreadsheet The third bill is operational. Your request path now contains a component that can fail on its own terms: a proxy can return a 503 for a connection it could not establish, and that response never appears in the application's logs because the application never saw the request. Debugging shifts from "read the app log" to "read the proxy's access log on both ends, then the application's". Engineers who have not internalised this lose hours to a service that is definitely healthy and definitely returning errors. ## What to do with all this Budget the mesh the way you budget any middleware: per hop, not per system. Before enabling it broadly, measure a representative call chain end-to-end and look at the tail. Size the proxy's CPU so it is not throttled at peak — an under-provisioned sidecar shows up as latency that looks like the application slowing down. Scope configuration visibility early, because retrofitting it across hundreds of workloads is unpleasant. And when the numbers do not justify the bill for a given class of traffic, remember that not all traffic has to be in the mesh: high-throughput internal paths are sometimes deliberately left out, or handled at L4 only, precisely to skip the parsing cost.
- Why does the sidecar's memory not shrink when the pod's traffic drops to nearly zero?Because the footprint is configuration, not buffers. The proxy holds the routes, clusters and endpoint lists the control plane pushed to it, and by default that is most of the mesh regardless of what the workload calls. Idle pods hold nearly as much as busy ones. Restricting the set of destinations a proxy is told about is what reduces it.
- A user-facing request fans out through four internal services and the team measured the mesh as adding one millisecond per hop. What is the end-to-end cost?Roughly six milliseconds at the median, because three internal calls cross six sidecars — and materially more at the tail, since each proxy adds its own queueing under load. Deep synchronous chains are where mesh latency becomes visible, which is an argument for shortening chains as much as for tuning proxies.
- Where do you look first when a service starts returning 503s that never appear in its application log?The proxy's access log at both ends. A response the application never saw was generated by a proxy — typically because it could not connect to the upstream, the upstream had no healthy endpoints, or a policy rejected the request. Comparing the caller's and callee's proxy logs tells you which side produced it.
saying these in an interview costs you the question
- Counts only one proxy per call instead of both ends
- Assumes sidecar memory scales with the pod's request rate
- Quotes a single millisecond figure as universally true
- Ignores that deep call chains multiply the per-hop cost
- Says the sidecar is free because it is 'just a proxy'