skip to content

Service Mesh

Push retries, timeouts, mTLS and telemetry out of every service and into a sidecar proxy managed by a control plane, as Istio and Linkerd do. The interview conversation is whether that uniformity is worth the operational weight of the mesh itself.

part ofMicroservices architectureoverview, primer and where to startread it →
on this pageshow

questions

6

In a sidecar-based service mesh like Istio or Linkerd, what is a sidecar proxy and how does it change the way one microservice's network traffic reaches another?

level: juniorimportance: must knowfreq 70%

answer

  1. sidecar = second container per pod
  2. iptables/eBPF traffic redirection
  3. app talks to localhost
  4. Envoy (Istio) vs linkerd2-proxy (Linkerd)
  5. no app code changes needed

basics

~10 s

A sidecar is a small helper program that runs next to each service and handles all its network traffic — sending, receiving, encrypting — so the service's own code doesn't have to.

solid answer

~30 s

In a sidecar-based mesh, every service instance gets a companion proxy (Envoy for Istio, linkerd2-proxy for Linkerd) deployed in the same pod. All inbound and outbound traffic is transparently intercepted (via iptables rules or eBPF) and routed through this proxy instead of going directly between application containers. The proxy handles load balancing, retries, timeouts, mTLS, and metrics collection uniformly, without any code in the service itself. The app just talks to localhost; the sidecar does the rest. This decouples network behavior from application logic and lets platform teams change routing/security policy without redeploying services.

go deeper

for a junior

Should describe the sidecar as a helper process per service and know it intercepts traffic without app code changes; doesn't need iptables internals.

for a middle

Should know traffic redirection mechanics (iptables/CNI) and name at least one concrete proxy (Envoy or linkerd2-proxy).

for a senior

Should discuss resource/latency overhead, startup ordering issues, and debugging implications of the extra hop.

for a principal

Should weigh sidecar footprint against fleet size/density when advising on mesh adoption and compare against sidecar-less alternatives.

## The problem it solves The sidecar pattern solves a specific packaging problem: how do you attach cross-cutting network behavior — **retries, encryption, load balancing, telemetry** — to every service in a fleet without baking a library into each service's codebase in every language teams happen to use? ## How traffic actually flows Mechanism: in Kubernetes, sidecar injection (a mutating webhook in Istio, `linkerd inject` in Linkerd) adds a second container to each application pod running the proxy binary. An init container (or, in newer versions, a CNI plugin) rewrites the pod's `iptables` rules so every outbound connection from the app container is transparently redirected to the sidecar's outbound listener port, and every inbound connection first passes through the sidecar's inbound listener before reaching the app. The application is unaware this redirection is happening — it opens a socket to what it believes is the destination service and the kernel silently reroutes the packets to a local port (Envoy's default outbound port in Istio is 15001). The sidecar then does the real work: 1. It **looks up the destination** in its locally cached configuration (pushed down from the control plane). 2. It **picks a healthy upstream endpoint** via its load-balancing algorithm. 3. It **opens the actual connection** (upgrading to mTLS if configured). 4. It **applies any retry/timeout/circuit-breaking policy**, and forwards the request. 5. On the way back it **emits metrics** (request count, latency histograms, response codes) scraped by Prometheus or exported via OpenTelemetry. ## Where it came from Why it exists: before service meshes, this network logic lived in language-specific client libraries — hand-rolled retry loops or resilience libraries in whatever language a service happened to be written in. That meant: - every language runtime a company used needed its own maintained library; - upgrades required redeploying every service; - and there was **no consistent way** to observe or enforce policy across the whole fleet. Moving that logic into an out-of-process sidecar makes it **language-agnostic** (any service, in any language, that speaks HTTP or gRPC gets the same behavior) and **centrally upgradable** — you roll a new Envoy image without touching application code. ## What it costs Trade-offs: the sidecar model buys uniformity and separation of concerns at the cost of resource duplication and operational complexity. - **Resource duplication.** Every pod now runs at least two containers, so a fleet of a thousand small services duplicates the sidecar's baseline CPU/memory footprint a thousand times — Envoy alone commonly needs tens of megabytes of memory and a non-trivial CPU reservation per pod, which for very small, high-density services can dwarf the resource cost of the actual application. - **There's also an added network hop.** A call from A to B now goes app to sidecar A, over the network, to sidecar B, to app, adding latency (typically low single-digit milliseconds per hop, but it compounds across deep call chains) and additional failure surface. - **Debugging gets harder.** A request timeout could originate in the app, in either sidecar's config, or on the wire, and engineers unfamiliar with the mesh often waste time looking in application logs for a problem that's actually a misconfigured retry policy or connection-pool limit in the sidecar. ## What goes wrong Failure modes in production: - **The most common is the sidecar not being ready yet.** If the application container starts before its sidecar has finished pulling policy from the control plane, outbound calls fail during pod startup; this is why Istio and Linkerd both support holding application start until the proxy is ready. - **Another is mismatched injection.** A namespace or pod that skips sidecar injection (a missing label, or a deploy that lacks the annotation) silently drops out of the mesh, meaning it gets no mTLS and no policy enforcement, which is a subtle security gap, not just an availability one. - **A third is sidecar resource starvation under load.** Because the sidecar sits in the hot path, if its CPU is throttled by a tight Kubernetes limit, every request through that pod slows down even though the application itself is healthy. ## Where it shows up Concrete example: Linkerd (built by Buoyant) deliberately ships an ultra-lightweight Rust-based proxy specifically to minimize per-pod overhead, positioning itself as the lower-footprint alternative to Istio's Envoy-based sidecar — the same proxy technology Lyft originally built and open-sourced to solve exactly this problem at their own microservice scale. Both prove the pattern works at scale, but both carry the same fundamental cost: - one extra process per pod, - one extra network hop per call, - and a new layer to understand when something goes wrong.

  • What happens to a pod's outbound traffic if the sidecar container crashes but the application container keeps running?
    Because iptables rules force all traffic through the sidecar's ports, the application effectively loses network connectivity — outbound calls hang or refuse until the sidecar restarts. This is why readiness/liveness probes for meshed pods are configured to consider sidecar health, and why some setups fix startup ordering so the app doesn't start before the sidecar is ready.
  • Why do many teams choose Linkerd over Istio specifically for the sidecar's resource footprint?
    Linkerd's proxy is written in Rust and deliberately minimal, aiming for a much smaller memory/CPU baseline per pod than Envoy, a general-purpose C++ proxy with a much larger feature surface. For fleets with thousands of small pods, that per-pod delta adds up to meaningful cluster-wide cost.
  • How does sidecar injection actually happen in Kubernetes without editing every deployment manifest?
    A mutating admission webhook intercepts pod creation requests for namespaces labeled for injection and rewrites the pod spec on the fly to add the sidecar container plus an init container or CNI hook for traffic redirection and shared volumes. Developers don't touch their deployment YAML at all.

Like giving every employee a personal assistant who intercepts all their calls and mail, handles security checks and routing, while the employee just talks normally without ever seeing the assistant do it.

saying these in an interview costs you the question

  • thinks the mesh requires rewriting application code to call a library
  • doesn't know traffic is redirected via iptables/eBPF, thinks it's DNS-based
  • assumes sidecar adds zero latency or resource cost
  • can't explain why a non-injected pod silently loses mTLS
  • confuses the sidecar with an API gateway that sits at the cluster edge

context

open as a page

In Istio's architecture, what specifically does the control plane (istiod) do versus what the data plane (the Envoy sidecars) does, and how does configuration get from one to the other?

level: middleimportance: must knowfreq 65%

basics

~10 s

The control plane is the 'brain' that decides the rules for routing and security; the data plane is all the sidecar proxies that actually carry traffic following those rules.

open as a page

How does a service mesh like Istio establish mutual TLS (mTLS) between two services automatically, without either service's code doing any TLS handshake itself, and what is PERMISSIVE mode for?

level: middleimportance: must knowfreq 60%

basics

~10 s

The sidecars on both ends do the encryption handshake for the services automatically, using certificates the mesh hands out and rotates — the app code just sends plain traffic to its own sidecar.

open as a page

When a platform team configures retry and timeout policy in a service mesh's data plane (e.g. Istio's VirtualService/DestinationRule) instead of in each service's application code, what do they gain, what do they give up, and where can this go wrong?

level: seniorimportance: must knowfreq 55%

basics

~20 s

Moving retries/timeouts into the mesh means one shared config controls every service's behavior instead of each team writing its own — easier to standardize, but harder for a single service to fine-tune its own special cases, and blind retries can make an outage worse.

open as a page

A service mesh is often sold on giving you 'free' observability — uniform metrics, distributed tracing, and access logs for every service without code changes. What exactly does the mesh actually provide here, what's the catch, and what production failure mode commonly trips up teams relying on it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

The mesh's sidecars automatically record how long calls took and whether they succeeded, for every service, without any app code — but they can't fully trace a request's journey unless the app helps pass along a few tracing headers.

open as a page

A principal engineer is asked whether a 15-service platform running on Kubernetes should adopt a sidecar-based service mesh like Istio. What factors would make this a bad idea, and what does a 'sidecar-less' or 'ambient' mesh architecture change about that calculus?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

For a small number of services, a full mesh can cost more in complexity and per-pod overhead than the traffic-management and mTLS benefits are worth; newer ambient-mesh designs try to get similar benefits without a sidecar in every pod.

open as a page