skip to content

Linkerd

Linkerd is the deliberately small service mesh: a purpose-built Rust micro-proxy, automatic mTLS between meshed pods, and golden metrics (success rate, RPS, latency percentiles) with almost no configuration. Interviewers reach for it as the explicit contrast to Istio — I am expected to argue what the narrower feature set buys in memory footprint, latency overhead and operational risk, and where it runs out.

on this pageshow

questions

5

Linkerd ships its own data-plane sidecar, linkerd2-proxy, written in Rust, instead of reusing Envoy the way most service meshes do. What does that choice buy you operationally, and what do you give up?

level: middleimportance: must knowfreq 70%

answer

  1. one job, not a general-purpose proxy
  2. no GC, no dynamic filter chain
  3. cost is multiplied by every pod
  4. the trade is extensibility, not speed
  5. no Wasm or Lua escape hatch

basics

~20 s

Linkerd's linkerd2-proxy is a Rust sidecar built for one job, so each pod costs far less memory and CPU and exposes a smaller attack surface. The price is extensibility: no Wasm or Lua filters, and a narrower protocol and feature set.

solid answer

~50 s

Envoy is a general-purpose proxy — edge gateway, mesh sidecar, standalone load balancer — configured through a very large API. Linkerd instead wrote `linkerd2-proxy` in Rust for exactly one role: the per-pod sidecar between two Kubernetes workloads. Because it does less, it is dramatically smaller in memory and has no garbage-collected runtime and no dynamic filter chain, so per-pod overhead and tail-latency variance are lower, and there is much less configuration surface to get wrong. Memory safety without a GC also shrinks the class of proxy CVEs you have to chase across every pod in the fleet. What you give up is the extension model: there is no Lua or Wasm filter API, no arbitrary protocol filters, and no raw proxy config escape hatch — non-HTTP traffic is marked with `config.linkerd.io/opaque-ports` and proxied as plain TCP. If your requirement is a custom L7 filter, Linkerd has nowhere to put it.

go deeper

for a junior

Know that Linkerd puts a small proxy container next to each application container, and that this proxy — not your code — handles mTLS, retries and metrics. Be able to say it is written in Rust and is intentionally minimal.

for a middle

Explain the mechanics: no garbage collector, no dynamic filter chain, configuration pushed by the control plane instead of hand-written, and L7 features limited to HTTP/1.x, HTTP/2 and gRPC with other protocols marked opaque.

for a senior

Show you have costed it. Talk about per-pod memory multiplied across the fleet, tail-latency stability, and the operational cost of fleet-wide proxy upgrades driven by CVEs, and name the requirement that would push you off Linkerd.

for a principal

Own the framing that this is a scope decision, not a language decision. Be ready to argue when a platform should standardise on one general-purpose proxy for edge and mesh versus accepting two technologies to keep the per-pod component minimal.

## What "micro-proxy" actually means A service mesh puts a proxy next to every workload instance. That proxy sees every request in and out of the pod, so its cost is multiplied by the number of pods, and its bugs are reachable from anywhere traffic flows. Linkerd's central design decision is that this per-pod proxy is a *different product* from an edge proxy, and should be built as such. Envoy is a general-purpose L4/L7 proxy: it serves as an edge gateway, an API gateway, a standalone load balancer, and a mesh sidecar, and it is configured through a large, dynamic filter-chain and xDS API. Linkerd instead maintains `linkerd2-proxy`, written in Rust on the Tokio/Hyper/Tower stack, that implements only what a sidecar between two Kubernetes workloads needs: transparent HTTP/1.x, HTTP/2 and gRPC proxying, mTLS with automatic identity, load balancing, retries and timeouts, and metrics. Its configuration does not come from a user-authored config file at all — it is pushed by the control plane's `destination` and `identity` services, and shaped by a handful of pod annotations. ## What the narrow scope buys **Footprint.** The proxy's resident memory is small enough that meshing thousands of pods does not visibly change cluster capacity planning; a heavier sidecar can be the single largest per-pod cost in a fleet of small services. Always measure with your own traffic rather than quoting a benchmark, but the direction is not in dispute — that is the point of the design. **Latency behaviour.** Rust has no garbage collector, so there is no per-proxy GC pause contributing to tail latency, and the request path has no dynamically composed filter chain to walk. You still pay two extra hops per call (client sidecar, server sidecar), so a mesh is never free — but the added latency is comparatively stable, which is what matters when you are defending a p99 SLO. **Security surface.** Rust's memory safety removes most buffer-overflow class vulnerabilities from a component that terminates untrusted traffic on every pod. Fewer features also means fewer parsers, and every parser is an attack surface. Fleet-wide proxy upgrades are the most disruptive operation a mesh operator performs, so having fewer of them forced on you by CVEs is a real operational saving. **Less to configure wrong.** There is no equivalent of hand-writing a listener filter chain. The knobs are annotations: ```yaml kind: Namespace metadata: name: payments annotations: linkerd.io/inject: enabled config.linkerd.io/proxy-cpu-limit: "1" config.linkerd.io/opaque-ports: "3306" ``` That is roughly the whole surface for a normal workload. The corresponding Envoy-based configuration is orders of magnitude larger, and most mesh outages are configuration outages. ## What you give up **Extensibility.** There is no Lua filter, no Wasm extension point, and no way to inject arbitrary proxy configuration. If a team needs a custom header-signing filter, a bespoke authentication step, or protocol translation inside the sidecar, Linkerd's answer is "do it in your application, or put a different proxy at the edge" — not "write a filter". **Protocol breadth.** Linkerd's L7 features apply to HTTP/1.x, HTTP/2 and gRPC. Other protocols still get mTLS and TCP-level proxying, but they are declared opaque with `config.linkerd.io/opaque-ports` so the proxy does not try to parse them as HTTP, and you get connection-level metrics rather than per-request ones. **Edge duties.** Linkerd deliberately does not ship a full-featured ingress gateway; the documented pattern is to run a normal ingress controller in front and mesh it. If you wanted one proxy technology for edge and mesh alike, Envoy-based meshes fit that ambition better. ## How to talk about the trade The honest framing is not "Rust is faster than C++" — it is *scope*. Linkerd bought a smaller, more predictable, less configurable component by refusing to be a general-purpose proxy, and pays for it whenever a requirement falls outside the sidecar's job description. When you evaluate it, list the L7 behaviour you actually need at the pod boundary. If that list is mTLS, golden metrics, load balancing, retries and timeouts, the micro-proxy is strictly the cheaper way to get it. If the list contains a custom filter or an exotic protocol, you have found the boundary.

  • If Linkerd's proxy only speaks HTTP, HTTP/2 and gRPC at layer 7, what happens to a meshed pod that talks to a MySQL database?
    It still works. The proxy handles non-HTTP traffic as plain TCP and can still apply mTLS between meshed peers, but you should mark the port with `config.linkerd.io/opaque-ports` so the proxy does not attempt protocol detection on it. You get connection-level metrics rather than per-request success rate and latency.
  • Does a smaller sidecar mean a mesh adds no meaningful latency?
    No. Every meshed call crosses two extra proxies, so you always add hops, connection setup and mTLS work. The claim is that the added latency is small and, more importantly, stable — no GC pauses, no heavy filter chain. You should still measure the delta against your own p99 before meshing latency-critical paths.
  • Linkerd deliberately does not ship a fully featured ingress gateway. How are meshed clusters normally exposed then?
    You run an ordinary ingress controller — nginx, Envoy-based, or a cloud load balancer in front of one — and mesh the controller's pods so traffic entering the cluster is inside the mesh from that point on. Edge concerns like TLS certificates for public hostnames, WAF rules and path rewriting stay in the ingress tier.

saying these in an interview costs you the question

  • Claims Linkerd is just Envoy with a different control plane
  • Says Rust makes it faster, missing that the real trade is scope
  • Assumes you can write Wasm or Lua filters for linkerd2-proxy
  • Thinks a light sidecar means zero added latency
  • Believes Linkerd replaces the ingress controller

context

open as a page

Your platform runs Linkerd today, and a team asks for a mesh capability Linkerd does not have. How do you decide between working around it, putting the capability somewhere else in the stack, or migrating the mesh to Istio?

level: principalimportance: must knowfreq 50%

basics

~20 s

Start from what is genuinely missing. Push edge and gateway concerns into an ingress tier, keep Linkerd for in-cluster mTLS and metrics, and treat migration to Istio as justified only by structural gaps such as non-Kubernetes workloads or per-request token policy.

open as a page

In Linkerd, what are the "golden metrics" that `linkerd viz stat` reports for a meshed workload, where do those numbers come from, and why might a meshed service show no success rate at all?

level: juniorimportance: should knowfreq 62%

basics

~20 s

linkerd viz stat reports success rate, requests per second and latency percentiles, computed from metrics that the linkerd2-proxy sidecars expose and the viz extension's Prometheus scrapes. A workload shows none of them when its traffic is not HTTP or gRPC.

open as a page

Linkerd's automatic mTLS is built on a trust anchor certificate and an issuer certificate. Walk through how a meshed pod obtains its identity, and explain what happens across the cluster when the issuer certificate — or the trust anchor itself — expires.

level: seniorimportance: should knowfreq 55%

basics

~20 s

Each proxy gets a short-lived leaf certificate from Linkerd's identity service, proving its Kubernetes ServiceAccount identity and signed by the issuer under the trust anchor. If the issuer or the anchor expires, issuance stops and mTLS fails across the whole mesh.

open as a page

A Linkerd ServiceProfile lets you mark a route `isRetryable` and configure a `retryBudget` rather than a fixed number of attempts. What does a retry budget actually enforce, and why is it preferred to "retry three times"?

level: middleimportance: nice to knowfreq 40%

basics

~20 s

A retry budget caps retries as a proportion of ordinary requests — Linkerd's default allows roughly 20% extra plus a small floor — so a struggling dependency cannot be hit with a multiple of its normal load the way a fixed per-request retry count allows.

open as a page