skip to content

How does a service mesh like Istio establish mutual TLS (mTLS) between two services automatically, without either service's code doing any TLS handshake itself, and what is PERMISSIVE mode for?

level: middleimportance: must knowfreq 60%

answer

  1. SPIFFE identity from k8s service account
  2. sidecar-to-sidecar TLS, app sees plaintext
  3. istiod CA issues short-lived certs, auto-rotates
  4. STRICT vs PERMISSIVE PeerAuthentication
  5. PERMISSIVE = migration path, accepts both

basics

~10 s

The sidecars on both ends do the encryption handshake for the services automatically, using certificates the mesh hands out and rotates — the app code just sends plain traffic to its own sidecar.

solid answer

~40 s

Each workload's sidecar is issued a short-lived X.509 certificate by the mesh's built-in certificate authority (istiod's CA in Istio), tied to an identity derived from its Kubernetes service account. When service A calls service B, A's sidecar intercepts the outbound plaintext call, upgrades the connection to mTLS using its own cert, and B's sidecar terminates that mTLS connection, verifies A's certificate against the mesh's trust root, and forwards the request to B's app container as plain HTTP. Neither app ever handles a certificate. Istio supports STRICT mode, which requires mTLS mesh-wide, and PERMISSIVE mode, which accepts both mTLS and plaintext on the same port — used as a migration path so non-meshed clients aren't broken while a mesh rolls out.

go deeper

for a junior

Should know the mesh encrypts service-to-service traffic automatically and that app code isn't handling certs.

for a middle

Should describe the sidecar-to-sidecar handshake and know a CA issues rotating certs.

for a senior

Should explain STRICT vs PERMISSIVE, the SPIFFE-style identity model, and the migration rationale for PERMISSIVE.

for a principal

Should reason about CA-availability blast radius, governance risk of a shared trust root, and catching stale PERMISSIVE namespaces as a security audit concern.

## The two problems it solves Mutual TLS in a service mesh solves two problems at once: - **encrypting traffic** between services so anyone sniffing the network can't read it; - **authenticating both ends** of every connection so a service can cryptographically prove which other service is calling it, not just trust the network. Doing this at the application layer historically meant every service team had to manage its own TLS certs, rotate them before expiry, and implement client and server TLS correctly in every language — error-prone and inconsistent. Pushing it into the sidecar makes it uniform, automatic, and invisible to application code. ## How the handshake happens Mechanism, step by step: 1. **Every workload in the mesh has an identity** — not a person-style identity but a machine identity following the SPIFFE standard (a URI derived directly from its Kubernetes service account and namespace). 2. **The sidecar requests a certificate.** When a pod starts, its sidecar requests a certificate for that identity from the mesh's certificate authority — istiod's built-in CA in Istio, though it can be swapped for an external CA like Vault or cert-manager. 3. **The CA validates and issues.** The CA validates the request, typically using the pod's bound service account token as proof of identity, and issues a short-lived X.509 certificate, commonly valid for 24 hours by default, which the sidecar caches and automatically rotates well before expiry without any operator intervention or application restart. 4. **The calling sidecar opens a TLS connection.** When service A's sidecar sees an outbound connection destined for service B, instead of forwarding plaintext, it opens a TLS connection to B's sidecar and presents its own certificate. 5. **The receiving sidecar verifies.** B's sidecar, playing the server role in that handshake, requests and verifies A's certificate against the mesh's shared trust root — this is the 'mutual' part, both sides authenticate, not just the client verifying the server as in normal one-way HTTPS. 6. **The app sees plain HTTP.** Once the handshake succeeds, B's sidecar decrypts the request and forwards it to B's application container as plain HTTP on localhost — the app never sees a certificate, a TLS library, or even knows encryption happened. ## STRICT and PERMISSIVE modes Because rolling this out fleet-wide instantly would break any traffic still coming from non-meshed sources — a legacy service without a sidecar, a health-check probe, an external client — Istio supports two modes controlled via `PeerAuthentication`. | Mode | What the sidecar does | |---|---| | **STRICT** | rejects any connection that isn't mTLS, appropriate once every client of a workload is confirmed to be in the mesh | | **PERMISSIVE** | the default during a mesh rollout, makes the sidecar accept both mTLS and plaintext connections on the same port, auto-detecting which protocol an incoming connection uses | This lets operators enable the mesh namespace-by-namespace without a flag-day cutover — services already migrated get automatic encryption for calls from other meshed services, while calls from not-yet-migrated clients keep working over plaintext, and only once everything is confirmed meshed does the team flip a namespace to STRICT. ## What it costs Trade-offs: automatic mTLS removes an enormous amount of per-service TLS management toil and enforces a consistent security posture, but it isn't free. - **A CPU and latency cost per connection.** There's a real, usually small, cost for the TLS handshake and ongoing encryption/decryption on both sidecars, which matters more for very high-throughput, latency-sensitive paths. - **A hard dependency on the CA.** It also introduces a hard dependency on the mesh's CA staying available for certificate issuance and rotation — if that path breaks, existing certs keep working until they expire, but no new workload can join the mesh with a valid identity. - **A governance dimension.** Because trust flows entirely through the mesh's CA and its root of trust, compromising that CA or a misconfigured `PeerAuthentication/AuthorizationPolicy` that's more permissive than intended is a fleet-wide security risk, not a single-service one. ## What goes wrong Failure modes seen in production: - **The most common is a namespace accidentally left in PERMISSIVE mode** long after migration was believed complete, silently allowing an unauthenticated plaintext caller to reach a service operators assumed was mTLS-only — a security gap that doesn't show up as an outage, so it's easy to miss without an explicit audit. - **Another is certificate expiry during a control-plane outage.** If istiod's CA is down for longer than a workload's cert TTL, that workload can't rotate and its mTLS connections start failing once the old cert expires, presenting as a sudden wave of failures that looks unrelated to the actual root cause. - **A third is mixed-mode confusion during migration** — a team enables STRICT too early on a namespace still receiving traffic from an unmeshed legacy service, instantly cutting that traffic off. ## Where it shows up A concrete real-world example of exactly this rollout pattern: Istio's own documented migration path explicitly recommends starting new installations in PERMISSIVE mode mesh-wide and only tightening individual namespaces to STRICT once their traffic sources are confirmed fully meshed, precisely to avoid the flag-day breakage described above.

  • Why is a workload's mTLS identity tied to its Kubernetes service account rather than, say, its IP address or pod name?
    IPs and pod names are ephemeral and get reused constantly as pods are rescheduled, so they make a poor basis for a stable identity to authenticate against. A service account is a durable, meaningful identity that maps to which application or team a workload belongs to, and Kubernetes already issues bound tokens for it, giving the mesh CA a trustworthy way to verify the request without inventing a separate identity system.
  • What happens to already-established mTLS connections if the mesh's certificate authority becomes unavailable for an hour, assuming certs are valid for 24 hours?
    Nothing breaks immediately — existing certificates remain cryptographically valid and sidecars keep using them for both new and existing connections without needing to contact the CA again. The risk only materializes for workloads whose certs are close to their 24-hour expiry during that outage window, or for brand-new workloads trying to join the mesh, since they can't get a cert issued while the CA is down.
  • Why might a team deliberately leave a namespace in PERMISSIVE mode indefinitely instead of moving to STRICT?
    If that namespace still needs to accept legitimate traffic from clients outside the mesh — health checks from infrastructure that isn't sidecar-injected, or a legacy service being migrated slowly — STRICT would break that traffic outright. PERMISSIVE lets them keep receiving both mTLS from meshed callers and plaintext from those unmeshed sources simultaneously, accepting a known, deliberate security trade-off rather than an oversight.

Like a building where security guards at every door check each other's badges and escort visitors through, while the people inside offices never see a badge check at all.

saying these in an interview costs you the question

  • thinks the application code performs the TLS handshake
  • doesn't know PERMISSIVE mode exists and thinks the mesh is binary encrypted-or-not
  • can't explain what makes it 'mutual' vs regular one-way HTTPS
  • assumes certs never need rotation once issued
  • thinks mTLS identity comes from IP address

context