skip to content

In an Istio mesh, where does a workload's mutual-TLS certificate come from, what identity does that certificate carry, and how is it kept fresh?

level: middleimportance: should knowfreq 50%

answer

  1. issued per pod, not per node
  2. the request is authenticated by a platform token
  3. identity is coarser than the pod
  4. short lifetime instead of revocation
  5. namespace and service account in a URI

basics

~20 s

The agent in each pod generates a key and certificate request, authenticates to istiod's built-in CA with the pod's service-account token, and receives a short-lived certificate whose SPIFFE identity encodes the trust domain, namespace and service account. The agent rotates it automatically.

solid answer

~50 s

Each pod's istio-agent generates a private key locally and sends a certificate signing request to istiod, which runs the mesh CA. The agent proves who it is with the pod's projected service-account token, and istiod validates that token against the Kubernetes API before signing. The certificate it returns carries a SPIFFE identity as a URI subject alternative name, shaped `spiffe://<trust-domain>/ns/<namespace>/sa/<serviceaccount>` — so the unit of identity in the mesh is the **service account**, not the pod, the Deployment or the Service. The agent hands the key and certificate to the proxy over the local secret discovery API, so the private key never leaves the pod and is never written into a Kubernetes Secret. Certificates are short-lived — 24 hours by default — and the agent renews them well before expiry, typically around half of their lifetime.

go deeper

for a junior

Know that each pod gets its own short-lived certificate from the control plane automatically, and that the identity inside it is built from the namespace and service account.

for a middle

Walk the issuance flow end to end — local key generation, a signing request authenticated by the pod's service-account token, validation by istiod, delivery to the proxy — and state the SPIFFE identity format.

for a senior

Draw the operational consequences: identity granularity equals service-account granularity, short lifetimes replace revocation, and a control-plane outage degrades new joins long before it breaks established traffic.

for a principal

Own the trust-domain and root-of-trust decisions — self-signed per cluster versus a plugged-in intermediate chained to the organisation's hierarchy — and what identity means across cluster or mesh boundaries.

## Why this matters before any policy does Every identity-based rule you can write in the mesh — a principal in an authorization policy, a peer shown in telemetry — is only as meaningful as the certificate underneath it. Understanding the issuance flow tells you what an identity actually asserts and, more importantly, what it does not. ## The issuance flow 1. **Key generation, in the pod.** The istio-agent process running alongside the proxy generates a private key inside the pod's memory. Nothing external ever sees it. 2. **A certificate signing request goes to istiod.** Istiod embeds the mesh certificate authority. The agent authenticates the request using the pod's projected service-account token, mounted by the platform with an audience scoped for this purpose. 3. **Istiod validates the token** against the Kubernetes API — confirming that this token really belongs to a live service account — and only then signs. 4. **The signed certificate comes back** and the agent serves it to the proxy over a local secret discovery service on a Unix domain socket. The proxy loads it into its TLS transport configuration and starts using it for both inbound and outbound mesh connections. Two properties fall out of this design that are worth stating explicitly. The private key **never crosses the pod boundary**, so there is no Secret object holding workload keys to leak, and a compromise of one pod does not yield another pod's key. And issuance is **online and continuous**, not a bootstrap step — a workload that cannot reach istiod will keep serving with its current certificate until it expires, then start failing. ## What the identity actually says The identity is encoded as a URI subject alternative name in the SPIFFE format: ``` spiffe://cluster.local/ns/payments/sa/ledger ``` `cluster.local` is the mesh's trust domain, configurable at install time and important once you federate meshes. The remainder is the namespace and the service account. In authorization rules the same identity is written **without the scheme**: ```yaml from: - source: principals: ["cluster.local/ns/payments/sa/ledger"] ``` The consequence that catches teams out: **the granularity of mesh identity is exactly the granularity of your service accounts.** If five Deployments in a namespace all run under the namespace's default service account, they share one identity and are completely indistinguishable to any policy you write. Introducing identity-based authorization to an existing cluster therefore usually begins with a boring prerequisite — give every workload its own service account — long before anyone writes a policy. Equally, the identity says nothing about *which Service* was addressed; it identifies the caller, not the route. ## Rotation Workload certificates are deliberately short-lived — a 24-hour lifetime is the default — and the agent renews well before expiry, conventionally around half of the remaining lifetime, so that a temporary control-plane outage has hours of headroom rather than minutes. Because the swap happens through the secret discovery API, the proxy picks up the new material without restarting and without dropping connections, and the application is never involved. Short lifetimes are the mesh's substitute for revocation: rather than distributing revocation information, an identity that should no longer exist simply stops being reissued. The root of trust is separate. By default istiod generates a self-signed root and holds it, which is convenient and appropriate for a single cluster. For anything beyond that — multi-cluster meshes, or an organisation that wants mesh identities chained to its own hierarchy — you plug in an intermediate by creating a secret named `cacerts` in the istiod namespace containing the certificate, key, root certificate and chain. Rotating that root is a genuine operational exercise, because every workload's chain must remain verifiable throughout. ## Diagnosing it When mesh TLS is failing, the questions in order are: can the agent reach istiod, is the certificate present and unexpired in the proxy's secret configuration, and does the identity in it match the service account you expected. The last one is where most surprises live — a workload running under an inherited default service account is the usual reason a carefully written principal rule never matches.

  • Two Deployments in the same namespace run under the same service account. What does that mean for authorization policies?
    They share one mesh identity and are indistinguishable. Any principal-based rule that admits one admits the other, so the policy cannot express a boundary between them. Fixing it means giving each workload its own service account and reissuing — which is why service-account hygiene is a prerequisite for identity-based authorization, not a follow-up.
  • Why does Istio use short-lived certificates instead of publishing revocation information?
    Because revocation distribution is the weak part of the model and clients often fail open. With a 24-hour default lifetime and automatic renewal, an identity that should disappear simply stops being reissued, and the window of exposure is bounded by the lifetime. It also keeps the data plane free of revocation lookups on the request path.
  • What happens to running workloads if istiod becomes unavailable for an hour?
    Existing traffic continues, because proxies keep serving with the certificates and configuration they already hold. What stops is anything that needs the control plane: new workloads cannot obtain a certificate and join the mesh, and configuration changes do not propagate. Only if the outage outlasts certificate lifetimes does established traffic start failing.

saying these in an interview costs you the question

  • Thinks workload private keys are stored in Kubernetes Secrets
  • Believes the identity distinguishes individual pods or Deployments
  • Says certificates are issued once at install and last for a year
  • Assumes rotation requires restarting the pod or the application
  • Thinks the certificate identifies the service that was called

context