Linkerd's automatic mTLS is built on a trust anchor certificate and an issuer certificate. Walk through how a meshed pod obtains its identity, and explain what happens across the cluster when the issuer certificate — or the trust anchor itself — expires.
answer
- three tiers: anchor, issuer, leaf
- identity comes from the ServiceAccount token
- leaves are short-lived and auto-renewed
- the anchor is baked into every pod spec
- failure arrives hours after the expiry
basics
~20 sEach proxy gets a short-lived leaf certificate from Linkerd's identity service, proving its Kubernetes ServiceAccount identity and signed by the issuer under the trust anchor. If the issuer or the anchor expires, issuance stops and mTLS fails across the whole mesh.
solid answer
~50 sOn startup the `linkerd2-proxy` generates a private key in memory, builds a CSR, and calls the control plane's `identity` service authenticating with a projected ServiceAccount token. The identity service validates that token against the Kubernetes API and returns a short-lived leaf certificate — the default issuance lifetime is 24 hours — whose identity is derived from the pod's ServiceAccount and namespace. The proxy renews it well before expiry. That leaf is signed by the **issuer** certificate, which is itself signed by the **trust anchor** whose PEM is baked into every proxy at injection time. If the issuer expires, the identity service can no longer issue: existing pods keep working until their 24-hour leaf lapses, then mTLS fails mesh-wide, and any pod that restarts fails immediately. An expired trust anchor is worse, because rotating it means restarting every meshed pod. Automate the issuer with cert-manager and alert on both with `linkerd check --proxy`.
go deeper
Know that Linkerd encrypts pod-to-pod traffic automatically and that each proxy holds a certificate identifying its workload. Be able to say the certificates are Linkerd's own, not public web certificates.
Explain the anchor/issuer/leaf hierarchy, that identity derives from the pod's ServiceAccount token validated against the Kubernetes API, and that leaves are short-lived and auto-renewed.
Diagnose the delayed outage: name why an expired issuer breaks the mesh hours later, why an anchor mismatch breaks it immediately, and how you would confirm each with linkerd check --proxy.
Own the PKI lifecycle as a platform commitment: automated issuer rotation, an alarm on the anchor with months of lead time, a rehearsed bundled-anchor rotation, and a clear decision on whether the mesh's trust domain nests under a corporate CA.
## The three-level certificate hierarchy Linkerd's mTLS rests on three tiers, and confusing them is the root of most mesh certificate outages. 1. **Trust anchor** — the root of the mesh's own PKI. Its public certificate is distributed to *every* proxy: the injector writes it into the pod spec at injection time (in Helm terms, `identityTrustAnchorsPEM`). Every proxy validates its peers against this anchor. 2. **Issuer** — an intermediate certificate and key held by the control plane's `identity` component. It is signed by the trust anchor and is what actually signs workload certificates. 3. **Leaf** — a short-lived per-proxy certificate. The default issuance lifetime is 24 hours. This is a private PKI with a trust domain of its own (`cluster.local` by default); it has nothing to do with the public web PKI, and the certificates are not for browsers. ## How a pod gets its identity When an injected pod starts: 1. The proxy generates a private key **in memory**. The key never touches disk and never leaves the pod. 2. It constructs a CSR whose identity is derived from the pod's Kubernetes ServiceAccount and namespace — of the form `<serviceaccount>.<namespace>.serviceaccount.identity.linkerd.cluster.local`. 3. It calls the `identity` service, presenting a **projected ServiceAccount token** as proof of who it is. 4. `identity` validates that token against the Kubernetes API server. Kubernetes, not Linkerd, is the source of truth for workload identity — that is the elegant part of the design. 5. `identity` signs and returns a leaf certificate with the default 24-hour lifetime; the proxy refreshes it well before expiry. The consequence is worth stating explicitly in an interview: **workload identity is ServiceAccount identity**. Two Deployments sharing one ServiceAccount are indistinguishable to the mesh, so if you intend to write identity-based authorization policy, give workloads distinct ServiceAccounts first. ```bash linkerd check --proxy # includes certificate expiry checks linkerd viz edges deploy -n app # shows the peer identity on each edge ``` ## What expiry does **Issuer expiry.** The identity service can no longer sign. Nothing breaks at the instant of expiry — running proxies still hold valid leaves. The failure arrives on a delay: any pod that starts, restarts or is rescheduled cannot get a certificate and fails to establish mTLS, and within the leaf lifetime (24 hours by default) every remaining proxy's certificate lapses too. The symptom is a mesh-wide connection failure that appears to have no trigger, hours after the actual event, and often overnight. This is exactly why it is a favourite interview scenario: the cause and the symptom are separated in time. **Trust anchor expiry.** Worse, because the anchor's public certificate is embedded in every pod spec. Replacing it is not a control-plane-only operation — you must roll every meshed workload so the new anchor reaches every proxy. The safe procedure is to *bundle*: create the new anchor, install a PEM bundle containing both the old and the new anchor so proxies trust either, roll the workloads, issue a new issuer signed by the new anchor, roll again, then drop the old anchor from the bundle. Attempting a hard cutover instead partitions the mesh into pods that trust the old anchor and pods that trust the new one, and they cannot talk to each other. ## How to not be in this position - **Automate the issuer.** The documented approach is cert-manager: it holds the trust anchor as an issuer, mints the Linkerd issuer certificate into the secret the identity service reads, and rotates it on a schedule. With that in place, issuer expiry stops being an event. - **Give the anchor a long life and an alarm.** A trust anchor is typically valid for years. Set a calendar reminder *and* an alert — `linkerd check --proxy` reports certificate expiry well ahead of time, so run it on a schedule in CI or a cron job and fail loudly, rather than reading it only when something is already broken. - **Practise the anchor rotation** on a non-production cluster before you need it. The bundling procedure is the part teams get wrong under pressure. ## The distinguishing insight A good answer separates *cannot issue new certificates* from *cannot validate existing ones*. Issuer expiry is the first: a slow-motion failure with a delay equal to the leaf lifetime. Anchor mismatch is the second: an immediate, symmetric failure between any two pods that do not share a trust anchor. Naming which one produces which symptom is what tells an interviewer you have operated this rather than read about it.
- Why does rotating the trust anchor require restarting every meshed pod, when rotating the issuer does not?The anchor's PEM is injected into each pod's spec at injection time, so a proxy learns it once at startup and never re-reads it. The issuer, by contrast, lives only in the control plane's identity component, which reads it from a Secret — replacing it affects newly issued leaves without touching workloads.
- Two Deployments in the same namespace share one ServiceAccount. What does that mean for Linkerd's identity model?They have the same mesh identity and are indistinguishable to any identity-based authorization policy — you cannot allow one and deny the other. Since Linkerd derives identity from the ServiceAccount, giving each workload its own ServiceAccount is a prerequisite for meaningful policy, not a cosmetic detail.
- You need to rotate an expiring trust anchor with no downtime. What is the sequence?Bundle, don't cut over. Generate the new anchor, update the trust anchor value to a PEM containing both old and new so every proxy trusts either, roll all meshed workloads, issue a new issuer signed by the new anchor, roll again, then remove the old anchor from the bundle. A direct swap splits the mesh into two mutually untrusting halves.
saying these in an interview costs you the question
- Thinks Linkerd uses public CA certificates for mesh mTLS
- Assumes an expired issuer breaks traffic instantly
- Believes rotating the trust anchor is a control-plane-only change
- Says identity comes from the pod name or IP rather than the ServiceAccount
- Treats certificate expiry as something the mesh silently handles itself