skip to content

Consul's service mesh issues each sidecar its own certificate. What does that certificate carry, how long is it valid by default, and what happens across the fleet when you rotate the mesh root or move the CA to the Vault provider?

level: seniorimportance: should knowfreq 35%

answer

  1. SPIFFE identity in the certificate's URI SAN
  2. leaves measured in hours, not months
  3. 72 hours is the default
  4. cross-signing is what makes rotation seamless
  5. expired Vault token fails slowly, not instantly

basics

~20 s

Each sidecar gets a short-lived leaf certificate whose URI SAN is a SPIFFE identity naming the service, valid for LeafCertTTL — 72 hours by default — and renewed automatically. Rotating the root signs a new intermediate and re-issues leaves; cross-signing is what keeps connections working during the rollover.

solid answer

~60 s

Consul runs its own mesh CA. Every sidecar holds a leaf certificate whose URI SAN is a SPIFFE-style identity — trust domain, namespace, datacenter and service name — which is exactly what intentions authorize on. Leaves are deliberately short-lived: `LeafCertTTL` defaults to 72 hours, and the agent or dataplane renews well before expiry, so there is no manual certificate management. Rotating the root or changing CA configuration (`consul connect ca set-config`, or `/v1/connect/ca/configuration`) makes Consul generate a new intermediate and re-issue leaves across the fleet. Where the provider supports **cross-signing**, the new root is signed by the old one so certificates from both generations validate and the rollover is seamless; where it does not, there is a window in which peers holding old material cannot verify new certificates. Moving to the Vault provider means Consul stops holding the private key and instead uses configured `RootPKIPath` and `IntermediatePKIPath` mounts. The classic outage is the Vault token expiring: signing stops, and connections fail *gradually* as existing leaves age out over the next `LeafCertTTL`.

code

bash · 3 lines
bash
consul connect ca get-config

consul connect ca set-config -config-file=vault-ca.json

go deeper

for a junior

Know that Consul issues each sidecar its own short-lived certificate automatically, and that the service's identity is carried inside that certificate rather than configured by hand.

for a middle

Explain the SPIFFE URI SAN, the default 72-hour LeafCertTTL with automatic renewal, and that changing CA configuration triggers a new intermediate and fleet-wide re-issue.

for a senior

Talk through a rotation as an operation: cross-signing, the convergence window bounded by the leaf TTL, and the Vault-credential failure that surfaces hours late as a slow cascade of handshake errors.

for a principal

Own the PKI decision — Consul's own root versus an existing Vault PKI, who audits it, how short a leaf lifetime the control plane can sustain, and what monitoring makes a signing failure visible before it becomes an outage.

## What each proxy actually holds Consul's mesh CA issues every sidecar an X.509 leaf certificate. The distinguishing feature is the URI Subject Alternative Name, a SPIFFE-style identifier of the shape: ``` spiffe://<trust-domain>.consul/ns/<namespace>/dc/<datacenter>/svc/<service> ``` Every mesh connection is mutual TLS, so both ends present one, and the destination proxy reads the service name straight out of the verified peer certificate. This is the substrate everything else in the mesh rests on: intentions authorize that name, metrics are labelled with it, and a workload that cannot obtain a certificate cannot participate in the mesh at all. ## Short lifetimes, automatic renewal `LeafCertTTL` defaults to **72 hours**. Consul renews each leaf well before it expires, so operators never touch certificates by hand. The short lifetime is deliberate: it bounds the value of a stolen key and removes the need for a working revocation path, since a compromised identity stops being usable within a bounded window. Shortening it is not free. Every renewal is a signing request against the Consul servers, and signing is rate-limited — the CA configuration exposes `CSRMaxPerSecond` (50 by default) and `CSRMaxConcurrent` precisely because a large mesh renewing aggressively can otherwise saturate the leader. Halving the TTL doubles steady-state signing load, and a mass simultaneous renewal (say, after everything restarted together) can push against those limits. ## Rotating the root CA configuration is changed through `consul connect ca set-config` or the `/v1/connect/ca/configuration` endpoint. When the root changes, Consul generates a new intermediate, begins signing new leaves under it, and pushes the updated trust bundle out. The risk during rollover is obvious: for a period, some proxies hold certificates chaining to the old root while their peers have been told to trust the new one. **Cross-signing** is what closes that gap. Where the provider can do it, the new root is signed by the old one, so a peer trusting either root can validate both generations, and the fleet converges without dropped connections. Where the provider cannot cross-sign, you get a genuine window in which mismatched pairs fail to handshake — which is why a provider change is an operation you schedule and stage rather than one you do casually at peak traffic. A useful mental check when planning a rotation: the fleet is fully converged only once every proxy has renewed at least once, and the worst case for that is one full `LeafCertTTL`. Anything you do that depends on "everyone has the new material" has that as its floor. ## The built-in provider versus Vault The **built-in** provider keeps the root private key in Consul's own state. It is the zero-setup option and it is fine for many environments, but it means the key's security is Consul's security, and the mesh's trust root is not visible to whatever else manages your PKI. The **Vault** provider hands root and intermediate signing to Vault PKI mounts configured with `RootPKIPath` and `IntermediatePKIPath`, using a token or auth method with policy over those paths. Consul no longer holds the root key; Vault does, with its own audit trail, and the same PKI can serve non-mesh consumers. The operational failure mode that follows is the one worth being able to describe in an interview: **the Vault credential expiring or losing its policy**. Nothing breaks at that moment. Existing leaves are still valid, so traffic flows and dashboards stay green. Then, over the following hours, renewals begin failing one workload at a time, and connections start dropping in a slow, seemingly random pattern until the whole mesh is down — up to `LeafCertTTL` after the actual cause. Anyone debugging from the symptom timestamps looks in entirely the wrong place. The defences are to alert on CA signing errors and on leaf-renewal failures rather than only on request errors, to use a renewable auth method rather than a static long-lived token, and to treat "can Consul still sign" as a first-class health signal. ## What this leaves you to decide Whether the trust root lives in Consul or in an existing PKI; how short a leaf lifetime you can afford against the signing load it creates; and whether your provider choice supports cross-signing, because that single property determines whether a future rotation is a non-event or a scheduled maintenance window.

  • Why does an expired Vault token produce a gradual outage rather than an immediate one?
    Because nothing revokes the certificates already in use. Signing fails at the moment the credential dies, but every proxy keeps serving on a leaf that is still valid. Failures appear one workload at a time as each leaf reaches renewal, spread over up to `LeafCertTTL`. That lag is why you should alert on signing and renewal errors directly, not on downstream request failures.
  • What is the cost of lowering LeafCertTTL from 72 hours to a few hours?
    Signing load scales inversely with the TTL, and the CA is rate-limited by `CSRMaxPerSecond` and `CSRMaxConcurrent`. A much shorter TTL means far more CSRs against the leader, and it makes any mass-restart event more likely to bunch renewals together. You buy a tighter compromise window and pay in steady-state control-plane load.
  • What does cross-signing actually buy you during a root rotation?
    It lets certificates from the old and new generations validate against each other while the fleet converges. The new root is signed by the old one, so a peer that trusts either can verify both, and no connection fails purely because the two ends rotated at different times. Without it, a rotation has a real failure window that must be planned around.

saying these in an interview costs you the question

  • Thinks mesh certificates are long-lived and rotated by hand
  • Says a failed CA causes an instant, total outage
  • Believes the mesh relies on CRLs or OCSP for revocation
  • Assumes switching CA providers is always seamless
  • Confuses the certificate identity with the source IP address

context