Why must a Kubernetes admission webhook be served over TLS, how does the API server decide to trust that server's certificate, and what exactly breaks when the certificate expires?
answer
- HTTPS only; caBundle in the webhook config, not the system trust store
- SAN must be <service>.<namespace>.svc
- expired cert = webhook failure = failurePolicy decides the blast radius
- cert-manager inject-ca-from keeps caBundle in sync
- rotate cert and caBundle together
basics
~20 sThe API server only calls webhooks over HTTPS, and it verifies the server certificate against the caBundle in the webhook configuration, requiring a SAN matching <service>.<namespace>.svc. On expiry the handshake fails, so requests are rejected under failurePolicy: Fail or silently unchecked under Ignore.
solid answer
~50 sAdmission requests carry the full object being created plus the requesting user's identity, and the response can rewrite that object — so the channel must be confidential and, more importantly, authenticated: a plaintext or untrusted endpoint would let anyone impersonate the policy engine and approve whatever they like. The API server therefore always uses HTTPS and verifies the certificate against the **`clientConfig.caBundle`** — a base64-encoded PEM CA bundle stored in the webhook configuration itself, not the cluster trust store. The certificate's SAN must match the name the API server dials: `<service>.<namespace>.svc` for a `service` reference, or the hostname of an external `url`. When the certificate expires (or the CA is rotated without updating `caBundle`), every call fails with an x509 error. That is a webhook *failure*, so `failurePolicy: Fail` starts rejecting all matching API writes, while `Ignore` silently stops enforcing. This is why teams use cert-manager with a CA-injection annotation, or a controller that reissues the Secret and patches `caBundle` together.
code
yaml · 37 linesapiVersion: cert-manager.io/v1
kind: Certificate
metadata:
name: policy-serving-cert
namespace: policy-system
spec:
secretName: policy-serving-cert
duration: 2160h
renewBefore: 720h
dnsNames:
- policy.policy-system.svc
- policy.policy-system.svc.cluster.local
issuerRef:
name: policy-ca-issuer
kind: Issuer
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
name: policy
annotations:
cert-manager.io/inject-ca-from: policy-system/policy-serving-cert
webhooks:
- name: policy.example.com
admissionReviewVersions: ["v1"]
sideEffects: None
clientConfig:
service:
name: policy
namespace: policy-system
path: /validate
# caBundle is written here by the cert-manager CA injector
rules:
- operations: ["CREATE"]
apiGroups: [""]
apiVersions: ["v1"]
resources: ["pods"]go deeper
Know that webhooks are HTTPS-only and that the API server needs the CA in caBundle to trust them.
Explain the SAN requirement for <service>.<namespace>.svc, that caBundle is per webhook configuration, and that expiry manifests as a webhook failure.
Connect expiry to failurePolicy blast radius, describe automated rotation with cert-manager CA injection or operator self-rotation, and say how you alert on expiry before it becomes an outage.
Treat the webhook endpoint as a trusted policy decision point: reason about impersonation risk, whether mutual TLS is warranted, who may edit webhook configurations, and how certificate lifecycle is owned so a bootstrap self-signed cert does not become an anniversary outage.
## Why TLS is not optional An `AdmissionReview` request contains the entire object being admitted — which for a Secret or a Pod with environment variables can include credentials — plus the authenticated identity of the user making the request. The response can rewrite the object arbitrarily via a JSON Patch. So the channel needs two properties: - **Confidentiality**, because the payload can carry secrets; - **Server authentication**, because whoever answers the API server *is* the policy decision point. If an attacker could impersonate the webhook endpoint, they could approve any object a `Fail`-policy webhook was supposed to block, or inject a container into every Pod. The API server enforces this by refusing plain HTTP entirely: `clientConfig.url` must use the `https` scheme, and `clientConfig.service` calls are always HTTPS on the given port (default 443). ## How trust is established The API server does **not** use the host's system trust store, and it does not use the cluster CA automatically. Trust comes from the webhook configuration itself: ``` clientConfig: service: name: policy namespace: policy-system path: /validate port: 443 caBundle: <base64 of one or more PEM CA certificates> ``` If `caBundle` is empty, the API server falls back to its own configured trust roots (`--client-ca-file`-adjacent system roots), which almost never validates a self-signed in-cluster certificate — so in practice you always set it. Two naming rules trip people up: 1. For a `service` reference, the API server dials the in-cluster DNS name **`<name>.<namespace>.svc`**. The serving certificate must carry that exact name in its **Subject Alternative Name** list. A certificate issued only for `policy-system/policy` or for a pod IP will fail hostname verification even if the CA is trusted. Common practice is to include `<name>.<namespace>.svc` and `<name>.<namespace>.svc.cluster.local`. 2. For a `url` reference, the target must be reachable **from the API server's network namespace**, which on managed control planes is not inside your cluster network. That is why in-cluster `service` references are the norm. Note the direction of authentication: this is server-side TLS. The webhook does not, by default, know that the caller is the API server. If you need that, configure the API server with an `AdmissionConfiguration` file (`--admission-control-config-file`) providing a client certificate for the `MutatingAdmissionWebhook`/`ValidatingAdmissionWebhook` plugins, and have the webhook require and verify client certs. Many webhooks skip this and rely on network policy plus the fact that a rogue caller can only ask for an opinion, not change cluster state. ## What expiry actually does The serving certificate typically lives in a Secret mounted into the webhook Pod. On expiry — or when the CA is rotated and the Secret is reissued from a new CA while `caBundle` still holds the old one — the TLS handshake fails and the API server records an error such as `x509: certificate has expired or is not yet valid` or `x509: certificate signed by unknown authority`. Crucially this is indistinguishable, to the API server, from any other webhook failure, so it is governed by `failurePolicy`: - With **`Fail`**, every API write matching the webhook's rules starts erroring. If the rules match Pods cluster-wide, the cluster stops being able to create Pods — including the webhook's own replacements. This is one of the classic self-inflicted cluster outages, and it happens at a predictable time: exactly one year after install, if you generated a 365-day certificate at bootstrap and never automated renewal. - With **`Ignore`**, nothing visibly breaks; the policy silently stops being enforced, which can go unnoticed for weeks. A subtlety: renewing the certificate is only half the job. If the CA changed, the `caBundle` in the `ValidatingWebhookConfiguration`/`MutatingWebhookConfiguration` (and in any CRD `conversion` webhook config) must be updated in the same operation, otherwise you have swapped one x509 error for another. Because the webhook configuration is a cluster-scoped object, that update needs cluster-level RBAC. ## How teams operate it The common patterns are: - **cert-manager**: a `Certificate` issues the serving Secret, and the `cert-manager.io/inject-ca-from` annotation on the webhook configuration makes the CA-injector controller keep `caBundle` in sync automatically. Renewal and bundle update stay coupled. - **Self-managed rotation in the operator**: many operators (controller-runtime has helpers for this) generate a self-signed CA and leaf at startup, write them into a Secret, and patch their own webhook configurations' `caBundle`. Simple, no dependency, but the operator needs `update` on webhook configurations. - **Kubernetes CertificateSigningRequests**: possible via the `kubernetes.io/kubelet-serving`-style flow, but the cluster CA is not exposed for arbitrary server certs in most distributions, so this is the least common route. Whatever you pick, alert on certificate expiry (`certmanager_certificate_expiration_timestamp_seconds`, or a blackbox probe of the webhook endpoint) rather than discovering it through a cluster-wide admission outage.
- The certificate is valid and signed by the CA in caBundle, but calls still fail with an x509 error mentioning the name. What is wrong?Almost certainly the Subject Alternative Name does not match the DNS name the API server dials. For a service reference that name is <service>.<namespace>.svc, so a certificate issued for the pod IP, for localhost, or only for the fully qualified cluster.local name without the .svc form will fail hostname verification even though the chain is trusted.
- Does the webhook know that the caller really is the API server?Not by default — TLS here authenticates the server to the API server, not the other way round. If you need mutual authentication you configure the API server with an admission-control config file that supplies a client certificate for the webhook plugins, and have the webhook require and verify that client certificate. Otherwise you rely on NetworkPolicy and on the fact that a rogue caller can only obtain a decision, not mutate cluster state.
The caBundle is like naming, in advance, the single notary whose seal you will accept: if the notary changes seals without telling you, every document they stamp is refused — not because it is wrong, but because you no longer recognise the seal.
saying these in an interview costs you the question
- Assuming the API server trusts the cluster CA or the system trust store automatically
- Issuing a certificate for the pod IP or the Service's ClusterIP instead of <service>.<namespace>.svc
- Renewing the serving certificate but forgetting to update caBundle when the CA changed
- Thinking an expired certificate is a special case rather than an ordinary webhook failure governed by failurePolicy
- Using clientConfig.url pointing at an in-cluster address on a managed control plane that cannot reach the pod network