A Kubernetes CRD's conversion webhook becomes slow or unreachable; what breaks across the cluster, and how would you run that webhook so it cannot stall the API?
answer
- inside the request path
- nothing to fail open
- informers, namespaces, migrations
- replicas, PDB, every replica serves
- caBundle expiry takes it down
basics
~20 sEvery request needing conversion fails or slows: lists, watches, controllers' informers, writes, namespace deletion and migrations. There is no failurePolicy to skip it, so run the webhook redundantly with a PodDisruptionBudget, pure fast code and rotated TLS.
solid answer
~50 sWith `strategy: Webhook` the conversion call sits inside the API server's request path, and `spec.conversion.webhook` has no `failurePolicy` or `timeoutSeconds` field, so it cannot be skipped or given its own timeout. When it is down, any get, list, watch or write that needs conversion fails with `conversion webhook for ... failed`: operators' informers cannot sync, a Namespace holding such objects can stick in `Terminating`, and migrations fail. When it is slow, that shows up as API latency, like a 310 ms p99 on `SuggestIndex` lists. Run it like control-plane infrastructure: two or more replicas spread across nodes with a PodDisruptionBudget, served from every replica rather than only the leader, pure field mapping with no outbound calls, `caBundle` and serving certificates rotated with overlap, and alerts on `apiserver_crd_conversion_webhook_duration_seconds`. Then shrink exposure by migrating stored objects and reading at the storage version.
code
bash · 13 lines# Does a list that needs conversion fail or crawl?
time kubectl get --raw '/apis/autocomplete.example.com/v1/suggestindexes?limit=1'
# Where does the CRD send conversion requests?
kubectl get customresourcedefinition suggestindexes.autocomplete.example.com \
-o jsonpath='{.spec.conversion.webhook.clientConfig.service}'
# Does the webhook Service have ready endpoints?
kubectl -n search-system get endpointslices \
-l kubernetes.io/service-name=suggestindex-conversion
# What latency does the API server see for the webhook?
kubectl get --raw /metrics | grep apiserver_crd_conversion_webhook_duration_secondsgo deeper
Remember that a CRD using webhook conversion depends on that service for every request that needs conversion, and that there is no setting to bypass it.
Explain which requests trigger conversion and why lists and watches fail when even some stored objects are in an older version.
Diagnose from symptoms to Service endpoints to webhook latency, then harden with replicas, a PodDisruptionBudget, pure conversion code and managed TLS.
Weigh whether a schema change is worth a new availability dependency in the API path, and set a platform rule for migrating and removing conversion webhooks.
## Where conversion sits A CustomResourceDefinition (CRD) with `spec.conversion.strategy: Webhook` makes an external HTTPS service part of the API server's **read and write path** for that resource. Whenever a request involves an object whose stored version differs from the version needed — a `v1` read of an object stored as `v1alpha1`, or a `v1alpha1` write that must be stored as `v1` — the API server sends a `ConversionReview` and waits for the answer before it can respond. Two facts make this different from an admission webhook: - **There is no `failurePolicy`.** `spec.conversion.webhook` holds only `clientConfig` and `conversionReviewVersions`. The call cannot be skipped, because the webhook produces the data itself; without it the API server has no correct object to return. - **There is no `timeoutSeconds` field.** You cannot give the call its own timeout, so a slow webhook turns directly into API latency. ## What an outage breaks | Symptom | Why | |---|---| | `kubectl get` and lists fail with "conversion webhook for ... failed" | Each list converts every object not already in the requested version | | Operators and other controllers stop reconciling | Their informers cannot list or watch the resource | | Writes in a non-storage version fail | The body must be converted before it is persisted | | A Namespace stays `Terminating` with `NamespaceDeletionContentFailure` | The namespace controller cannot finish removing the custom resources | | A storage version migration fails | Every re-encoding write goes through conversion | | API latency rises for that resource | Each conversion adds a network round trip | A partial failure is the same story at lower volume: requests that happen to need conversion fail, the rest succeed, which makes the symptom look random. ## A worked incident The search-autocomplete team runs its operator on a 12-node GPU cluster for model serving. After a node drain, lists of `suggestindexes.autocomplete.example.com` show a **310 ms p99** latency spike, and some fail outright. 1. Time a raw list of the resource at the served version with `kubectl get --raw`, to confirm the slowness is in the API path and not in the client. 2. Read `spec.conversion.webhook.clientConfig` from the CRD to find the Service. 3. Check the Service's EndpointSlices: here only one of two webhook Pods was ready, and it sat on a node with heavy load. 4. Read `apiserver_crd_conversion_webhook_duration_seconds` from the API server's metrics to confirm the time is spent in the webhook call. 5. Restore capacity, then fix the causes below so a single drain cannot repeat it. ## Hardening the webhook - **Run at least two replicas** behind the Service, spread with pod anti-affinity or topology spread constraints, protected by a **PodDisruptionBudget** so a drain never removes all of them. - **Serve conversion from every replica**, not only the leader-elected one; leader election is for reconcile loops, not for request serving. - **Keep conversion pure and fast**: no calls to other services, no reads from the API server, just field mapping. Give the Pods requests high enough that they are not starved. - **Automate TLS**: the API server trusts the serving certificate through `caBundle`. When it expires, every conversion fails. Rotate with overlap — put old and new CA in `caBundle` before switching the serving certificate — or let a certificate controller inject it. - **Give the Pods a suitable priority** so they are not preempted before workloads that depend on the API. - **Alert** on conversion latency and on conversion failures, not only on Pod readiness. ## Reducing how often it is called - Have your own controllers and tooling request the **storage version**, so already-migrated objects need no conversion. - **Migrate stored objects** to the current storage version; after that, requests at the storage version skip the webhook entirely. - Once only one version remains in `spec.versions`, switch back to `strategy: None` and remove the webhook. ## Pitfalls 1. **Uninstalling the operator first.** Deleting the webhook Deployment while the CRD still says `Webhook` breaks every request that needs conversion — including the deletes needed to clean up. Remove the custom resources and the CRD before the webhook, or switch the strategy first. 2. **Webhook in a Namespace being deleted.** If the webhook's Namespace is deleted along with custom resources that need conversion, deletion can deadlock. 3. **A certificate lifecycle nobody watches.** Conversion and admission webhooks use the same `caBundle` plumbing; if an operator serves both from one certificate, its expiry takes both down at once.
- Why does a Kubernetes CRD conversion webhook have no failurePolicy when admission webhooks do?An admission webhook passes judgement on a request that is otherwise complete, so skipping it is a policy trade-off the API server can make. A conversion webhook produces the object in the requested version; if it is skipped, the API server has nothing correct to return or store. Falling back to the unconverted object would silently hand clients wrong data, so the request fails instead.
- How does uninstalling an operator break a Kubernetes CRD that still uses Webhook conversion?The CRD keeps pointing at a Service with no backends, so every request that needs conversion fails, including the deletes a cleanup needs. Custom resources and Namespaces can then get stuck. Delete the custom resources and the CRD before removing the webhook, or first migrate storage and switch the CRD to `strategy: None`.
- How do you rotate a Kubernetes CRD conversion webhook's serving certificate without an outage?Add the new CA to the CRD's `spec.conversion.webhook.clientConfig.caBundle` alongside the old one, wait until the API servers pick it up, then switch the webhook Pods to the new certificate, and remove the old CA later. A certificate controller can inject `caBundle` automatically, but alert on expiry anyway.
saying these in an interview costs you the question
- Set failurePolicy Ignore on the conversion webhook to ride out outages.
- A conversion outage hurts only clients still asking for v1alpha1.
- The API server calls the conversion webhook once per CRD update, not per request.
- Serving the webhook only from the leader-elected operator replica is enough.
- An expired webhook serving certificate only produces API server warnings.