You are about to roll out a cluster-wide admission webhook that intercepts every Pod creation in a production Kubernetes cluster. What latency and availability risks does that introduce, and how would you scope, size and roll it out so that a problem with the webhook cannot take the cluster down?
answer
- webhook = your service on the control-plane critical path
- timeout 2–5s, not the 10s default or 30s max
- namespaceSelector: NotIn kube-system + own namespace
- HA replicas, PDB, spread, no dependency on what it gates
- Ignore + audit + metrics → then Fail; break-glass = delete the config
basics
~20 sEvery matching write now waits on a network call, so webhook latency becomes API latency and webhook downtime can block Pod creation. Scope with rules and selectors, exclude kube-system and the webhook's own namespace, keep timeouts at a few seconds, run it HA, and ship with failurePolicy Ignore before flipping to Fail.
solid answer
~50 sRegistering a webhook puts your service on the **critical path of the control plane**. Three risks follow. **Latency**: every matching write blocks on the call. Validating webhooks run in parallel, mutating ones serially, so mutating chains add up. With `timeoutSeconds` at the 30s maximum, a wedged webhook stalls every Pod create for 30 seconds — controllers back off, rollouts crawl. **Availability**: with `failurePolicy: Fail`, webhook downtime means no Pod creation. If the webhook's own Pods match its rules, it cannot restart itself — a hard deadlock. **Correctness under load**: mass events (node failure, cluster autoscaler scale-up) generate admission bursts far above steady state. Mitigations: narrow `rules`, `namespaceSelector` excluding `kube-system` and its own namespace, `objectSelector` for opt-in labels, `timeoutSeconds` of 2–5, multiple replicas with a PDB and topology spread, no dependency on gated workloads, and in-memory decisions with no synchronous external calls. Roll out with `Ignore`, watch the admission metrics, then flip to `Fail`. Document deleting the webhook configuration as break-glass.
code
yaml · 30 linesapiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
name: pod-policy
webhooks:
- name: pods.policy.example.com
admissionReviewVersions: ["v1"]
sideEffects: None
failurePolicy: Ignore # flip to Fail only after the soak
timeoutSeconds: 3
matchPolicy: Equivalent
namespaceSelector:
matchExpressions:
- key: kubernetes.io/metadata.name
operator: NotIn
values: ["kube-system", "kube-node-lease", "policy-system"]
objectSelector:
matchLabels:
policy.example.com/enforce: "true"
rules:
- operations: ["CREATE"]
apiGroups: [""]
apiVersions: ["v1"]
resources: ["pods"]
clientConfig:
service:
name: pod-policy
namespace: policy-system
path: /validate
caBundle: LS0tLS1CRUdJTi...go deeper
Recognise that a webhook is called on every matching API write, so if it is slow or down, those writes are slow or fail.
Name the concrete knobs: timeoutSeconds, failurePolicy, namespaceSelector and objectSelector, plus running more than one replica.
Give an operational plan — scoping, HA topology, the Ignore-then-audit-then-Fail rollout, the specific API server metrics you alert on, and the break-glass command.
Frame it as extending the control plane's failure domain: decide which rules deserve unbypassable enforcement, what the organisation's admission latency budget is, who owns the webhook's on-call, and when a CEL policy in-process is the better engineering trade than a service you must keep alive.
## What changes the moment you register the webhook A `MutatingWebhookConfiguration` or `ValidatingWebhookConfiguration` inserts your HTTP service into the API server's synchronous write path. From that moment your service shares the availability and latency requirements of the control plane itself, for every request its `rules` match. Interviewers ask this because it is one of the few extension points where an application-level mistake produces a cluster-level outage. ## The latency budget Every matching write pays: DNS/TCP/TLS to the Service, request serialization of the full object, your processing, and the response. Validating webhooks run concurrently, so their cost is roughly the slowest one; mutating webhooks run serially and their costs add, and `reinvocationPolicy: IfNeeded` can double a call. `timeoutSeconds` defaults to 10 with a maximum of 30. That default is dangerous on a hot path: a webhook stuck on a slow dependency turns every Pod creation into a 10-second wait, and controllers that create Pods (Deployments, Jobs, the cluster autoscaler's churn) start to queue. Pick **2–5 seconds** and make your handler meet it comfortably. The deeper rule is: **make admission decisions from memory**. A webhook that synchronously queries an external policy service, a database, or even the Kubernetes API on every call has imported that dependency's tail latency into cluster writes. Use an informer-backed cache for cluster state, preload policy, and refresh asynchronously. Burst behavior matters as much as steady state. A node failure or an autoscaler scale-up can produce hundreds of Pod creations in seconds; a webhook sized for average traffic will queue, hit its timeout, and — with `Fail` — convert a recoverable node failure into a cluster that cannot reschedule. ## The availability risk and the self-deadlock With `failurePolicy: Fail`, unavailability of the webhook equals unavailability of the operation it gates. The pathological case is circular: the webhook matches `pods/CREATE` in all namespaces, its own Pods are evicted during a node drain or upgrade, and they cannot be admitted because the admitter is down. Defensive scoping: - **`namespaceSelector`** excluding `kube-system`, `kube-node-lease` and the webhook's own namespace, using the automatic `kubernetes.io/metadata.name` label. This is the single most important line in the configuration. - **Narrow `rules`** — only the resources, API groups and operations you actually inspect. `resources: ["*"]` with `Fail` is an outage waiting for the next rollout. - **`objectSelector`** so workloads opt in by label, which also gives you a per-workload escape hatch. - **`matchPolicy: Equivalent`** (the default) if you want the rule to still match when a client uses a different but equivalent API version — important so that policy is not bypassed by version choice, at the cost of matching more traffic. And defensive topology: at least three replicas, a `PodDisruptionBudget`, `topologySpreadConstraints` across nodes and zones, `priorityClassName: system-cluster-critical` so it is not the first thing evicted, and readiness probes that fail *before* the process degrades. The webhook must not depend on anything it gates — not on a mesh sidecar it injects, not on a database in the same cluster gated by itself. ## Rollout sequence A disciplined rollout looks like: 1. Deploy the backend with no webhook configuration and load-test it at expected burst rate. 2. Register with **`failurePolicy: Ignore`**, a tight `namespaceSelector` (one non-production namespace), short timeout. 3. Run in "audit" mode — the webhook logs the decision it would make and always allows — so you can measure how much would break. 4. Watch the API server's own metrics: `apiserver_admission_webhook_admission_duration_seconds` (per webhook, p99), `apiserver_admission_webhook_rejection_count`, and overall `apiserver_request_duration_seconds` for the gated verbs. Add SLO alerts before enforcement. 5. Widen the selector namespace by namespace; start enforcing (real denials) while still on `Ignore`. 6. Only then flip to `Fail`, and only for the rules that genuinely must not be bypassed. ## Break-glass and governance Write the recovery procedure down before you need it: `kubectl delete validatingwebhookconfiguration <name>` restores writes immediately. Ensure at least one credential that can do this does not itself depend on the cluster's gated workloads (a static admin kubeconfig, not an SSO flow that runs in-cluster). Because that same permission also disables enforcement, `update`/`delete` on webhook configurations is a privileged, audited right. Finally, consider whether you need a webhook at all. Rules expressible in CEL can run as an in-process `ValidatingAdmissionPolicy` — no network hop, no TLS, no availability story, and failure modes confined to the API server. Reserve webhooks for logic that genuinely needs to run your code, and accept the operational burden knowingly.
- How do you keep a Fail-policy webhook from blocking its own Pods from starting?Exclude its namespace with a namespaceSelector on kubernetes.io/metadata.name, so its own Pod creations never route through it, and exclude kube-system for the same reason. Back that with multiple replicas spread across nodes, a PodDisruptionBudget, and a system-cluster-critical priority class so a drain or upgrade never removes all of them at once.
- What would push you to use an in-process ValidatingAdmissionPolicy instead of a webhook?If the rule is expressible in CEL over the object and a few cluster resources, the in-process policy removes the network hop, the TLS certificate lifecycle, and the entire availability story — there is no backend to be down. Webhooks earn their operational cost only when the decision needs your own code: calls into custom libraries, complex state, or logic that cannot be expressed declaratively.
- Your webhook's p99 latency is 40 ms but Pod creation p99 jumped by 900 ms. What would you look at?Look at whether the latency appears only during bursts — node failure or autoscaler events create admission storms that queue behind limited replicas or a saturated CPU limit. Check the API server's per-webhook duration histogram rather than the webhook's own timer, since it includes connection setup and TLS, and check whether a mutating chain plus reinvocation is calling the handler more than once per request.
It is like adding a mandatory security desk at the only entrance of a building: fine when it is staffed and fast, but if the guard is slow everyone queues, and if the guard is locked outside the building nobody — including the replacement guard — can get in.
saying these in an interview costs you the question
- Leaving timeoutSeconds at the default or maximum on a hot path
- Matching all resources and all namespaces because it is simpler to configure
- Running a single replica of the webhook backend and calling it highly available
- Making synchronous calls to an external policy service or database inside the admission handler
- Going straight to failurePolicy: Fail on day one with no audit period or metrics
- Having no written break-glass procedure, or one that requires the very workloads the webhook gates