skip to content

Each webhook entry in a Kubernetes MutatingWebhookConfiguration or ValidatingWebhookConfiguration has a `failurePolicy` field set to either `Fail` or `Ignore`. Explain what each value does when the webhook backend is unreachable, times out or returns an error, and how you decide which one to use.

level: seniorimportance: must knowfreq 55%

answer

  1. Fail = reject on error, Ignore = admit unchecked
  2. v1 defaults to Fail (v1beta1 defaulted to Ignore)
  3. allowed:false is not a failure — always honored
  4. namespaceSelector excludes kube-system and self
  5. break-glass: delete the webhook configuration

basics

~20 s

Fail means a webhook error, timeout or unreachable backend causes the API request to be rejected; Ignore means the API server logs it and admits the request as if the webhook had approved it. Fail is safe for correctness, Ignore is safe for availability.

solid answer

~50 s

`failurePolicy` decides what the API server does when it cannot get an answer: connection refused, TLS failure, HTTP 5xx, a malformed AdmissionReview, or exceeding `timeoutSeconds`. - **`Fail`** (the default in `admissionregistration.k8s.io/v1`) rejects the API request with a 500-class error. The policy cannot be bypassed, but an unhealthy webhook now blocks writes to everything its rules match. - **`Ignore`** admits the request unchecked. Writes keep working; unvalidated or unmutated objects get through, and clients are not told. The decision is a security-versus-availability trade. Security controls that must not be bypassed — image provenance, privilege escalation checks — justify `Fail`, but only if the webhook is highly available, tightly scoped and excluded from itself. Convenience mutations such as label defaulting or sidecar injection are usually `Ignore`. In practice: ship new webhooks with `Ignore`, watch rejection and latency metrics, scope with `namespaceSelector` so `kube-system` is never affected, then flip to `Fail` once you trust it. Keep deleting the configuration as the documented break-glass step.

code

yaml · 26 lines
yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
  name: image-policy
webhooks:
  - name: images.example.com
    admissionReviewVersions: ["v1"]
    sideEffects: None
    failurePolicy: Fail
    timeoutSeconds: 3
    namespaceSelector:
      matchExpressions:
        - key: kubernetes.io/metadata.name
          operator: NotIn
          values: ["kube-system", "kube-node-lease", "policy-system"]
    rules:
      - operations: ["CREATE", "UPDATE"]
        apiGroups: [""]
        apiVersions: ["v1"]
        resources: ["pods"]
    clientConfig:
      service:
        name: image-policy
        namespace: policy-system
        path: /validate
      caBundle: LS0tLS1CRUdJTi...

go deeper

for a junior

State the two values and their effect: Fail rejects the request when the webhook cannot be reached, Ignore lets it through.

for a middle

Add what counts as a failure (timeouts, TLS, 5xx, malformed response), that v1 defaults to Fail, and that an explicit denial is unaffected by the setting.

for a senior

Lead with blast radius: which requests break when the webhook is down, namespaceSelector exclusions, HA and short timeouts, and the Ignore-then-Fail rollout with metrics.

for a principal

Frame it as a policy-integrity versus control-plane-availability decision: which controls must be unbypassable, who holds RBAC on webhook configurations (that is the break-glass and the bypass), and when to move rules into in-process CEL policies instead.

## What counts as a failure `failurePolicy` governs the case where the API server gets **no usable answer** from a webhook. That includes: the Service has no ready endpoints, the TCP connection is refused or reset, the TLS handshake fails (expired certificate, unknown CA, hostname mismatch), the call exceeds `timeoutSeconds` (default 10, maximum 30), the webhook returns a non-2xx status, or the response body is not a well-formed `AdmissionReview` with a matching `uid`. It does **not** cover a webhook that answers correctly with `allowed: false` — that is a deliberate rejection and is always honored regardless of `failurePolicy`. ## The two values **`Fail`** is the default in `admissionregistration.k8s.io/v1` (the older v1beta1 defaulted to `Ignore`, which caused a lot of silently unenforced policy). The API server converts the failure into an error response, so the client sees something like `Internal error occurred: failed calling webhook "policy.example.com": ... context deadline exceeded`. The guarantee you buy is **no bypass**: nothing gets into etcd without the webhook's opinion. **`Ignore`** makes the API server record the failure (metrics and audit) and continue admission as though the webhook had approved and made no changes. Writes keep flowing. The guarantee you lose is enforcement, and the loss is silent from the user's point of view — a Pod that should have been rejected is created, or one that should have been injected with a sidecar starts without it and behaves subtly wrong. ## Blast radius is what actually decides The honest framing is not "which is safer" but "what breaks when the webhook is down". A `Fail` webhook matching `rules: pods, CREATE` across all namespaces means: if the webhook is unavailable, **no Pod can be created anywhere**. That includes Deployments scaling up, DaemonSets landing on new nodes, evictions being rescheduled, and — critically — the webhook's own Pods if they get rescheduled. That last one is a genuine deadlock: the webhook cannot start because the webhook must approve it starting. The standard mitigations are: - **Exclude system and self namespaces.** Kubernetes labels every namespace with `kubernetes.io/metadata.name`, so a `namespaceSelector` with `NotIn [kube-system, <webhook-namespace>]` keeps the control plane and the webhook itself outside the trap. - **Scope the rules narrowly.** Match only the resources and operations you actually inspect. A webhook matching `resources: ["*"]` with `Fail` is an outage waiting for a rollout. - **Run it highly available.** Multiple replicas, a PodDisruptionBudget, topology spread across nodes/zones, readiness probes, and no dependency on the very workloads it gates. - **Keep timeouts short.** With `Fail`, `timeoutSeconds: 30` means a wedged webhook makes every matching API call hang for 30 seconds before failing — far worse for the cluster than failing in 2–5. ## Choosing per webhook, not per cluster `failurePolicy` is set on each webhook entry, so you can mix. A useful split: - **`Fail`** for controls whose bypass is a security incident: blocking `privileged: true`, enforcing signed images, denying host-path mounts, validating that a tenant cannot claim another tenant's namespace. - **`Ignore`** for ergonomics: defaulting labels and annotations, injecting a logging sidecar, filling in a default storage class. Missing these degrades convenience, not safety. Be aware of the compliance argument against `Ignore`: an attacker who can make your webhook unreachable (crash it, exhaust it, delete its Service) has disabled the control. If your control is genuinely security-critical, `Ignore` is not a real control — invest in webhook availability instead, or move the rule into an in-process `ValidatingAdmissionPolicy` where there is no network hop to break. ## Rollout and break-glass A safe lifecycle is: deploy with `Ignore` and a tight `namespaceSelector`; watch `apiserver_admission_webhook_rejection_count` and `apiserver_admission_webhook_admission_duration_seconds` to see what *would* have been rejected and how slow you are; widen the selector; then flip to `Fail`. The break-glass procedure must be written down before you need it: an admin with rights to delete the `ValidatingWebhookConfiguration` / `MutatingWebhookConfiguration` object can restore cluster writes in one command. Because that also removes the enforcement, the RBAC for editing webhook configurations is itself sensitive and should be tightly held and audited.

  • Your webhook matches Pod CREATE cluster-wide with failurePolicy: Fail, and its own Pods get evicted during a node drain. What happens?
    The webhook Pods cannot be admitted because admission requires the webhook that is now down, so the cluster deadlocks on Pod creation. The fix is to exclude the webhook's own namespace (and kube-system) with a namespaceSelector, spread replicas across nodes with a PodDisruptionBudget, and keep the documented break-glass step of deleting the webhook configuration.
  • A webhook explicitly returns allowed: false. Does failurePolicy: Ignore let the request through?
    No. failurePolicy only applies when the API server cannot obtain a valid answer — network errors, TLS problems, timeouts, 5xx or malformed responses. A well-formed denial is a successful call, so the request is rejected regardless of failurePolicy.
  • How would you find out what a webhook is doing before you switch it from Ignore to Fail?
    Run it in Ignore and observe the API server metrics: apiserver_admission_webhook_rejection_count broken down by webhook and error type shows what would have been denied and how often calls fail, while apiserver_admission_webhook_admission_duration_seconds shows the latency you are about to add to every matching write. Combine that with the webhook's own logs of its decisions before flipping.

saying these in an interview costs you the question

  • Believing Ignore also bypasses an explicit allowed:false denial
  • Saying the v1 default is Ignore (it is Fail)
  • Setting Fail on a cluster-wide pods rule without excluding kube-system and the webhook's own namespace
  • Treating an Ignore webhook as a real security control when an attacker can simply make it unreachable
  • Leaving timeoutSeconds at the maximum with Fail, turning a slow webhook into cluster-wide API latency

context