skip to content

Policy Enforcement via Admission

Admission is where policy bites: in-tree ValidatingAdmissionPolicy evaluates CEL inside the API server, while an engine sits behind a webhook whose failurePolicy trades safety for availability. Blocking privileged pods cluster-wide is the stock question.

part ofKubernetesoverview, primer and where to startread it →
on this pageshow

questions

5

When you register an admission webhook with Kubernetes, the configuration has a failurePolicy field set to either Fail or Ignore. What does each value do when the webhook is unreachable, and how do you decide which one a security policy webhook should use?

level: seniorimportance: must knowfreq 50%

answer

  1. Fail = fail-closed, outage risk
  2. Ignore = fail-open, silent bypass
  3. only applies to errors/timeouts, not explicit denials
  4. exclude kube-system + own namespace, else bootstrap deadlock
  5. HA replicas + low timeoutSeconds + matchConditions

basics

~20 s

failurePolicy: Fail rejects any matching API write when the webhook errors or times out — policy holds, but an outage of your webhook blocks writes cluster-wide. failurePolicy: Ignore admits the request unchecked — the cluster keeps working, but policy silently stops being enforced. Security policies want Fail plus a highly available webhook and careful scoping.

solid answer

~60 s

`failurePolicy` decides what happens when the API server cannot get an answer: connection refused, TLS failure, or exceeding `timeoutSeconds` (default 10, max 30). - **Fail** — the request is rejected. Your policy is fail-closed and cannot be bypassed by killing the webhook, which is the whole point of a security control. The risk is a self-inflicted cluster outage: if the webhook is down, every matching write fails, and if it matches broadly you can deadlock — the webhook's own pods can't be scheduled because admitting them requires the webhook. - **Ignore** — the request proceeds unvalidated. The cluster stays available, but enforcement is silently off and an attacker who can make the webhook unreachable has bypassed it. For a security webhook the answer is `Fail`, made survivable by engineering rather than by weakening the policy: run multiple replicas with a PodDisruptionBudget and anti-affinity, keep the handler fast with a low `timeoutSeconds`, narrow the blast radius with `rules`, `objectSelector`, `namespaceSelector` (exclude `kube-system` and the webhook's own namespace), and `matchConditions`. Non-security, best-effort webhooks can reasonably use `Ignore`.

code

yaml · 31 lines
yaml
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingWebhookConfiguration
metadata:
  name: pod-policy
webhooks:
  - name: pods.policy.example.com
    admissionReviewVersions: ["v1"]
    sideEffects: None
    failurePolicy: Fail
    timeoutSeconds: 3
    rules:
      - apiGroups:   [""]
        apiVersions: ["v1"]
        operations:  ["CREATE"]
        resources:   ["pods"]
        scope: Namespaced
    namespaceSelector:
      matchExpressions:
        - key: kubernetes.io/metadata.name
          operator: NotIn
          values: ["kube-system", "policy-system"]
    matchConditions:
      - name: skip-node-and-system-agents
        expression: "!(request.userInfo.username.startsWith('system:node:'))"
    clientConfig:
      service:
        namespace: policy-system
        name: policy-webhook
        path: /validate
        port: 443
      caBundle: <base64-ca>

go deeper

for a junior

Know the two values and their literal effect: Fail rejects the request when the webhook can't be reached, Ignore lets it through unchecked.

for a middle

Add that it only covers errors and timeouts, and that Ignore means enforcement silently stops — so security rules lean towards Fail.

for a senior

Lead with the trade-off and then the engineering that makes Fail survivable: HA replicas, tight rules/selectors/matchConditions, short timeouts, namespace exclusions, cert rotation, and detection of enforcement gaps.

for a principal

Treat it as an availability-versus-integrity policy decision made per rule class, with an explicit blast-radius budget for the control plane, a break-glass procedure, and monitoring that proves enforcement is actually on.

## What the setting actually controls A `ValidatingWebhookConfiguration` or `MutatingWebhookConfiguration` tells kube-apiserver: for these resources and verbs, make a synchronous HTTPS call to this endpoint and use the reply to admit, reject, or patch the object. `failurePolicy` covers only the case where **no usable reply arrives** — connection refused, DNS failure, TLS handshake error, HTTP 5xx, or the call exceeding `timeoutSeconds`. It does not apply when the webhook answers "denied"; an explicit denial always rejects regardless of failurePolicy. ## Fail: fail-closed With `failurePolicy: Fail` an unreachable webhook turns into an API error on every matching write. Security-wise this is what you want: a control you can disable by crashing a pod is not a control. Availability-wise it is the single most common way people take down their own cluster. The classic failure modes: 1. **Bootstrap deadlock.** The webhook matches `pods` in all namespaces, including its own. Restart the cluster or drain the node and the webhook pods cannot be admitted because admitting them requires the webhook that is not running. The fix is to exclude the webhook's own namespace (and usually `kube-system`) via `namespaceSelector`. 2. **Cascading control-plane failure.** A webhook matching `*/*` also intercepts writes by controllers, node registration, lease renewals, and CRD writes. A slow handler adds its latency to every one. 3. **Certificate expiry.** The `caBundle` stops matching the serving cert; every call fails TLS and the cluster becomes read-only for matching resources. Cert rotation for webhooks is a real operational duty. ## Ignore: fail-open `failurePolicy: Ignore` keeps the cluster writable when the webhook is down, at the cost of an invisible enforcement gap. Two things make this worse than it looks. First, it is silent — the API server records a failure metric and (with `matchPolicy`) may emit warnings, but the user's `kubectl apply` succeeds normally, so nobody notices. Second, it is attacker-controllable: anyone who can pressure the webhook (delete its pods, exhaust it with load, break its network path) has disabled the policy for the duration. For an image-provenance or privileged-container policy that is a complete bypass. ## How to actually decide The honest framing is not "which value is safer" but "what is the cost of a false reject versus an unchecked admit for this specific rule". - **Security-critical, narrowly scoped** (block privileged pods, require signed images): `Fail`. Earn the availability back with replicas across nodes/zones, a PodDisruptionBudget, readiness probes, low `timeoutSeconds` (1–3s for a cheap check), and tight `rules`/`matchConditions` so the webhook is on the path of as few requests as possible. - **Advisory or convenience mutation** (inject a sidecar, add a default label): `Ignore` is defensible; a missing sidecar is usually better than a broken deployment pipeline. Note the subtle trap: a mutating webhook that injects a security sidecar with `Ignore` yields workloads silently running without it. - **Never** use `Fail` with a match rule of all resources and all namespaces and no exclusions. ## Reducing the blast radius The fields that matter alongside `failurePolicy`: - `rules` — the narrowest apiGroups/resources/operations that satisfy the policy. - `namespaceSelector` — exclude `kube-system` and the policy engine's own namespace; many clusters label those with `control-plane` or `admission.example.com/exempt`. - `objectSelector` — only objects carrying (or lacking) a label. - `matchConditions` — CEL predicates evaluated by the API server *before* the call, so exempted requests never touch the network. Excluding system service accounts here is cheap and effective. - `timeoutSeconds` — with `Fail`, this is your added tail latency on every write and your outage trigger; keep it small. - `sideEffects: None` and honouring `dryRun` — required so `kubectl --dry-run=server` and quota evaluation behave. ## Detecting the gap Watch `apiserver_admission_webhook_rejection_count` and the webhook request-duration/failure metrics, alert on non-zero failure rates even when `failurePolicy: Ignore` keeps writes flowing, and run a periodic "canary" object that the policy should reject — if it gets admitted, enforcement is off. Policy engines' audit modes complement this by re-scanning existing objects, catching anything admitted during a gap.

  • Your cluster suddenly rejects every pod creation with a message about calling a webhook. What do you check, and how do you restore service?
    Confirm it is admission by reading the error — it names the webhook. Check the webhook Deployment's pods, its Service endpoints, and whether the serving certificate or caBundle expired. To restore service immediately you can edit the ValidatingWebhookConfiguration to failurePolicy: Ignore, or delete the configuration object (it is recreated by its operator), then fix the backend and restore Fail. Keep the window short and re-audit anything admitted during it.
  • Why is matchConditions preferable to doing the same exemption check inside your webhook handler?
    matchConditions are CEL predicates evaluated by the API server before the network call, so exempted requests never leave the API server: no latency, no dependency on the webhook being up, and no risk that a Fail policy blocks them during an outage. Filtering inside the handler still requires the handler to be reachable, which is exactly what you were trying to avoid.

Fail is a turnstile that locks when the power dies; Ignore is one that swings free. For a vault door you choose locked and then buy a generator — you don't leave it swinging.

saying these in an interview costs you the question

  • Saying failurePolicy also governs what happens when the webhook explicitly denies — it does not; denials always reject.
  • Reflexively choosing Ignore "to be safe", not recognising that it is a silent, attacker-triggerable bypass of the control.
  • Setting Fail with a cluster-wide rule and no namespace exclusions, then being surprised by a bootstrap deadlock.
  • Ignoring timeoutSeconds — a 30s timeout with Fail turns a slow webhook into a cluster-wide write stall.
  • Believing a down webhook with Ignore is visible to users; the request succeeds normally and nothing is logged for them.

context

open as a page

OPA Gatekeeper and Kyverno are the two common policy engines for Kubernetes clusters. How does each one express a policy, and what practical differences would drive you to pick one over the other?

level: middleimportance: should knowfreq 38%

basics

~20 s

Both install as admission webhooks with CRDs. Gatekeeper wraps Open Policy Agent: you write rules in the Rego language inside a ConstraintTemplate, then instantiate Constraint objects. Kyverno uses plain Kubernetes YAML policies with validate, mutate, generate and verifyImages rules — no new language, and it can also create and clone resources.

open as a page

Kubernetes ships a built-in ValidatingAdmissionPolicy resource whose rules are written in CEL (Common Expression Language) and evaluated inside the API server. How does it work, and when would you pick it over an external validating admission webhook?

level: middleimportance: should knowfreq 45%

basics

~20 s

ValidatingAdmissionPolicy holds CEL expressions the API server evaluates in-process on incoming objects. A ValidatingAdmissionPolicyBinding scopes it to namespaces or resources and picks the action: Deny, Warn, or Audit. No external service means no serving certificates, no network hop, no availability risk. But CEL cannot mutate or call out.

open as a page

How would you enforce, at Kubernetes admission time, that workloads may only run container images from approved registries, pinned by digest, and carrying a valid signature?

level: seniorimportance: should knowfreq 34%

basics

~20 s

Use an admission policy engine. A simple CEL or YAML rule can require the image reference to start with an approved registry prefix and to contain an @sha256: digest. Signature and provenance checks need an engine that can reach the registry — for example Kyverno's verifyImages with cosign — which verifies the signature and rewrites the tag to the verified digest.

open as a page

You must introduce a cluster-wide rule that rejects Pods running as UID 0 across dozens of teams' namespaces in a running cluster. How do you roll that out without breaking production workloads, and how do you handle the teams that legitimately cannot comply?

level: principalimportance: should knowfreq 28%

basics

~20 s

Never start at deny. Ship the rule in audit/warn mode first, measure violations per namespace from the reports, give owners a deadline and help, then enforce namespace by namespace starting with low-risk ones. Handle exceptions as explicit, labelled, time-boxed, individually-reviewed exclusions — not by weakening the global rule.

open as a page