skip to content

OPA Gatekeeper and Kyverno are the two common policy engines for Kubernetes clusters. How does each one express a policy, and what practical differences would drive you to pick one over the other?

level: middleimportance: should knowfreq 38%

answer

  1. Gatekeeper = OPA + Rego; ConstraintTemplate → Constraint CRD
  2. Kyverno = plain YAML; validate / mutate / generate / verifyImages
  3. generation + cosign verification: Kyverno's differentiators
  4. referential policies need Gatekeeper's synced cache
  5. both: webhook in the path, audit mode, background scans

basics

~20 s

Both install as admission webhooks with CRDs. Gatekeeper wraps Open Policy Agent: you write rules in the Rego language inside a ConstraintTemplate, then instantiate Constraint objects. Kyverno uses plain Kubernetes YAML policies with validate, mutate, generate and verifyImages rules — no new language, and it can also create and clone resources.

solid answer

~60 s

**Gatekeeper** is Open Policy Agent packaged for Kubernetes. A `ConstraintTemplate` contains **Rego** (OPA's declarative query language) and defines a new Constraint CRD; you then create Constraint objects that scope the rule and supply parameters. It syncs selected cluster objects into OPA's cache so policies can be *referential* ("no duplicate Ingress hosts"). It runs an audit loop that flags existing violations, supports `enforcementAction: dryrun|warn|deny`, and mutation is a separate, more limited feature. **Kyverno** is Kubernetes-native by design: a `ClusterPolicy` is YAML with `match`/`exclude` blocks and rules of type `validate`, `mutate`, `generate`, or `verifyImages`. No new language to learn, and it does things Gatekeeper does not do as naturally — generating a NetworkPolicy into every new namespace, cloning secrets, verifying cosign signatures out of the box. It also has audit/enforce modes, background scanning, and PolicyReport output. Practically: Gatekeeper if you already invest in Rego, want that expressive power, or need OPA elsewhere; Kyverno if you want low adoption cost, generation, or built-in image verification. Both increasingly complement rather than replace ValidatingAdmissionPolicy.

code

yaml · 55 lines
yaml
# Gatekeeper: template (Rego) + constraint
apiVersion: templates.gatekeeper.sh/v1
kind: ConstraintTemplate
metadata:
  name: k8srequiredlabels
spec:
  crd:
    spec:
      names: {kind: K8sRequiredLabels}
      validation:
        openAPIV3Schema:
          properties:
            labels: {type: array, items: {type: string}}
  targets:
    - target: admission.k8s.gatekeeper.sh
      rego: |
        package k8srequiredlabels
        violation[{"msg": msg}] {
          required := input.parameters.labels[_]
          not input.review.object.metadata.labels[required]
          msg := sprintf("missing required label: %v", [required])
        }
---
apiVersion: constraints.gatekeeper.sh/v1beta1
kind: K8sRequiredLabels
metadata:
  name: ns-must-have-team
spec:
  enforcementAction: deny
  match:
    kinds:
      - apiGroups: [""]
        kinds: ["Namespace"]
  parameters:
    labels: ["team"]
---
# Kyverno: one policy, no new language
apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: ns-must-have-team
spec:
  validationFailureAction: Enforce
  rules:
    - name: check-team-label
      match:
        any:
          - resources:
              kinds: ["Namespace"]
      validate:
        message: "missing required label: team"
        pattern:
          metadata:
            labels:
              team: "?*"

go deeper

for a junior

Know that both are add-on policy engines installed as admission webhooks, that Gatekeeper policies are written in Rego and Kyverno policies in ordinary YAML.

for a middle

Explain Gatekeeper's ConstraintTemplate/Constraint two-level model versus Kyverno's rule types, and name generation and image verification as Kyverno-specific capabilities.

for a senior

Argue selection from team skills and operational load — Rego familiarity, cache memory, request-path latency, audit/reporting workflow — and describe the audit-first rollout.

for a principal

Position it as governance architecture: which invariants live in-tree as CEL, which need an engine, how policies are versioned and reviewed as code, exception handling, and the cost of a second control-plane dependency.

## What both are Neither engine is part of Kubernetes. Both install into the cluster as a Deployment plus a `ValidatingWebhookConfiguration` (and usually a `MutatingWebhookConfiguration`), so every matching API write is routed to them for a decision. Both add CRDs so that policies themselves are cluster objects — versioned in Git, reviewed, and applied like anything else. Both report violations of *existing* objects through a background scan, because admission control only ever sees new writes. ## Gatekeeper's model Gatekeeper is the Kubernetes front end for **Open Policy Agent (OPA)**, a general-purpose policy engine used well beyond Kubernetes. Policies are written in **Rego**, a declarative query language. The structure is two-level: 1. A **ConstraintTemplate** contains the Rego and declares a schema for parameters. Creating it registers a *new CRD* — e.g. `K8sRequiredLabels`. 2. A **Constraint** is an instance of that CRD: it says which kinds/namespaces the rule applies to, supplies parameters (`labels: [team]`), and sets `enforcementAction`. The payoff is reuse: one template, many constraints, each with different scope and parameters. The cost is Rego, which is genuinely a new language with its own evaluation model — powerful, but a real learning curve and a common source of subtly wrong policies. Gatekeeper can also **sync** chosen resource kinds into OPA's in-memory cache (`Config` object), which is what enables referential policies that compare the incoming object against others in the cluster. That cache is eventually consistent, so referential rules are best-effort against races. Mutation exists (`Assign`, `AssignMetadata`, `ModifySet` CRDs) but is deliberately narrower than validation. ## Kyverno's model Kyverno takes the opposite bet: **no new language**. A `Policy` (namespaced) or `ClusterPolicy` contains rules; each rule has a `match`/`exclude` block and exactly one of: - `validate` — a `pattern` (YAML overlay with wildcards and anchors), a `deny` block with expressions, or a CEL expression; `validationFailureAction: Enforce|Audit`. - `mutate` — a strategic-merge patch or JSON patch applied to the object, optionally to existing resources too. - `generate` — create or clone a resource when a trigger appears. This is Kyverno's signature capability: every new namespace automatically gets a default-deny NetworkPolicy, a ResourceQuota, and a pull secret, kept in sync. - `verifyImages` — verify cosign/Notary signatures and attestations against a key or keyless identity, and (importantly) mutate the image reference to its resolved digest so the verified artifact is the one that runs. Because policies are YAML overlays, simple rules are very short and readable; complex logic can get awkward, which is why CEL support was added. ## Comparing them where it matters **Learning cost.** Kyverno wins for a team with no Rego background; a working "require these labels" policy is a dozen lines of familiar YAML. Gatekeeper wins if OPA/Rego is already used for API authorization or CI policy, since the language and tooling are shared. **Expressiveness.** Rego handles complex, set-oriented, cross-object logic more comfortably. Kyverno covers the common 90% and now offers CEL for the rest. **Beyond validation.** Only Kyverno makes resource *generation* and cloning first-class, and only Kyverno ships image signature verification as a built-in rule type; with Gatekeeper you add an external provider (e.g. an image-verification service it calls out to). **Operational shape.** Both run in the request path with the same `failurePolicy` trade-off and both need HA and namespace exclusions. Gatekeeper's cache-syncing adds memory proportional to the synced resources. Kyverno's controllers doing generation add write load and reconciliation of their own. **Reporting.** Both emit violation reports for existing objects (Kyverno via the PolicyReport CRDs, Gatekeeper via constraint status and audit); this matters because you inherit clusters full of non-conforming workloads and need to measure before enforcing. ## Where ValidatingAdmissionPolicy fits With CEL-based policies built into the API server, the cheapest field checks no longer need either engine. The mature stance is a split: in-tree CEL for simple, high-volume invariants; a policy engine for mutation, generation, image provenance, and referential rules. Both projects support emitting or authoring in CEL to ease that split.

  • You adopt a policy engine on a cluster that already runs hundreds of non-conforming workloads. How do you avoid breaking them?
    Start in audit mode — Gatekeeper's enforcementAction: dryrun (or warn), Kyverno's validationFailureAction: Audit — so existing and new objects are only reported. Use the background scan / PolicyReport output to quantify violations per team, drive remediation with owners, and only flip to deny/Enforce once the report is clean or the remaining cases have explicit, time-boxed exclusions.
  • Why can a Gatekeeper referential policy such as 'no two Ingresses may claim the same hostname' produce a wrong decision?
    Referential policies read from OPA's synced cache of cluster state, which is populated asynchronously by watches and is therefore eventually consistent. Two conflicting Ingresses created at nearly the same moment can each be evaluated against a cache that does not yet contain the other, so both are admitted. Admission control is not a distributed lock; the audit loop is what catches the resulting violation afterwards.

saying these in an interview costs you the question

  • Saying Gatekeeper or Kyverno is part of Kubernetes — both are add-ons installed as webhooks.
  • Claiming a policy engine secures existing workloads; admission only sees new writes, so you also need background scanning.
  • Thinking Rego is just YAML with extra steps, and underestimating the learning curve for correct policies.
  • Believing referential policies are transactional/consistent — they read an eventually-consistent cache.
  • Assuming an engine must be chosen instead of ValidatingAdmissionPolicy rather than alongside it.

context