How does a Kyverno cleanup policy remove completed Jobs, and when do you use a TTL label instead?
answer
- A rule whose outcome is deletion
- Cron schedule plus a match block
- Conditions narrow it to finished work
- A label can carry a per-object expiry
- Deleting needs granted permissions too
basics
~20 sA Kyverno ClusterCleanupPolicy matches a kind, narrows it with conditions such as a succeeded status, and carries a cron schedule; the cleanup controller deletes every match on each tick. A cleanup.kyverno.io/ttl label instead expires one specific object.
solid answer
~50 sA `CleanupPolicy`, or its cluster-scoped form `ClusterCleanupPolicy`, is a rule whose outcome is deletion rather than admission. It has a `match` block selecting the kind, optional `exclude` and `conditions` expressions to narrow it - for a completed Job, a condition on the succeeded count - and a `schedule` in cron format. On each tick the cleanup controller lists the matches and deletes them, so it is a fleet-wide sweep expressed once. The `cleanup.kyverno.io/ttl` label is the per-object alternative: put it on a resource with a duration or an absolute timestamp and that one object is deleted when it expires, which suits sandbox namespaces where every tenant gets a different lease. The controller deletes under its own ServiceAccount, so it needs RBAC for every kind you point it at - the usual reason a schedule appears to run and nothing goes away.
code
yaml · 16 lineskind: ClusterCleanupPolicy
metadata:
name: cleanup-completed-jobs
spec:
match:
any:
- resources:
kinds:
- Job
conditions:
all:
- key: "{{ target.status.succeeded }}"
operator: Equals
value: 1
schedule: "*/10 * * * *"
...go deeper
Know that Kyverno can delete resources on a schedule as well as gate them, and that a policy needs to say which kind, which condition makes it finished, and how often to sweep.
Explain the anatomy - match, conditions, cron schedule - contrast it with a per-object TTL label, and note that the cleanup controller deletes under its own granted permissions.
Show the operational care: prove the match selects only what you meant before enabling, scope narrowly, and argue for reporting instead of deleting when a wrong removal is unrecoverable.
Own the boundary between lifetimes the platform enforces centrally and lifetimes teams set on their own objects, and the standard that the least reversible automation gets the most review.
## Deletion as a policy outcome Most of a policy engine's outcomes are decided at admission: allowed, denied, mutated, or a companion object generated. Cleanup is the outcome that happens later, on the engine's own clock. Kyverno expresses it with a dedicated resource: `CleanupPolicy` in a namespace, `ClusterCleanupPolicy` across the cluster. The anatomy is deliberately similar to a normal rule: - **`match`** (and optional **`exclude`**) - which kinds and which objects, by name, namespace or selector. - **`conditions`** - expressions that narrow the set further, evaluated against the candidate resource, which cleanup policies expose as `target`. This is where you say *completed*, not merely *a Job*. - **`schedule`** - a cron expression. Each tick, the cleanup controller lists what matches and deletes it. So the sweep is declarative and fleet-wide: one object states that finished Jobs do not accumulate, everywhere, forever, without anybody writing a cron container. ## The TTL label: expiry attached to the object The second mechanism is a label, `cleanup.kyverno.io/ttl`, carried by the resource itself. Its value is either a duration or an absolute timestamp, and the same controller deletes the object once it expires. The distinction is which side owns the expiry: - **A cleanup policy** is a *class* rule. Every object matching this description is transient. The platform team writes it, and no tenant has to know it exists. - **A TTL label** is a *per-object lease*. This particular namespace goes away on Friday; that demo environment lasts four hours. The lifetime varies per object and is decided by whoever created it, which makes it the right fit for sandbox namespaces and short-lived test fixtures. They compose. A cleanup policy can sweep an entire class of leftovers while individual sandboxes carry their own TTL, and a generate rule that stamps the TTL label onto whatever it creates is a tidy way to give generated fixtures a bounded life without a second policy. ## RBAC, again The cleanup controller deletes under its own ServiceAccount. It only has the delete rights an administrator granted, so a policy pointed at a kind nobody granted appears to run on schedule and removes nothing. When a candidate says the first thing they check is the controller's permissions and its events, they have operated this rather than read about it. That identity is also worth thinking about as a threat surface, because it is deletion authority over whatever you scope it to. A cleanup policy is the one policy type whose *bug* destroys data rather than blocking a deploy, and a match block that is broader than intended - the kind without the conditions, an empty selector, a label that turns out to be on more objects than you thought - is a fleet-wide delete with no approval step. Two habits follow: scope match blocks as narrowly as the intent allows, and prove the selection before the deletion by listing what the same match would select today. ## Judgment: should this be a policy at all The interviewing point is not that Kyverno can delete things; it is when you should let it. Some cleanup is already native to the platform, and reaching for a policy to do what the workload API does for you adds a moving part, a controller dependency and a set of permissions for no benefit. So the honest ordering is: 1. If the resource's own API expresses the lifetime, use that. 2. If the lifetime belongs to the individual object and varies, use the TTL label. 3. If it is a class of leftovers nobody will remember to clean, use a scheduled cleanup policy. 4. If deleting it wrongly would lose data someone needs, do not automate the delete at all - report on it and let a human decide. Rule four is the one that separates a thoughtful answer from an enthusiastic one. Cleanup automation is the place where a policy engine has the least reversible outcome available to it, and treating it with less care than a blocking gate gets the priorities exactly backwards.
- The schedule fires but nothing is deleted. Where do you look?First at the cleanup controller's RBAC: it deletes under its own ServiceAccount and needs delete rights on that kind, so an ungranted kind produces a policy that ticks and removes nothing. Then at the conditions - an expression referencing a field that is absent on the candidate will not select it - and at the policy's own events and status.
- When would you refuse to write a cleanup policy at all?When a wrong deletion loses something a human needs and cannot recreate, or when the match block cannot be narrowed to exactly the intended class. Cleanup is the least reversible outcome a policy engine has. In those cases produce a report of what would be deleted and let someone act on it, rather than automating the delete.
- How do you sanity-check a cleanup policy's match block before enabling it?Run the equivalent selection as a read - the same kinds, namespaces, selectors and field conditions - and look at what comes back and how much of it. Then start narrow: one namespace or one label, verify the deletions were exactly the intended objects, and widen the scope only after a full cycle has run clean.
saying these in an interview costs you the question
- Thinks cleanup happens at admission time
- Forgets the controller needs delete permissions
- Writes a match block on kind alone with no conditions
- Treats an automated delete as lower risk than a block
- Confuses a per-object TTL with a fleet-wide schedule