Why would the API server reject a ValidatingAdmissionPolicy for exceeding its CEL cost budget?
answer
- it fails before any traffic exists
- estimated at apply time, not measured
- the estimator reads the schema, not your data
- unbounded list inside unbounded list
- make the allowlist a map, not a list
basics
~20 sKubernetes statically estimates each expression's worst-case evaluation cost when the policy object is written, and refuses it if the estimate exceeds the limit. Nested iteration over collections the schema does not bound is the usual cause.
solid answer
~60 sCEL evaluation is bounded so admission cannot become a denial-of-service vector, and Kubernetes enforces that in two places. At runtime there is a per-request budget. But the rejection you hit first happens at *apply* time: when you create or update the policy object, the API server computes a static worst-case cost estimate for each expression and refuses to store a policy whose estimate exceeds the limit. The estimator has no data, so it reasons from the schema — where a list or string declares `maxItems` or `maxLength` it uses that bound, and where the field is unbounded it assumes a very large worst case. A single pass over containers is fine. A nested `all()` over two unbounded collections is a product of two very large numbers, and that is what blows the budget. The fixes are structural: avoid the nested scan by making the allowlist a map keyed by the value so membership is a key lookup, split one dense expression into several cheaper validations, and hoist repeated sub-expressions into `variables` so they are computed once.
code
yaml · 7 linesvalidations:
- expression: >-
!has(object.spec.template.spec.tolerations) ||
object.spec.template.spec.tolerations.all(t,
params.spec.allowedTaintKeys.exists(k, k == t.key))
message: "toleration key is not approved"
...go deeper
Know that CEL expressions in admission policies have a cost limit and that a policy can be refused when you apply it, not just when it evaluates.
Explain the difference between the static estimate at apply time and the runtime budget per request, and why an unbounded list in the schema makes the estimate explode.
Diagnose it concretely: find the nested iteration, reshape the parameter data so the inner scan becomes a key lookup, split into separate validations with distinct messages, and hoist shared work into variables.
Read a persistently unaffordable expression as a venue signal rather than a puzzle — the write path lends every rule a shared compute budget, and a requirement that will not fit inside it is telling you where it should run.
## Two budgets, not one Kubernetes bounds CEL in admission twice. 1. **A static estimate at write time.** When you `kubectl apply` the `ValidatingAdmissionPolicy`, the API server parses each expression and computes an upper bound on how much work it could possibly do. If that exceeds the per-expression or per-policy limit, the policy object is rejected outright — you never get to test it, because it was never stored. 2. **A runtime budget per request.** During actual evaluation there is also a cost budget across the policy's expressions, so a pathological input cannot make a stored policy consume unbounded CPU on the write path. The second is a safety net. The first is the one that surprises authors, because it fires before any traffic exists and its arithmetic is not about your real data at all. ## How the estimator reasons The estimator has no cluster and no sample object. It works from the OpenAPI schema of the resource being validated. Where a field declares a size bound — `maxItems` on a list, `maxLength` on a string, `maxProperties` on a map — it uses that number. Where the field declares no bound, it substitutes a very large worst-case size, because a client really could submit a list that long. The consequences follow directly: - **Iterating one unbounded list is linear in that large number.** Usually affordable. - **Iterating one unbounded list inside another is the product of two large numbers.** Almost never affordable. - **Regular-expression matching against an unbounded string is expensive**, and it is expensive per element if you do it inside a loop. - **Repeating the same sub-expression in several validations pays for it several times**, because each expression is estimated on its own. A rule that reads perfectly well to a human — "every toleration's key must appear in the approved list" — is a nested scan, and that is why it is refused. ## Fixing it **Change the data shape so the inner scan disappears.** If the parameter resource holds the approved keys as a map keyed by the key name rather than as a list, membership becomes `t.key in params.spec.allowedTaintKeys`, which is a key lookup rather than a scan of an unbounded list. One loop instead of two. This is the single most effective fix and it costs you nothing but a slightly odd-looking parameter resource. **Split one expression into several.** Each entry in `validations` is estimated separately against the per-expression limit, and splitting also gives each failure its own `message`, which developers reading the denial will thank you for. **Hoist with `variables`.** A named variable is computed lazily and reused, so a sub-expression referenced by three validations is not paid for three times. It also makes the expressions readable, which matters more than it sounds when someone has to review the rule. **Narrow what you iterate.** Reach for the specific field rather than mapping over a whole collection to find it. `object.spec.template.spec.nodeSelector['node-class']` is a lookup; building a list of every selector value and searching it is not. **Note what does not help.** `matchConditions` narrow which *requests* the policy applies to, which is excellent for correctness and for runtime work, but the static estimate is computed per expression on its worst case — narrowing the audience does not lower the estimate for the expression that was refused. ## When the fix is not a fix Sometimes the expression is expensive because the requirement genuinely involves a cross-product — every container against every approved entry, with a per-element pattern match. If restructuring the parameter data cannot collapse it, that is a real signal about venue: the requirement is asking for more computation than the API server's write path is willing to lend you, and it belongs somewhere with its own compute budget rather than the shared one every write shares. ## The interview shape This is a diagnosis question. The candidate who says "the policy is too complicated, I simplified it" has not learned anything transferable. The candidate who says "the schema does not bound tolerations or the allowlist, so the estimate is their product; I turned the allowlist into a map so membership is a key lookup, and I split the two conditions into separate validations with their own messages" has.
- Would adding a matchCondition bring the estimated cost back under the limit?No. `matchConditions` decide which requests the policy is evaluated for, which reduces real runtime work, but the static estimate is computed per expression against its own worst case. An expression refused at apply time is still refused. Fix the expression's shape, not its audience.
- How do `variables` help with cost?A named variable in `spec.variables` is evaluated lazily and its result reused, so a sub-expression shared by several validations is computed and charged once rather than once per validation. It is a readability win too — the validations end up reading like the requirement instead of like a chain of field paths.
- What is the runtime budget for, if the policy already passed the static check?The static estimate bounds what a policy could cost; the runtime budget bounds what a specific request actually costs while it is being evaluated. It stops a single pathological object from consuming disproportionate CPU on the write path, and an expression cut short by it is treated as a policy evaluation error rather than a pass.
saying these in an interview costs you the question
- Thinks the rejection means the expression has a syntax error
- Assumes the cost is measured against real objects
- Adds a matchCondition to lower a static cost estimate
- Does not know unbounded schema fields get worst-case sizes
- Keeps allowlists as lists and scans them inside a loop