skip to content

Before adopting an in-cluster policy engine, what do you require of its upgrade story and who owns it?

level: principalimportance: nice to knowfreq 40%

answer

  1. who carries this pager in a year
  2. the engine sits in the write path
  3. rehearse the upgrade without a cluster
  4. a back-out on-call can run alone
  5. rule authors versus engine operators

basics

~20 s

Name the team that will upgrade it and carry its pager before you install it. Require a rehearsable upgrade: pinned versions, documented schema changes to the policy resources, rule tests re-run first, and a back-out on-call can execute.

solid answer

~50 s

The engine is not a security tool sitting to one side; it is infrastructure on the write path of every API request its rules match, so a bad upgrade is a cluster availability event rather than a policy bug. Before adopting one I want: a pinned version and a predictable release cadence; documented changes to the policy custom resources between versions, because those are the API my rules are written against; the ability to re-run the whole rule test suite offline against the new engine binary before it touches a cluster; a staged rollout through a non-production cluster; and a written back-out that on-call can execute at 3am — relax or remove the webhook configuration, accept that the cluster is briefly unenforced, and reconcile afterwards with the background report. Then ownership: if security authors the rules and platform runs the engine, write down who approves a rule, who approves an upgrade, and who may disable it in an incident.

go deeper

for a junior

Understand that the policy engine is a running service someone maintains and upgrades, not a one-time install, and that when it is unhealthy it can affect ordinary deploys.

for a middle

Be able to describe what an upgrade actually touches: the engine's workloads, its webhook configuration and certificates, and the custom resource definitions your rules are written against.

for a senior

Show that you would rehearse the upgrade by re-running the rule suite against the new version, stage it through a non-production cluster, and keep a back-out that on-call can execute alone, followed by a reconciliation scan.

for a principal

Own the ownership split before adoption: who approves a rule, who approves an upgrade, who may disable the engine in an incident, and whether an expressiveness gain is worth an engine your on-call rotation cannot debug.

## Why the upgrade story is an adoption criterion An in-cluster policy engine is usually evaluated on what its rules can express. The chair that matters after adoption is different: someone on the platform team will be paged for this component and will move it through minor versions every quarter for years. From that chair the engine is infrastructure sitting in the path of API writes, and its properties as a running service outrank its properties as a language. The reason is structural. The engine is wired in by a webhook configuration, and if that configuration is set to fail closed, then while the engine's pods are unavailable, matching API requests are rejected. An engine that is down is therefore not merely "not enforcing" — depending on how it is wired, it can be the reason a deploy, a controller reconcile or an autoscaling event fails. That is why the upgrade is a change-managed event and not a routine bump. ## What to require before you install it **A pinned version and a known cadence.** You want to choose when to move, and you want to know roughly how often you will have to. An engine that ships breaking changes unpredictably makes the quarterly upgrade a research project each time. **Documented changes to the policy resources.** The engine's custom resource definitions are the API your rules are written against. An upgrade that changes their schema, tightens validation or deprecates a field is a change to every rule you own. You want release notes that state this plainly, and you want to know whether existing stored policies remain valid after the CRDs are updated. **A rehearsal that does not need a cluster.** This is where the criteria compose: because rules can be evaluated offline, you can run the entire rule suite against the **new** engine version before it goes anywhere near production, and see whether any verdict changed. An engine without offline evaluation cannot be rehearsed this way, so its upgrades are validated in a live cluster or not at all — that is the compounding cost of failing the first criterion. **A staged rollout.** Non-production cluster first, then production, with a soak in between long enough for the periodic background scan to run and for the normal deploy traffic of a working day to pass through. **A back-out an on-call engineer can execute.** Written down, tested, and short: how to relax or remove the webhook configuration so the API server stops depending on the engine, what that means (the cluster is unenforced until it is restored), and how to reconcile afterwards. The background report is what closes that loop: after enforcement is restored you scan to find whatever was created while the gate was off. ## The ownership question The more interesting half of the question is organisational, and it has a common failure shape: the security team authors and ships the rules, the platform team runs the engine, and the platform team's pager fires because of a rule they did not write and cannot immediately interpret. Adoption is the moment to settle this rather than discovering it during an incident. Concretely, write down: - **Who approves a new rule**, and against what bar — including whether it lands in a reporting-only state before it enforces anything. - **Who approves an engine upgrade**, and what evidence they expect (the rule suite re-run against the new version). - **Who may disable the engine during an incident**, without needing to find the rule's author at 3am. This authority has to sit with on-call. - **Who is accountable for the rules that exist**, because rules accumulate and the platform team should not inherit an unowned pile of them. ## The tradeoff a lead is expected to own This is where the expressiveness argument gets resolved honestly. A more expressive engine can encode rules the simpler one cannot, and that is a real benefit. Against it: an engine nobody on the on-call rotation can debug at 3am is a worse engine for that organisation regardless of what it can express, because the failure mode is not "a rule was hard to write" but "nobody could tell whether the engine or the workload was the problem". Where the split is severe, the usual resolution is not to pick expressiveness but to invest in the operability side — a runbook, a training session, dashboards for the engine's own health, and rules whose match blocks are narrow enough that the engine sees a fraction of the traffic. The related call is how many engines to run. Two engines double the on-call surface, the upgrade calendar and the number of places a denial could have come from. There are reasons to do it — a capability one has and the other lacks — but it should be a decision with a named owner and a plan to converge, not an accumulation. ## What a strong answer sounds like It names the team before it names the product, treats the upgrade as a rehearsed change with a back-out, ties the rehearsal to offline rule evaluation and the recovery to the background report, and states the decision rights between rule authors and engine operators explicitly.

  • Security wants the more expressive engine; platform will run it. How do you decide?
    Weight operability heavily, because the binding constraint is who can debug it during an incident, not who can write the cleverest rule. If the expressive engine wins on merits anyway, the price is explicit: a runbook, health dashboards, on-call training and narrow match blocks so the engine sees less traffic. What you do not do is let the team that will never be paged make the call alone.
  • The engine is the incident at 3am. What does the runbook's first step look like?
    Establish whether API requests are failing because the engine is unreachable, then execute the documented back-out — relax or remove the webhook configuration so the API server no longer depends on it. Record the window. Restore the engine, then run the background scan to find what was created while the gate was open and reconcile it. On-call must be able to do all of that without paging the rule's author.
  • How does offline rule evaluation change the upgrade itself?
    It turns the upgrade into something you can rehearse. Run the full rule suite against the new engine version before it touches a cluster and compare verdicts; a rule whose result flipped is found on a laptop rather than in production. Without that property the only place to validate an upgrade is a live cluster, which is why the two criteria are usually evaluated together.

saying these in an interview costs you the question

  • Chooses the engine on rule expressiveness alone
  • Has no documented way to disable the engine during an incident
  • Assumes the team writing rules will also operate the engine
  • Upgrades in production without re-running the rule suite first
  • Treats engine downtime as merely not enforcing, ignoring rejected requests
  • Runs two engines with no named owner or plan to converge

context