A Gatekeeper constraint blocked a Deployment for a missing NetworkPolicy applied seconds earlier — why, and what do you do?
answer
- the rule was right about what it saw
- the cache trails the API server
- watch events are not synchronous
- retry succeeds, audit is clean
- fix provisioning order, not the Rego
basics
~10 sReplication lag. A Gatekeeper rule reads data.inventory, a watch-fed replica rather than a live query, so the NetworkPolicy existed in the API server but had not reached the cache when the decision was made.
solid answer
~40 sThis is replication lag, not a broken rule. Gatekeeper's sync controller watches the kinds you configured and mirrors them into an in-memory document; a rule reading `data.inventory` sees that mirror, which trails the API server by however long the watch and the write take. Applying a namespace's NetworkPolicy and its workloads in one batch is enough to lose the race. The signature is diagnostic: the identical manifest succeeds on an immediate retry, and the constraint's audit results are clean. The fixes are operational, not Rego. Make the default-deny NetworkPolicy part of namespace provisioning so it exists long before any workload; sequence the pipeline so the policy is applied first; and treat an inventory-backed rule as an eventually-correct control, keeping the periodic sweep as the backstop for what admission read too early.
go deeper
Know that a rule reading replicated cluster data sees a copy that can be slightly behind reality, so an object created a moment ago may not be visible yet.
Explain the mechanism: a watch-driven sync controller fills an in-memory document, evaluation reads that document, and nothing in the request path waits for the mirror to catch up.
Diagnose it from the signature — succeeds on retry, clean audit, same-batch apply — and fix it where it belongs: provisioning order, a non-blocking rollout, and the sweep as the backstop.
Decide how much eventual correctness the organisation will accept from a blocking gate, and set the pattern that dependencies like a namespace's default-deny policy are provisioned by the platform rather than shipped with each application.
## What actually happened Gatekeeper answers `data.inventory` questions from a cache it fills itself. A sync controller opens watches on the kinds listed in the Config's `syncOnly` and in any SyncSet, and mirrors those objects into an in-memory document that the Rego evaluation reads. Nothing in that path is synchronous with the request being admitted. So the timeline of the incident is: 1. `kubectl apply -f ./namespace/` creates the NetworkPolicy. The API server persists it and acknowledges. 2. Milliseconds later, the same apply creates the Deployment. The API server calls Gatekeeper's webhook. 3. Gatekeeper evaluates the rule against its cache. The watch event for the NetworkPolicy has not been processed yet, so the inventory branch for that namespace is empty. 4. The helper that looks for a default-deny policy does not resolve, the violation fires, and the Deployment is rejected with a message that is *true of the cache and false of the cluster*. This is the ordinary behaviour of an eventually consistent replica. It is not a bug in the rule, and rewriting the Rego will not fix it — there is no expression that can wait, retry or reach past the cache. ## How to recognise it quickly Three signals, in the order they are cheapest to check: - **The same manifest succeeds on retry**, with no change to the constraint, the template or the cluster. A logic error would reject it every time. - **The constraint's audit results are clean** for that namespace on the next sweep — the periodic sweep runs later, by which time the cache has caught up, and reports no violation for the workload that admission rejected. - **The rejected object was applied in the same batch** as the object it depends on, or within a second or two of it. If instead every namespace is being rejected, including ones whose policy has existed for weeks, you are not looking at lag: you are looking at a kind that is not replicated at all, or a wrong path in the rule. ## The same property in the other direction Lag is symmetric, and the reverse case is the one that matters for security. Delete the default-deny NetworkPolicy and it remains visible in the inventory for a moment afterwards. A workload created in that window is admitted on the strength of a fact that is already gone. No rule reading a replica can close that gap, which is why an inventory-backed control is best understood as **eventually correct**, and why admission alone is not the whole control. ## What to actually do **Fix the ordering, not the rule.** The durable answer is that a namespace's default-deny NetworkPolicy should be created as part of provisioning the namespace, not shipped alongside the workloads that depend on it. Then it predates every workload by minutes or days and the window never opens. Where the policy really does travel with the application, sequence the delivery so it is applied and observed before the workloads go out — most delivery tooling can express that as a wave or a dependency, and a retry loop on the workload apply is a cruder version of the same idea. **Keep the periodic audit as the backstop.** Admission catches things at the moment of change; the sweep re-evaluates what is actually in the cluster afterwards. Anything admission got wrong because it read too early shows up there, and anything admitted during a delete window shows up there too. Treat the two together as the control, and cite them together when someone asks how you know the requirement holds. **Roll out non-blocking first.** When a new inventory-backed constraint goes into an estate, a period of reporting-only enforcement tells you how often the ordering problem exists before it becomes a wall of failed deployments at 5pm. **Say so in the message.** A violation message that reads *"no default-deny NetworkPolicy visible for namespace X — note this is read from a replicated cache and may lag a policy created moments ago; retry if you just applied one"* costs nothing and converts a support ticket into a self-service fix. Blocked developers do not know what your engine reads from, and the message is the only place you get to tell them. **Do not reach for a longer webhook timeout.** Nothing about this is slow; the cache is fast and confidently wrong. Extra time in the request would not change the answer. ## What not to conclude The wrong lesson is *"inventory rules are unreliable, disable the constraint."* The right one is that the replica's freshness is an operational property you design around: provision the dependency early, back admission with the sweep, and write the message so the person on the wrong end of the decision knows what they are looking at.
- Does the same lag work in the other direction, and does that matter?Yes, and it matters more. Deleting the NetworkPolicy leaves it visible in the cache briefly, so a workload created in that window is admitted on a fact that no longer holds. No rule reading a replica can close that window, which is why the periodic sweep over live objects is part of the control rather than a nice extra.
- The team asks you to make the rule wait for the cache to catch up. What do you say?It is not expressible. Rego in a ConstraintTemplate has no sleep, no retry and no API client — evaluation is a pure function of the request and the data document as they are at that instant. The fix lives outside the rule: provision the dependency before the workloads, sequence the apply, or let the client retry.
- How would you tell this apart from the kind simply not being replicated?By blast radius and repeatability. Lag hits a namespace that was just changed and clears on retry; an unsynced kind rejects every namespace forever, including ones whose policy is weeks old. Check whether a long-standing namespace also fails — if it does, look at the sync configuration and Gatekeeper's RBAC, not at timing.
saying these in an interview costs you the question
- Assumes data.inventory is a live read of the API server
- Rewrites the Rego to fix what is a timing problem
- Proposes a sleep or retry inside the rule
- Raises the webhook timeout to give the cache time
- Concludes the engine is broken and disables the constraint