skip to content

Your OPA gate denies an instance type approved this morning. How do you diagnose and unblock it?

level: seniorimportance: should knowfreq 55%

answer

  1. the rule is not the suspect
  2. ask the engine, not the source of truth
  3. which replica answered matters
  4. restart loses whatever was pushed
  5. stamp the facts with an age

basics

~20 s

Read what the running engine holds, not what the source of truth says: GET the approved list from that OPA and compare. A correct rule denying a newly approved value is a data-freshness failure, fixed on the data path.

solid answer

~50 s

First separate the rule from the facts. `GET /v1/data/approved/instance_types` against the instance that made the decision shows exactly what it holds; if the new class is missing, the rule is behaving correctly on old data. Then find which loading path owns that fact. If it comes in a bundle, check whether the bundle was rebuilt after the approval and whether this replica activated it — what the bundle service serves and what a replica has activated are different questions. If it is pushed over the Data API, check whether the job ran, whether it reached every replica, and whether any replica restarted since and lost it. Unblock by re-driving that path, never by hand-writing the value into one replica. Then fix the diagnosis time: stamp the facts with a generation time, surface it in the denial message, and alert when it ages past a threshold.

go deeper

for a junior

Know that a denial can come from the facts being old rather than the rule being wrong, and that you can read the facts back out of a running OPA over its Data API.

for a middle

Be ready to walk the data path backwards: which loading path owns the fact, whether the bundle was rebuilt and activated on this instance, whether the push job ran and whether the pod restarted since.

for a senior

Demonstrate the separation of rule from facts under pressure, an unblock that converges every replica rather than one, and the observability you add afterwards so staleness pages someone before it blocks anyone.

for a principal

Own the freshness contract with whoever owns the source of truth: the propagation target, who is paged when facts age out, and what developers are entitled to when the platform's copy of the truth lags.

### The shape of the failure The rule is right. The organisation's decision is right. The engine is answering correctly about a world that is six hours old. This is the characteristic failure of any guardrail whose facts are replicated into the decision point, and it is worth naming out loud in an interview because the instinct — go and read the Rego — wastes the first twenty minutes. ### Diagnose in order **1. Ask the engine what it holds.** `GET /v1/data/approved/instance_types` on the instance that produced the denial. This is the whole diagnosis in one request: either the newly approved class is there, in which case you have a rule or an input-shape problem, or it is not, in which case you have a data problem and the rule is exonerated. Note *the instance that produced the denial*: with several replicas or a sidecar per pod, asking a different one answers a different question. **2. Establish which path owns the fact.** Bundle or push. This determines everything downstream, and on a mature platform it should be written down rather than rediscovered under time pressure. **3. Follow that path backwards.** For a bundle: was it rebuilt after the approval landed in the source of truth, was the new revision published, and did *this* replica fetch and activate it? The activated revision on the instance and the revision the bundle service is serving are separate facts, and a failing activation is invisible unless you look — OPA keeps serving with what it already had. For a push: did the job run after the approval, did it write to every replica, and has any replica restarted since? A restarted OPA comes back with only what it loads at boot, so a pod that restarted after the last push holds pre-push facts indefinitely. **4. Rule out the total failure.** Consider a rule shaped `not data.approved.instance_types[input.instance_type]`. If the whole subtree is missing — a data file dropped from the bundle, a push that wrote to the wrong path — that reference is undefined for *every* value, the negation succeeds every time, and the gate denies everything. If the reports are "everyone is blocked" rather than "one class is blocked", check for absent data before anything else. ### Unblock without making it worse The tempting fix is a `PUT` of the missing value into the OPA that is blocking this developer. Resist it. It makes the decision depend on which replica served the request; it leaves no record of who changed the guardrail's facts; and it evaporates on the next bundle activation or restart, so the same developer is blocked again next week with the diagnosis already "solved". Re-drive the owning path instead — trigger the bundle build and publish, or re-run the sync job — so that every instance converges and the fix is reproducible. If the propagation genuinely cannot be made to happen in time, the honest move is to route the developer to the platform's documented exception process rather than to quietly mutate one engine. ### What the platform owes the developer From the blocked engineer's chair, `instance type m7g.large is not on the approved list` is a false statement, and being told a false statement by a system that will not let you ship is what turns guardrails into adversaries. Three cheap changes remove most of that pain: - **Stamp the facts.** Ship a `generated_at` (and, for bundles, the revision) alongside the list, and include both in the denial message: *not on approved list, revision 41, generated 14h ago*. The developer now knows in one glance whether to argue with the policy or ask about propagation, and your on-call gets a bug report that already contains the diagnosis. - **Alert on age, not just on errors.** A push job that silently stops still leaves a perfectly healthy OPA answering perfectly wrong questions. Monitor the age of the facts inside the engine against the expected refresh interval and page on it, because nothing else will notice. - **Canary the decision.** Periodically evaluate a known-good and a known-bad input against the live engine. A known-good input that starts being denied catches an empty or misplaced data document long before a human hits it. ### The sentence to say out loud "The rule enforced the policy correctly against the facts it had; the facts were stale, so the gate was wrong. That is a freshness bug in the data path, and the fix is propagation and observability, not a policy edit." Interviewers are listening for exactly that separation — because the candidate who edits the rule to unblock someone has just permanently weakened a control to fix a sync job.

  • What single field in the data document would have turned this into a five-minute diagnosis?
    A generation timestamp shipped with the facts. Surfaced in the denial message it tells the blocked developer immediately that the list is hours old, and monitored against the expected refresh interval it pages someone before anyone is blocked at all. Without it, staleness is indistinguishable from a correct denial.
  • The developer asks you to just push the value into OPA directly. Why is that a poor unblock?
    It writes to one replica, so the answer now depends on which pod serves the request; it leaves no reviewable record of who changed the guardrail's facts; and it disappears at the next bundle activation or pod restart, so the same block returns with everyone believing it was fixed. Re-drive the owning path instead.
  • Reports come in that the gate is denying every workload, not just one class. What do you suspect?
    That the data subtree the rule reads is absent rather than out of date. A lookup into missing data is undefined for every value, so a rule that denies when the lookup does not resolve denies universally. Check that the path exists in the running engine before looking anywhere else.

saying these in an interview costs you the question

  • Reading the Rego first instead of the engine's data
  • Trusting the bundle service's contents as what the replica holds
  • Hand-writing the fact into one replica and calling it fixed
  • Loosening the rule to unblock a stale-data denial
  • Denial messages that name no data revision or age

context