Your OPA gate reports healthy but is still enforcing last quarter's rules — how do you diagnose it?
answer
- last good bundle keeps deciding
- readiness is not freshness
- compare loaded against intended
- split download failure from activation failure
- alert on time since last activation
basics
~20 sA failed bundle download or activation leaves the last activated bundle deciding, so OPA answers confidently on old policy. Compare the active revision in status and decision logs with the revision the service intended to serve.
solid answer
~50 sStart from the design fact: when a bundle fetch or activation fails, OPA does not fail closed and does not clear policy — it logs the error and keeps serving the last bundle that activated successfully. Nothing about a correct-looking decision proves the policy behind it is current. So get the **active revision** out of the status plugin's report or a recent decision log entry and compare it with the revision the bundle service is actually serving. If they differ, read OPA's bundle errors: a download failure (auth, TLS, 404, a 5xx) or an activation failure (the new bundle does not compile, or its roots conflict with another bundle). If they match, the publishing side never shipped the revision you think it did. And note why health did not catch it: `/health?bundles=true` answers "have the bundles activated at least once", not "is the active bundle current".
go deeper
Know that OPA keeps serving the last bundle it loaded when a new one cannot be fetched, so an old rule set can still be answering while everything looks healthy.
Explain where to find the active revision — the status report and decision log entries — and why a download failure and an activation failure produce the same symptom but different log lines.
Drive the investigation: compare loaded revision against published revision first, split the failure by phase, check the whole fleet rather than one pod, then add an alert on time since last successful activation.
Own the standard that makes this class of failure impossible to sit on: freshness as a monitored property of every decision point, and a publish pipeline that cannot ship a bundle the engine will refuse.
## The scenario A rule banning a deprecated API version — say the long-deprecated `extensions/v1beta1` Ingress that the cluster upgrade is about to remove — was added to the policy bundle at revision N and published six weeks ago. This morning someone notices that workloads using the deprecated API are still being admitted. The OPA pods are `Ready`, the bundle service is up, the policy pipeline is green, and the dashboard is entirely calm. ## The design fact everything follows from When a bundle download or activation fails, **OPA keeps the last successfully activated bundle and carries on answering queries with it.** It records the error and moves on to the next poll. This is deliberate: an agent in a request path should not start erroring because a bundle service had a bad five minutes. But it produces the failure mode above — an engine that is confidently, correctly, and quietly enforcing last quarter's rules. And the obvious health signal does not cover it. `GET /health?bundles=true` reports whether the configured bundles have been **activated at least once**. Once the first activation succeeds, later download and activation failures do not flip it back. It is a readiness check, not a freshness check. ## The diagnosis, in order **1. Establish what is actually loaded.** Get the active revision per bundle from the status plugin's report (console or the configured HTTP sink), and cross-check with a recent decision log entry, which carries the revision that was active when that decision was made. The decision log is the stronger evidence because it ties a specific admitted request to a specific policy build. **2. Establish what should be loaded.** Ask the bundle service what it is serving at that resource path, and map that revision back to a commit. Now you have a comparison rather than a suspicion. **3a. Revisions differ — the agent is stuck.** Read OPA's logs and the status report's error field for that bundle, and split by phase: - **Download failed.** Expired credentials for the bundle service, a rotated CA or TLS name mismatch, a 404 after someone renamed the resource path, 5xx from the service, a network policy or egress rule added since. The last successful download timestamp tells you when it stopped. - **Download succeeded, activation failed.** The bundle came down but did not install: the new policy does not compile, its declared roots conflict with another configured bundle, or content falls outside the declared roots. This is the nastiest variant because the transport is healthy — the tarball arrives every poll and is thrown away every poll. **3b. Revisions match — the agent is fine.** Then the bundle you believe was published never was. The build ran, but the publish step failed or wrote to a different path or a different environment's service, and the revision string it advertises is stale too. The pipeline being green is not evidence; the artifact the service is serving is. **4. Check the fleet, not one pod.** Activation is per process. Replicas poll independently on jittered intervals, so a partially stuck fleet is normal during a rollout and a persistent split — some replicas on N, some on N-1 — points at something environmental about those pods rather than at the bundle itself. ## What to change afterwards The repair is easy once found; the prevention is the interesting answer: - **Alert on bundle age, not on process health.** Time since last successful activation crossing a threshold is the signal that would have paged someone in week one. Bundle download and activation error counters are the second. - **Assert the revision, not just the behaviour.** A post-deploy check that reads the active revision from status and compares it with the revision the publish job produced turns a silent divergence into a failed deploy. - **Put the revision where investigators already look.** Decision logs carrying the revision are what let you answer, afterwards, exactly which requests were judged by stale policy — which is the question that gets asked once the deprecated API finally breaks something. - **Fail the publish, not the agent, on a bad bundle.** A compile check in the publishing pipeline stops an unactivatable bundle from ever being served, which removes the worst variant of this failure entirely. ## The sentence that shows you have lived it "A green gate proves the engine answered, not that it answered with the policy we shipped." Everything above is the machinery for closing that gap.
- Why did the bundle health check not catch this?Because `/health?bundles=true` answers whether the configured bundles have activated at least once, not whether the activated bundle is current. After the first successful activation it stays green through every later download or activation failure. It is a readiness gate for startup, and treating it as a freshness signal is the mistake that lets this run for weeks.
- What single alert would have caught it in week one?Time since the last successful bundle activation, per agent, alerted above a threshold that is a small multiple of the poll interval. It catches every variant — download failures, compile failures, roots conflicts — because they all end in the same observable: activation stopped happening. Pair it with a bundle-error counter so the page tells you which phase failed.
- The status report shows the revision you expected, but the rule still is not blocking. Where do you look?At the publishing side, not the agent. The revision string is an opaque label the publisher sets; a build that advertises revision N while packaging the previous commit's policy is entirely possible. Resolve the revision back to a commit, confirm the rule is in that commit, and confirm the built tarball contains it — then check whether every replica reports it, since activation is per process.
- Half the replicas report the new revision and half report the old one. What does that tell you?That activation is per process and these agents poll independently on jittered intervals, so a brief split during a rollout is expected. A split that persists is environmental to the lagging pods — different credentials, a different service endpoint, an egress restriction, or a pod that has not restarted since a config change. Compare the last successful download timestamp on both sides.
It is a wall clock that stopped at a plausible hour: everyone who glances at it gets a confident answer, and nothing about the answer looks wrong until you compare it with a clock that is still running.
saying these in an interview costs you the question
- Assumes a failed bundle fetch makes OPA fail closed
- Treats a ready pod as proof policy is current
- Never compares the active revision to the published one
- Blames the rule before checking whether it loaded
- Ignores that a downloaded bundle can still fail activation