skip to content

Your policy service is healthy and answering fast but loaded an empty rule set - how do you detect that?

level: seniorimportance: nice to knowfreq 26%

answer

  1. up is not the same as enforcing
  2. no default fires - it did answer
  3. prove the engine can still say no
  4. an input that must always be denied
  5. an allowed canary means treat it as down

basics

~10 s

Health checks only prove the process is up. Detect it behaviourally: evaluate a synthetic input that must always be denied, continuously, and treat an allow on that probe as an outage.

solid answer

~50 s

This is the third failure the fail-open/fail-closed binary hides. The service is reachable, fast and confident, so no unavailability default ever fires - it simply answers allow to everything, because it has nothing to match against. Liveness and latency monitoring will all be green. The detection has to be behavioural: keep a canary input that violates a rule which is never legitimately absent - a workload requesting privilege escalation, say - and evaluate it continuously through exactly the path a real call takes. If the canary comes back allowed, the enforcement layer is not enforcing, and the caller should treat the service as unavailable rather than trusting it, so the fail-closed subset applies and someone is paged. Deny-rate anomalies are a useful secondary signal, but a genuinely quiet day looks identical, so the probe is the primary one.

go deeper

for a junior

Take away the core fact: a policy service can be up and fast while enforcing nothing, so a health check that only proves the process is running is not evidence that rules ran.

for a middle

Explain why standard monitoring misses this - liveness, latency and error rate are all green with an empty rule set - and what a synthetic must-deny input tests that none of them do.

for a senior

Design the probe: which rule it violates, that it runs through the real call path from the caller's side, how often, that allow is the failure condition, and what the gate does with the result.

for a principal

Frame it as an assurance problem: the organisation is claiming a control operated continuously, so it needs continuous evidence the control could still say no, not just evidence that a service was reachable.

## The failure the binary does not cover Fail open and fail closed both assume you can tell that no answer arrived. The nastier state is one where an answer *does* arrive - promptly, with a 200, in the right shape - and it is allow, because the engine has no rules loaded, or loaded a rule set that matches nothing you care about. Nothing in the unavailability design fires, because from every angle the service is healthy. The gate reports pass. The pipeline is green. Every change in the organisation ships unassessed, and it looks exactly like a period of unusually good compliance. How it happens is mundane: a rule set that failed to load and left the engine running with an empty one; a configuration change that points at an empty or wrong location; a deploy that shipped an unfinished rule set; a filter or selector that no longer matches the inputs being sent. None of these crash anything. ## Why ordinary monitoring misses it The usual signals answer the wrong question: - a liveness or readiness probe answers *is the process running*; - a latency alert answers *is it answering quickly* - and an engine with no rules is beautifully fast; - a request-rate metric answers *is anyone calling it* - they are; - an error-rate metric answers *is it throwing* - it is not. None of them answers the only question that matters: **is it still saying no to things it should say no to?** ## The canary deny probe The direct test is a synthetic input constructed so that it must always be denied, evaluated on a schedule. Properties that make it trustworthy: - **It violates a rule that is never legitimately absent.** Pick something structural to your policy - a workload asking for privilege escalation is the canonical choice, because no rule set you would ever ship permits it. If the rule can be legitimately removed, the probe generates false alarms and gets muted, which is worse than not having it. - **It runs through the real path.** Same endpoint, same input shape, same client library, same network route as a real evaluation, so it exercises loading, matching and response handling. A component that asks itself whether it is healthy proves much less. - **It runs continuously**, at a frequency that makes the undetected window shorter than the time it takes a change to travel through the gate. If a deploy takes ten minutes and you probe hourly, you can still ship an hour of unassessed changes. - **Allow is a failure.** The probe is inverted relative to normal health checks: the passing condition is a denial. Wire it that way explicitly, or somebody will "fix" the alert by asserting a 200. A useful extension is a small set of probes rather than one - a must-deny input and a must-allow input. The must-allow probe catches the opposite fault, where a broken or over-broad rule set denies everything, which is loud but often misdiagnosed as an engine outage. ## Secondary signals Deny-rate monitoring helps: if a gate normally denies a handful of changes a day and has denied nothing for two days, that is worth a look. But it is a weak signal on its own. Low-traffic gates and genuinely clean weeks look identical to a dead rule set, so it is an investigation trigger, never the alarm. Asserting a rule count at start-up is likewise a partial control. It catches the empty-load case, but a rule set can load a plausible number of rules and still match nothing, and a check made once at start-up says nothing about a reload six hours later. ## What the caller should do when the probe fails The answer is not "alert a dashboard and keep going". The evidence says the engine's answers are worthless, so the caller should treat it as unavailable and take the degraded path: the small fail-closed subset denies, everything else degrades with a loud signal, and a human is paged. The pathological outcome to avoid is a red monitor beside a gate that cheerfully keeps allowing. Afterwards, reconcile: the window between the last known-good probe and the alert is the set of changes that went out unassessed, and re-evaluating them against the restored rule set is the only way to close it. ## Why interviewers like this question It separates people who think about a gate as a configuration from people who have operated one. The insight it tests is that **up is not the same as enforcing**, and that the only proof of enforcement is a rule still firing on something you deliberately fed it.

  • Where should the probe run from, and how often?
    From the caller's side, through the exact path a real evaluation takes - same endpoint, same input shape, same client - so it exercises loading and matching the way a real call does. Something that probes itself from the inside proves much less. Run it often enough that the undetected window is shorter than the time it takes a change to travel through the gate, or you can still ship an unassessed window between probes.
  • What should the calling gate do when the canary comes back allowed?
    Treat the service as unavailable rather than trusting the answer. It is confident but the evidence says the rules are not there, so the degraded path applies: the small fail-closed subset denies, the rest degrades loudly, and someone is paged. The failure to avoid is a red dashboard next to a gate that keeps allowing changes because the service technically responded.
  • Why not just assert the rule count when the engine starts?
    It catches the empty-load case and is worth having, but it is weaker than a probe. A rule set can load a plausible count and still match nothing that is being sent to it, and a check made once at start-up says nothing about a reload hours later. A must-deny probe tests the behaviour you actually depend on, continuously.

A smoke alarm with the battery removed still hangs on the ceiling looking perfect. The only way to know is to press the test button, on a schedule.

saying these in an interview costs you the question

  • Assumes a passing health check proves enforcement
  • Monitors only uptime, latency and error rate
  • Reads a green pipeline as proof rules ran
  • Treats a week with zero denials as good news
  • Alerts on the failed probe but keeps allowing changes

context