skip to content

Acting on the Answer

A decision is not yet an outcome: something must happen when the engine is unreachable, when the answer is no, and when the rule cannot express what the reviewer actually cares about.

on this pageshow

explore

questions

11

A policy gate blocks your deploy with only 'denied by policy' — what should that denial have carried?

level: juniorimportance: must knowfreq 66%

answer

  1. the developer must act without asking anyone
  2. identity, location, requirement, remedy
  3. point at the exact array index
  4. all violations, not just the first

basics

~20 s

A usable denial names the rule that fired, the object and the exact field it failed on, what the rule requires, and one concrete fix. It should list every violation found, not only the first.

solid answer

~40 s

A bare `denied by policy` tells the developer that something is wrong and nothing about what. A denial has to answer four questions on its own: which rule fired (a stable identifier, not just a sentence), which thing failed it (the resource, and the exact field path inside it — `spec.template.spec.containers[2].resources.limits.memory`, not "a container"), what the rule actually requires, and what to change to satisfy it. It should also report every violation in the change rather than stopping at the first, otherwise the developer burns one submit-and-fail cycle per problem. The test is simple: can someone who has never read the policy fix their manifest from the denial text alone, at 17:00, without opening a ticket? If not, the gate has moved work onto the person least equipped to do it.

go deeper

for a junior

Be ready to list what you personally needed the last time a gate blocked you: which rule, which object, which field, what to change. Recall that the field path with its array index is the part that saves the most time.

for a middle

Explain why the payload is part of the rule's design, not an afterthought, and why an engine should evaluate every rule rather than short-circuit on the first failure.

for a senior

Show that you measure a gate by time-to-fix and support load, not just block rate, and that you test your own denials by breaking something deliberately and reading the output cold.

for a principal

Own the argument that unactionable denials are how guardrails lose consent: teams that cannot self-serve a fix start routing around the gate, and the exemption backlog becomes the real risk.

## The same block, two very different costs A policy gate makes a binary decision — this change goes through or it does not. But the decision is not the whole outcome. What the engine *emits* alongside "no" determines whether the block costs the developer thirty seconds or half a day. Both denials stop the same bad change; only one of them also tells the person how to unblock themselves. Take a concrete rule: **every container must declare a memory limit**. A developer submits a Deployment with five containers — an app, two sidecars, an init container and a log shipper. One of them has no `resources` block at all, so it has no memory limit. The gate refuses the change. Denial A: `denied by policy`. Denial B: > Rule `container-memory-limit-required` (high): container `log-shipper` in Deployment `payments/checkout-api` declares no memory limit. Offending field: `spec.template.spec.containers[3].resources.limits.memory`. Fix: set `resources.limits.memory` on that container (team default `512Mi`). Denial A makes the developer diff five containers by hand against a policy they have not read, guess which of several rules fired, and — most likely — ask someone. Denial B is a fix. ## The four things a denial must carry **1. Which rule fired.** Not just a sentence, but an identifier the developer, a bot and a dashboard can all refer to. Without it, two different problems that produce similar prose are indistinguishable, and there is no way to look the rule up, cite it in a discussion, or count how often it fires. **2. Which thing failed, and where inside it.** This is the field most often missing and the one that does the most work. Identity of the resource (kind, namespace, name — or file and line, if the input was a source document) tells you *which* object out of the many in a change. The **field path** tells you where inside that object. `spec.template.spec.containers[3].resources.limits.memory` is locatable in seconds; "a container is missing limits" is a search. Note that the array index matters: with five containers, "a container" is a one-in-five guess. **3. What the rule requires.** The developer needs the requirement, not just the fact of failure. "Must declare a memory limit" is the requirement. A one-line reason for the requirement ("an unlimited container can starve its neighbours on the node") is what stops the rule being read as arbitrary bureaucracy, and it is cheap to include. **4. How to fix it.** A remediation hint that names the field to set and, where the organisation has one, the value or the source of the value. "See the policy documentation" is not remediation; it is a redirect. A link *in addition to* a concrete instruction is fine. ## Report all of them, not the first An engine that short-circuits on the first failing rule turns a change with three problems into three submit-wait-fail cycles. If each cycle costs a pipeline run, that is minutes to tens of minutes of wall-clock time and three context switches, for information the engine already had on the first pass. Evaluate everything and return the full list. The one legitimate exception is a violation that makes the rest of the evaluation meaningless — an input that failed to parse, for example, where the remaining rules would only produce noise. ## The cost you do not see on the dashboard The cheapest thing to measure about a gate is its block rate. The expensive thing is **time-to-fix** — how long between the block and the corrected change — and the support load it generates. A bare deny does not reduce the number of blocks; it converts each one into a conversation. Those conversations land on the platform or security team that owns the gate, which is exactly the team that has the least context on the change being blocked. A team that keeps getting stopped by messages it cannot act on stops treating the gate as a helpful colleague and starts treating it as an obstacle to route around — which is how exemption requests, disabled checks and "just merge it, we'll fix it later" habits get started. ## The check to apply before you ship a rule Write the rule, then deliberately break something and read your own denial as if you had never seen the policy. Can you locate the offending field without opening the rule source? Can you fix it without asking anyone? If the answer is no, the rule is not finished — the payload is part of the rule, not decoration on top of it.

  • If a change trips five rules at once, should the gate report all five or stop at the first?
    Report all five. Stopping at the first turns one block into five submit-wait-fail cycles, each costing a pipeline run and a context switch, for information the engine already computed. The only sensible exception is a failure that invalidates the rest of the evaluation — an input that did not parse, say — where continuing would produce noise rather than findings.
  • How specific does the remediation hint need to be?
    Specific enough to act on without a second source: name the field to set and, where one exists, the value or where to get it — "set `resources.limits.memory`; the team default is `512Mi`". A bare link to policy documentation is a redirect, not remediation. A link alongside a concrete instruction is useful; a link instead of one is what generates the ticket.
  • Is a screenshot-friendly prose message enough, or does the denial need structure too?
    Prose alone is enough only for the human reading it right now. Anything else that wants to act on the denial — a bot annotating the pull request at the offending line, a dashboard counting which rules fire most, a metric on time-to-fix — needs discrete fields it can read. In practice a gate emits both: structured fields for machines, rendered text for people.

A compiler that said only "compile error" would be useless. The file, line, column and the expected token are what make an error message a fix instead of a hunt.

saying these in an interview costs you the question

  • Thinks 'blocked by security policy' is a sufficient message
  • Assumes the developer can find the offending container themselves
  • Offers a documentation link instead of a field path
  • Stops at the first violation and calls it a clean design
  • Puts the detail only in the gate's own server logs
  • Treats message quality as cosmetic rather than part of the rule

context

open as a page

A deploy tool calls a policy decision service and gets no answer - what do fail-open and fail-closed mean here?

level: juniorimportance: must knowfreq 66%

basics

~20 s

Fail open means the deploy proceeds when the decision service cannot answer; fail closed means it is rejected. Open protects delivery and lets unassessed changes through; closed protects the guardrail and can halt every deploy in the organisation.

open as a page

A build-time policy check verifies a container image's base image is on the approved list — what can it not establish?

level: juniorimportance: must knowfreq 61%

basics

~20 s

It establishes one fact: the recorded base image matches a list. It cannot establish whether an unapproved base was justified, whether the approved one was the right choice, or whether anything layered on top of it is safe.

open as a page

What fields should a policy violation record carry beyond the human-readable message?

level: middleimportance: should knowfreq 48%

basics

~20 s

A violation record needs discrete fields other systems can read: a stable rule identifier, severity, the outcome, the resource identity, the offending field path, and a remediation hint. Prose cannot be filtered, aggregated or placed on a line.

open as a page

What middle designs sit between fail-open and fail-closed when a policy decision service is unavailable?

level: middleimportance: should knowfreq 48%

basics

~20 s

Split rules by what they protect: a small named subset stays fail-closed, everything else degrades to a static allow. Or answer from a last-known-good rule set with an expiry and a loud signal that the fallback is in use.

open as a page

An auditor asks what your approved-base-image gate actually establishes — what controls do you name beside it?

level: middleimportance: should knowfreq 44%

basics

~20 s

Name the gate as a preventive check on one recorded fact, the human review that judges whether a deviation was necessary, and runtime detection that watches what the image does once it runs. Say what each decides.

open as a page

Your approved-base-image rule has grown a dozen special-case clauses and still blocks legitimate builds — what do you change?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Stop encoding judgement in the predicate. Split the rule: keep a hard block for the part that is mechanically decidable, and demote the rest to a signal that routes the build to a human review instead of failing it.

open as a page

At 02:00 your policy decision service is down and every deploy is blocked - do you flip the default to allow?

level: principalimportance: should knowfreq 38%

basics

~20 s

Only as a narrowly scoped, time-boxed change that expires by itself, and never for the small set of rules whose whole purpose is stopping unassessed privilege changes. Restoring service or serving last-known-good rules is the better move.

open as a page

Your bots and dashboards parse denial message text — what breaks when you reword the message?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

Anything keyed on the sentence breaks silently: suppressions stop matching, dashboards split one rule into two series, dedupe fails. Nothing errors — the wording was an undocumented API. Give consumers stable fields and declare the message unstable.

open as a page

Your policy service is healthy and answering fast but loaded an empty rule set - how do you detect that?

level: seniorimportance: nice to knowfreq 26%

basics

~10 s

Health checks only prove the process is up. Detect it behaviourally: evaluate a synthetic input that must always be denied, continuously, and treat an allow on that probe as an outage.

open as a page

Your policy program's human-review queue is the bottleneck — how do you decide which decisions stay with people?

level: principalimportance: nice to knowfreq 33%

basics

~20 s

Treat review capacity as a fixed budget. Automate every decision the artifact fully determines, spend the remaining review on the few classes where necessity genuinely matters, and explicitly accept the rest rather than queueing work nobody will do.

open as a page