skip to content

A policy rule you own blocks another team's deploy at 5pm: who answers, and what must the failure say?

level: seniorimportance: must knowfreq 58%

answer

  1. engine failure or working rule
  2. breadth and reproducibility separate them
  3. the message decides who is called
  4. name the team, not the person
  5. state the route to change the rule

basics

~20 s

Separate an engine failure from a working rule. The platform on-call owns the engine; the named rule owner owns the decision. The failure output must name the rule, the resource and property, the owning team and how to propose a change — otherwise every block routes to the platform team.

solid answer

~50 s

First triage which kind of failure this is. If the engine is erroring, timing out or denying broadly across unrelated rules and teams, it is an incident and the platform on-call owns it. If one rule denied one resource and reruns identically, the engine is working and the question — should this rule say that — belongs to the rule's named owner. The failure output is what makes that routing happen, so it has to carry the rule id, the exact resource and property that failed, the required value, one line on why the requirement exists, the owning team, a link to the rule's source, and the path for proposing a change to it. A message that says only "policy check failed" routes every block to whoever is visible, which is the platform team, and they cannot judge a rule they did not write.

go deeper

for a junior

Know that a blocked build should tell you which rule denied you, which resource and property failed, what value is required, and who owns that rule.

for a middle

Explain the two failure classes and how to distinguish them: a broad, erratic set of denials points at the engine, while one reproducible rule on one resource points at a working rule.

for a senior

Show that you design the failure output as the routing mechanism, and that you hold the line on the on-call diagnosing the engine rather than adjudicating a control they did not author.

for a principal

Own the operating agreement: a stated response window for rule owners, a runbook for the platform on-call, and metrics on where disputed denials actually land and whether they ever change a rule.

## Two very different failures wearing the same shirt At 5pm a deploy fails on a policy gate. Two causes look identical from the pipeline log: **The engine is the incident.** It cannot be reached, evaluation errors, it times out, or it is denying on rules and teams that have nothing to do with each other. The tell is breadth and irreproducibility: many rule ids, many teams, results that change between runs. This is a platform on-call problem — the decision service is down, and the correct response is an incident response, not a policy discussion. **The rule is working.** One rule id, one resource, one property, and a rerun against the same document produces exactly the same verdict. Nothing is broken. The question is whether the rule *should* say that about this template, and that is a question about the requirement, not about the system — so it belongs to the team named as the rule's owner, who chose the threshold and can judge whether this case is a genuine miss. Getting this triage backwards is expensive in both directions. Treating a working rule as an outage means the on-call is pressured to make a call about a control they did not author. Treating a genuine outage as "the rule blocked you" means an engine failure sits unfixed while a developer edits a template that was never the problem. ## The failure message is the routing mechanism Here is the part candidates miss: **who gets woken is decided by what the output says**, not by an org chart. A blocked engineer at 5pm contacts whoever the message makes it possible to contact. If the output is "policy check failed: 3 denials, exit 1", the only visible party is the platform team who put the gate in the pipeline, and they become the de facto owner of every rule in the estate. Ownership files do not fix this; the message does. A denial should carry, for each finding: - **The rule identifier**, stable and searchable, so the conversation has a subject. - **The exact resource and property.** The logical resource name in the template and the field, not "a resource in this stack". - **What was found and what is required** — retention of 1 day, at least 7 required. The engineer should be able to fix a genuine violation without asking anyone. - **Why the requirement exists**, in one line. "Recovery objective for production data stores" tells a team whether their case is an exception or a mistake. - **The owning team.** A team handle, never a personal name — people move. - **Where the rule's source lives**, so a disagreement can be read rather than argued. - **How to propose a change**, because that is the legitimate route when the rule is wrong: a change to the rule repository carrying a fixture that reproduces this template, reviewed by the owner. That last item is what converts a 5pm argument into a normal engineering workflow. The blocked team's evidence — the exact template — is also the fixture the rule owner needs. ## What the owner owes in return Being named has obligations, or naming is theatre. The owning team should have a stated response window for questions about their rule, and it should be honest: if it is next business day, say so, because a team that expects an answer in ten minutes and waits until morning learns to route around the gate instead. The owner also owes the estate a fix loop: a confirmed false positive becomes a fixture and a rule change, not a note in a chat thread. The platform team, correspondingly, owes a runbook that says how to tell the two failure classes apart, and holds the line that they diagnose the engine rather than adjudicating rules. Their most valuable contribution during a 5pm block is often a single sentence: "the engine is healthy, this denial is reproducible, here is the owning team". ## Measuring whether the routing works Two signals tell you the truth. First, what fraction of blocked builds end in a message to the central security or platform channel rather than in contact with the owning team — if it is most of them, your output is not naming owners usefully. Second, what fraction of disputed denials end in a change to the rule repository. Rules that are never changed after complaints are either perfect or feared, and one of those is far more likely than the other. ## Anti-patterns worth naming in an interview - A generic "failed policy" message with the detail only in a dashboard the blocked engineer cannot reach. - Owner recorded as an individual who has since changed teams. - The on-call being expected to relax the rule to clear the page; that is a decision about a control, made under time pressure, by the person least placed to judge it. - No stated route to change the rule, so the only paths available are argument or avoidance.

  • At 3am, how do you tell an engine failure from a correct denial?
    Breadth and reproducibility. An engine failure spreads across unrelated rules and teams, shows evaluation errors or timeouts, and often varies between runs. A working rule denies one resource on one rule id and reproduces exactly when you re-evaluate the same document. Ten minutes of that triage decides whether this is an incident or a conversation for the morning.
  • The rule's named owner has left and the team was reorganised. What now?
    Treat a stale owner as a defect in the rule, not an inconvenience. The rule repository's CI should fail or flag rules whose owner entry does not resolve to a live team, so orphans surface before a build hits them. Unclaimed rules get re-homed or demoted deliberately rather than blocking builds nobody can answer for.
  • The blocked team insists the rule is wrong about their template. What is the path?
    They open a change against the rule repository with their template reduced to a fixture that reproduces the denial. The owner reviews: either the fixture proves an over-block and the rule is corrected with that fixture kept permanently, or the requirement genuinely applies and the answer is a documented refusal with a reason. Both outcomes leave a record.

The failure output is the label on a fire door: if it does not say who to call and why the door is locked, everyone calls the building manager, forever.

saying these in an interview costs you the question

  • The platform team owns every policy failure by default
  • The message only needs to say the build failed policy
  • Whoever is on-call can relax the rule to clear the page
  • Naming an individual engineer as the rule's owner
  • Telling the blocked team to file a ticket and wait

context