skip to content

Owning and Shipping Rules

One rule enforced in three places is three copies that drift apart, published by whoever holds write access to the library. Interviewers probe how long the copies may disagree and who may change them.

on this pageshow

explore

questions

16

Your nightly encryption-at-rest sweep reports zero violations — what must you know before that means anything?

level: juniorimportance: must knowfreq 60%

answer

  1. green is a claim about a run
  2. which rules actually loaded
  3. scope enumerated, not assumed
  4. an empty ruleset also reports zero
  5. stamp the ruleset digest on results

basics

~20 s

Which ruleset produced the number and what it covered. Zero violations only means the rules that actually loaded found nothing in the resources that were actually enumerated. An empty or replaced ruleset reports exactly the same green.

solid answer

~40 s

A green sweep is a statement about a run, not about the estate. Two things have to travel with it before I believe it. First, the identity of the ruleset that was loaded — ideally a digest over the exact rule content, plus the list of rule ids that were evaluated — because a bundle that never contained the encryption rule, or one that was replaced wholesale, reports zero just as convincingly. Second, the scope that was actually enumerated: which accounts, regions and resource types, with counts, and any that could not be reached listed as errors rather than folded into the compliant total. With those, `0 violations` is a checkable claim. Without them it is unfalsifiable, and a dead sweep looks identical to a healthy one.

go deeper

for a junior

Be ready to say out loud that zero violations describes a run, not the estate, and to name the two unknowns: which rules loaded and which resources were looked at.

for a middle

Explain the mechanics that make zero ambiguous — a fetched ruleset that can be stale, empty or substituted, and an enumeration that can silently shrink. Describe the extra fields that resolve it.

for a senior

Show how you operate a control whose healthy state is silence: alert on the disappearance of expected output, watch per-rule evaluation counts, and keep errors out of the compliant total.

for a principal

Own the framing your organisation reports upward. Insist that assurance statements name the ruleset and the scope, so leadership is never handed a colour that cannot be traced to what was checked.

## What a green sweep actually asserts A periodic compliance sweep enumerates live resources, evaluates a loaded set of rules against each one, and emits findings. When it prints `0 violations`, the honest reading is: *the rules that were loaded found nothing wrong in the resources that were enumerated*. Both halves are variables. Neither of them is visible in the number zero, which is why the number on its own carries almost no information. ## The two questions hiding behind the number **Which rules ran?** The enforcer does not carry its rules in its binary; it fetches a ruleset from somewhere at start-up or on a refresh interval. That fetch can return a stale copy, a copy that never contained the encryption-at-rest rule, a partially written bundle, or — in the interesting case — a complete ruleset that somebody else published. Every one of those outcomes yields zero violations. The pathological version is an empty ruleset: it evaluates nothing, finds nothing, and produces a flawless report. In this domain the failure mode that most resembles success is *no policy at all*. **What was enumerated?** A sweep over managed volumes and databases has to list accounts or subscriptions, then regions, then resource types, and page through each. A credential that quietly lost read access to one account, a region nobody added to the config, a resource type the enumerator does not know about, or a pagination bug that stops after the first page — each shrinks the denominator without touching the numerator. Zero findings across three resources and zero findings across thirty thousand render identically on the dashboard. ## Why this bites harder on a sweep than on a blocking gate A gate that blocks changes generates friction: people argue with it, so its liveness is continuously, if grumpily, confirmed. A periodic detective sweep has silence as its expected output. Nobody investigates a quiet control. That asymmetry is what makes a sweep worth attacking and worth instrumenting: you have to be able to distinguish *nothing is wrong* from *nothing was checked*. ## What the report has to carry - **Ruleset identity** — a digest computed over the rule content that was actually loaded, not the version string someone hoped was deployed. - **Rule inventory** — the list of rule ids evaluated in this run, and for each, how many resources it was evaluated against. A rule that evaluated zero resources is a completely different event from a rule that evaluated four thousand and found none non-compliant, yet both report zero violations. - **Scope enumerated** — accounts, regions and resource types with counts, so the denominator is on the page. - **Errors as a separate category** — unreachable accounts, expired credentials, throttled API calls. An enumeration failure is not compliance. If it is folded into the green count, losing access to an account looks exactly like cleaning it up. ## How to read it once you have it The useful signal is usually comparative rather than absolute. Per-rule evaluation counts that drop from thousands to zero, a rule id that vanishes from the inventory, a scope count that halves overnight — these are loud even while violations stay reassuringly at zero. Alerting on the *absence* of expected output is the habit that separates people who have operated a detective control from people who have only configured one. ## Language that keeps you honest The result of a sweep is not `compliant`. It is `no evidence of non-compliance was produced by ruleset <digest> over scope <S>`. That sentence is longer and less satisfying, and it is the one you can defend when somebody asks what the control actually established. It also makes the gaps obvious: change the ruleset and the claim changes; shrink the scope and the claim shrinks; lose the ruleset identity and there is no claim left at all, only a colour.

  • The sweep skipped a whole account because its credential expired. Should that come back as zero violations?
    No. An enumeration failure is not compliance. Unreachable scope belongs in its own error category, reported alongside the results and alerted on. If it is absorbed into the compliant total, losing access to an account is indistinguishable from securing it, and estates silently drop out of coverage over months.
  • How do you make 'the rule was missing' visibly different from 'nothing was violating'?
    Report per-rule evaluation counts, not just violation counts. A run that lists the encryption rule and shows it evaluated four thousand volumes tells a different story from a run where that rule id is absent from the inventory, or present with zero evaluations. Both currently show zero violations.
  • What is the smallest addition to a report that improves it most?
    The digest of the loaded ruleset. It costs one field, it makes every run traceable to exact rule text, and it turns the most dangerous failure — running rules nobody approved, or none at all — from invisible into a comparison anyone can perform later.

A smoke alarm that never beeps is either a quiet house or a dead battery. The only difference is the test button.

saying these in an interview costs you the question

  • Treats a green run as proof the estate is compliant
  • Assumes the enforcer swept everything it was supposed to
  • Never checks which rules the run actually loaded
  • Counts an unreachable account as compliant
  • Ignores that an empty ruleset also reports zero

context

open as a page

Why is a policy rule repository reviewed, tested and released like application code?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A rule is production code: one bad rule blocks every team's builds at once. Review, tests and CI catch it before it reaches a gate, and give each change an author, a reviewer and a history.

open as a page

Why must a policy rule carry a stable id and an owner, not just a title?

level: juniorimportance: must knowfreq 58%

basics

~20 s

A rule's id is the handle everything outside the rule keys on: waivers, suppressions, control maps, dashboards. Titles get reworded, so they cannot be that handle. The owner tells a blocked engineer who to ask.

open as a page

A policy rule you own blocks another team's deploy at 5pm: who answers, and what must the failure say?

level: seniorimportance: must knowfreq 58%

basics

~20 s

Separate an engine failure from a working rule. The platform on-call owns the engine; the named rule owner owns the decision. The failure output must name the rule, the resource and property, the owning team and how to propose a change — otherwise every block routes to the platform team.

open as a page

Your CI check at rule library v2.3 blocks a manifest that a developer's pinned v1.9 pre-commit hook passed. How do you respond?

level: seniorimportance: must knowfreq 62%

basics

~20 s

Confirm it is version drift rather than a false positive by running both versions on the same manifest and reading the changelog between them. Then unblock by fixing the manifest, and close the window by moving the lagging hooks forward. The gate stays current.

open as a page

One resource-limits rule runs in a pre-commit hook, CI and admission — why do the three copies drift?

level: juniorimportance: should knowfreq 48%

basics

~20 s

Each decision point installs its own copy of the rule library and updates on its own schedule. Publishing a new version does not change what is already installed, so the three run different rule versions until each one is upgraded.

open as a page

How does a policy enforcer establish that it loaded the intended ruleset and not a substituted one?

level: middleimportance: should knowfreq 46%

basics

~20 s

By computing a digest over the whole rule set it loaded and comparing it against an expected value obtained through a different path than the ruleset itself. A mismatch must fail the run loudly, and the digest should be stamped on every result.

open as a page

What must a policy rule's test suite assert beyond denying the obviously bad input?

level: middleimportance: should knowfreq 62%

basics

~20 s

It must assert what the rule allows, not only what it denies — otherwise a rule that blocks everything passes. It also needs the awkward cases: the property missing entirely, and a value that is only resolved at deploy time.

open as a page

When should a family of near-copy policy rules become one parameterised rule?

level: middleimportance: should knowfreq 47%

basics

~20 s

Collapse near-copies when the logic is identical and only a data value differs - the list of allowed licences, say. Keep them separate when the denial message or the remediation genuinely differ, because that is different logic wearing the same shape.

open as a page

In a shared policy rule library, what makes a change a major SemVer bump rather than a minor one?

level: middleimportance: should knowfreq 42%

basics

~20 s

Anything that turns a previously passing input into a failure: a new blocking rule, a tightened condition, or a stricter default. Changes that cannot newly fail anything — a fix that only loosens a check, a rule shipped switched off — are minor or patch.

open as a page

Your whole policy ruleset was swapped for one that allows everything and sweeps stay green — how do you detect it?

level: seniorimportance: should knowfreq 50%

basics

~20 s

Not from the results — they look perfect. Detect it out of band: alert on every publish to the rule distribution point, reconcile the enforcer's reported ruleset digest against what was actually published, and keep a deliberately non-compliant canary whose finding must appear in every sweep.

open as a page

You renamed a policy rule's id last sprint and nothing failed - what silently broke?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Everything keyed on the old id stopped matching, silently. Waivers now exempt nothing, suppressions in other repositories are dead text, control-map rows name a check nobody emits, and the rule's violation trend fell to zero - which reads like success.

open as a page

An auditor asks which ruleset was live for each of last quarter's sweeps — what do you show?

level: seniorimportance: nice to knowfreq 35%

basics

~20 s

Each sweep record must carry the digest of the ruleset it loaded, and every published ruleset must still be retrievable by that digest so its rule text can be read back. A digest you can no longer resolve to content proves nothing.

open as a page

A central security team wrote every policy rule and cannot maintain them. How do you re-home ownership?

level: principalimportance: nice to knowfreq 36%

basics

~20 s

Split authorship, ownership and operation. Move each rule to the domain team closest to the resource, transferred with its fixtures, its reason and its current denial rate, and with real authority to change it. Keep contractual rules central. Give unowned rules an expiry, not indefinite life.

open as a page

Who approves widening a shared licence-policy rule's allowed list, and what does that silently re-open?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Widening a shared parameter is a scope change, not an exception: it applies to every artifact reading that set, retroactively and silently. Approval belongs with the set's accountable owner plus the rule owner, on a computed delta of flipped decisions.

open as a page

Your shared policy library ships 60 rules as one version. How do you roll one bad rule back everywhere?

level: principalimportance: nice to knowfreq 30%

basics

~20 s

Versions are library-granular; a bad rule is not. Ship the off switch as configuration every enforcer reads at evaluation time so pinned consumers are covered too, then fix the rule forward in a new version. Never republish changed content under an existing version.

open as a page