skip to content

Your shared policy library ships 60 rules as one version. How do you roll one bad rule back everywhere?

level: principalimportance: nice to knowfreq 30%

answer

  1. one bad rule, fifty-nine good ones
  2. pinned consumers will not upgrade today
  3. separate activation from rule code
  4. published versions are never mutated
  5. every flip needs owner and expiry

basics

~20 s

Versions are library-granular; a bad rule is not. Ship the off switch as configuration every enforcer reads at evaluation time so pinned consumers are covered too, then fix the rule forward in a new version. Never republish changed content under an existing version.

solid answer

~50 s

You need both levers. Fixing forward — publishing v2.3.1 with the rule corrected or removed — is the clean, auditable path, but it only reaches consumers who upgrade, and anyone pinned below it stays broken until they move, which can take days. So the library also needs a **per-rule switch that travels separately from the rule code**: a small owned configuration document, keyed on stable rule ids, that every decision point fetches when it evaluates. That reaches pinned consumers within their refresh interval, which is your real rollback time. What you never do is move the whole library back a version — that drops 59 working rules and any fixes released alongside — or republish the same version with different content. The judgment is governance: who may flip the switch, whether every flip is logged with a reason and an expiry, and how you stop the disabled list quietly becoming the ruleset.

go deeper

for a junior

Understand that a library version covers every rule in it at once, so moving the version to undo one rule also moves the other fifty-nine and any fixes released with them.

for a middle

Explain why fixing forward with a new patch version is the clean path, and why it reaches nobody who is pinned below that version until they raise their pin.

for a senior

Show that you would keep a per-rule switch separate from the rule code so it reaches pinned consumers at their refresh interval, and that you know that interval is your actual rollback time.

for a principal

Own the governance: who may disable a rule, how each flip is recorded and expired, and the argument that a fast rollback path is what makes the organisation willing to enforce strictly at all.

## Why the version is the wrong instrument A rule library is versioned as a unit but fails as individual rules. When one rule out of sixty starts blocking valid work, three instincts are available and two of them are wrong. - **Roll the library back to the previous version.** This drops the other fifty-nine rules to their older behaviour, including any fixes shipped in the same release, and it only affects consumers who actually change their pin. You have taken a broad, blunt action to solve a narrow problem, and it still does not reach anyone. - **Republish the version with the rule removed.** Now a pin that resolved one thing yesterday resolves something else today. Every reproducibility property you had is gone, and any record of the form "this change was evaluated under v2.3" no longer identifies a ruleset. Published versions are immutable; this is not negotiable. - **Fix forward and disable out of band.** This is the answer, and it is two mechanisms rather than one. ## Lever one: fix forward Publish v2.3.1 with the rule corrected, loosened or removed. This is the durable fix: the library's history stays honest, the changelog explains what changed, and consumers who upgrade get correct behaviour permanently. Its weakness is **reach**. Every pinned consumer keeps evaluating v2.3 until someone raises and merges the bump, and the consumers most likely to be badly stale are exactly the ones least likely to move quickly. If your estate is entirely floating or entirely bot-bumped within an hour, fix-forward alone may genuinely be enough — measure it before assuming. ## Lever two: a switch that is not part of the version Design the library so a rule's *code* and its *activation* travel separately. The activation lives in a small configuration document — owned by the platform team, itself versioned and reviewed — that names rule ids and whether each is active, optionally scoped to a subset of consumers. Every decision point fetches it at evaluation time rather than baking it in at install time. Then disabling a rule reaches even a consumer pinned three majors back, because that consumer still fetches the current activation document. Three design consequences follow: 1. **Rules must be individually addressable.** A stable rule id per rule is the precondition for any of this; a switch keyed on file paths or line numbers will not survive a release. 2. **The refresh interval is your rollback time.** If enforcers reread the document once an hour and cache aggressively, your fastest possible rollback is roughly an hour, and it differs per enforcer. Measure it per decision point, publish the number, and treat it as an operational commitment rather than an implementation detail. 3. **Decide what an unreachable switch document means.** The safe default is to keep the last known state rather than to fall open into "no rules" or slam shut into "deny everything". Whatever you choose, write it down before you need it. Note what this switch is *not*. It is a global off switch for a rule that is broken, not a per-resource exception for a workload that wants to be different. Conflating the two turns your incident tool into an ad-hoc permission system. ## The part that makes it a lead's call A switch that can silently disable a control is, by construction, a bypass. Everything that makes it safe is organisational rather than technical: - **A single named owner.** Who is allowed to flip it, and who is called when it is flipped wrongly. - **A record of every flip** — which rule, who, why, when — kept where an auditor and an on-call engineer can both read it, because "the gate was in force all quarter" is only true if you can also say which rules were switched off and for how long. - **An expiry on every flip.** A disable with no expiry is a permanent policy change made without review. Expiries force the fix-forward to actually happen. - **A visible disabled list**, reviewed on a cadence. A rule that has been off for a quarter with no owner is not disabled; it is deleted, and the library should say so. ## The strategic reason to build it The deeper argument is about appetite. A platform team that cannot undo a rule quickly ships rules slowly, weakly and with endless pre-negotiation, because every release carries the risk of blocking the estate with no way back. A credible, fast, per-rule rollback path is what buys the organisation the confidence to enforce strictly in the first place. When you are asked "how do you roll one rule back", the answer that lands is not just the mechanism but that sentence: rollback capability is what makes strict enforcement politically affordable. ## Before you reach for either lever One measurement decides the shape of the response: how many consumers are pinned below the fix? If the answer is "three, all bot-bumped nightly", fix forward and go home. If it is "forty, half of them months behind", the switch is the only thing that will actually reach them today, and the fix-forward is the follow-up.

  • Why not simply republish v2.3 without the offending rule?
    Because a pin must mean one thing forever. Mutating a published version makes the same build at the same pin produce different verdicts on different days, and it destroys any evidence of the form "this change was evaluated under v2.3". Roll forward to v2.3.1 instead; the history stays honest and consumers can see exactly what changed.
  • What bounds how fast a per-rule switch actually takes effect?
    The interval at which each decision point refetches the activation document, plus whatever it caches. That interval is your rollback time, and it differs per enforcer — a CI job may pick it up on the next run, a long-lived component only after its refresh. Measure it per decision point and publish the number as a commitment.
  • How do you stop the disabled list becoming the real ruleset?
    Put an expiry on every flip, a named owner on every entry, and a review of the disabled list on a fixed cadence. A rule that has been off for a quarter with nobody defending it should be deleted from the library outright, so the published ruleset and the enforced ruleset stay the same thing.

Recalling the whole catalogue because one page is wrong reaches nobody who already has a copy. A notice that readers check each morning reaches all of them.

saying these in an interview costs you the question

  • Rolls the entire library back a version to undo one rule
  • Republishes changed rule content under an already-published version
  • Assumes every consumer will upgrade within the hour
  • Adds a kill switch with no owner, log or expiry

context