skip to content

Signing-enforcement exceptions granted 'for two sprints' are still live three years on — how do you shrink the set?

level: principalimportance: should knowfreq 34%

answer

  1. the resting state is permanent
  2. new things enforce from birth
  3. inertia should remove, not preserve
  4. four causes, not a hundred rows
  5. publish age, name an owner, fund it

basics

~20 s

Change the defaults instead of chasing rows: every new namespace enforces from birth so the set can only shrink, lapsing happens without a reminder, remaining entries are batched by root cause, and each has a named owner.

solid answer

~50 s

The list is three years old because every incentive points at growth: granting is fast, keeping costs nothing, and nobody is measured on removal. Fix the structure in order. First a ratchet — every new namespace, repository and service enforces by default — so new work cannot join the set and the population becomes finite and closing. Second, make expiry the default state: an exception lapses unless someone renews it with a reason, so inertia removes entries rather than preserving them. Third, batch by cause; a hundred rows are usually four causes, one legacy builder, one vendor, one base image, and fixing a cause closes rows in blocks. Fourth, give every survivor a named owner, a published age and funding. What you must not do is announce one cutoff date and mass-revoke: that turns slow security debt into a simultaneous availability incident and costs you the mandate to finish.

go deeper

for a junior

Understand why an exception needs an owner, a reason and an end condition written down at the moment it is granted, and that 'for two sprints' with none of those is how a temporary gap becomes permanent.

for a middle

Explain the mechanics that make the set shrink: lapse-by-default rather than reminders, enforcement on every newly created namespace, and grouping entries by root cause instead of working them one at a time.

for a senior

Show you can run the reduction — cluster the list by cause, rank clusters by rows closed per unit of work, negotiate dates with owning teams, and keep enforcement expanding while the residue is worked.

for a principal

Own the incentives and the funding. Decide who accepts residual risk and at what level, resist the estate-wide cutoff that trades security debt for an availability incident, and protect the programme's mandate while the population closes.

## Why the list grew Nothing about a three-year-old "two sprint" exception is unusual — it is the predictable output of the incentives. Granting an exception unblocks a delivery today; removing one costs engineering effort tomorrow, usually from a team that inherited the problem. The reason is written in a ticket nobody reads again, the owner has changed roles, and the exception has no end state anyone can check. **Permanent-by-default is the resting state of every exception process that does not actively work against it.** So the interesting question is not "how do we review the list", which every organisation says and few sustain. It is "what would make the list shrink even when nobody is paying attention". ## The ratchet is the highest-leverage move Make enforcement the default for everything created from now on: every new namespace, every new repository, every new service starts enforcing, with no inherited exemption. This single change converts the problem from an unbounded one into a **finite, closing population**. Before it, the set can grow faster than you close it and the programme never ends. After it, the only direction is down, and you can put a credible date on completion because the denominator stops moving. It also fixes the fairness problem. Without a ratchet, new teams learn that the exception path exists and is cheap; with one, they never experience the gap, and the exception process is visibly reserved for genuine legacy. ## Make lapsing the default, not a reminder A review date that depends on someone remembering is a reminder, and reminders lose. The property you want is that an exception **lapses unless renewed**: inertia removes it rather than preserving it. Renewal should be possible — some exceptions are legitimately long-lived — but visible and slightly expensive: a stated reason, a named accountable owner, and a record that the renewal happened. When renewal costs a little and expiry costs nothing, the list drains on its own. This is a governance default, not a paperwork exercise: the point is to reverse which outcome requires effort. ## Batch by cause, not by row A list of a hundred exceptions is almost never a hundred problems. It is typically: - one legacy build system that cannot sign, - one or two vendor images, - one base image nobody rebuilds, - a handful of workloads whose owning team dissolved. Working the list row by row is demoralising and slow; fixing the legacy builder closes forty rows at once. Cluster the list first, rank the clusters by rows-closed-per-unit-of-work, and fund the top two. The residue after that is the honest exception set, and it is usually small enough to defend individually. ## Name an owner, publish the age, fund the work Every surviving entry needs a named team, a one-line reason, and an end condition — the thing that must become true for it to go. Publish the set with the **age** of each entry visible, because age is the metric that embarrasses; a count alone hides that half the list predates the current org chart. Then face the funding question honestly. If closing an exception requires a team to rebuild a service they inherited and were not resourced to touch, the exception is an unfunded mandate and it will not close, no matter how many review meetings it attends. Either the work is funded, or the risk is formally accepted at a level senior enough to own it and re-affirmed on a schedule. A permanently accepted risk with a named accepting executive is a legitimate outcome; a forgotten row pretending to be temporary is not. ## Why the cutoff-date approach fails The tempting executive move is to announce that all outstanding exceptions expire on a date. It fails for a specific reason: the exceptions are spread across teams who mostly did not create them, and a synchronised revocation converts a slow-accumulating security debt into a **simultaneous availability incident**. The rollback that follows does more damage than the original gap, because it costs the security programme the organisational credit it needs to finish the job. Deadlines work per-cluster, negotiated with the owning team and sequenced — not estate-wide and unilateral. ## What good looks like a year later - New namespaces have never had an exception; the population is closed. - The list is a third of its old size, and the reduction came from three root-cause fixes rather than a hundred tickets. - Every remaining entry has an owner, an age, and an end condition. - The handful that will never close are written down as accepted risk, re-affirmed annually by someone who can actually accept it. The measure of success is not zero exceptions. It is that the set is closed, owned, and shrinking without anyone having to chase it.

  • Why not just announce that every outstanding exception expires at the end of the quarter?
    Because the exceptions sit with teams who mostly inherited them, so a synchronised revocation turns slow security debt into one simultaneous availability incident. The rollback that follows costs more credibility than the original gap cost risk. Deadlines work when they are negotiated per root-cause cluster with the owning team and sequenced, not imposed estate-wide on one date.
  • Some exceptions genuinely will never close. How do you represent those?
    As formally accepted risk with a named accepting owner senior enough to carry it, a written statement of what could go wrong, the compensating controls in place, and a re-affirmation on a schedule. That is an honest permanent entry. What you must not have is the same thing labelled temporary, because it hides real exposure inside a list everyone assumes is transient.
  • What single metric would you report to leadership on this?
    The count split by age, with the trend — not the raw count. Age exposes the entries that predate the current organisation, and the trend shows whether the ratchet is working. Add the share of new namespaces created enforcing, which should be one hundred percent; if it is not, the population is still open and every other number is temporary.

It is the difference between mopping a floor and closing the tap: reviewing rows removes water, but only the default-enforce ratchet stops more arriving.

saying these in an interview costs you the question

  • Sets one cutoff date and mass-revokes across teams
  • Tracks exceptions in a list reviewed 'quarterly' by security
  • Counts exceptions without naming an owner for each
  • Assumes renewal-with-justification alone shrinks the set
  • Treats a permanent exception as acceptable while labelling it temporary

context