skip to content

In a pull-based estate, how do you keep the advisory repo check and the authoritative cluster gate from drifting apart?

level: principalimportance: nice to knowfreq 30%

answer

  1. two copies of one rule
  2. which direction may they differ?
  3. no false blocks before merge
  4. tighten the early copy first
  5. count the post-merge surprises

basics

~20 s

Keep one definition and generate both copies from it, decide which direction they may differ, and always tighten the early copy first. The tolerable skew is the early check passing something the cluster later denies, never the reverse.

solid answer

~50 s

Accept that you are running the same rule twice and design the relationship rather than hoping they match. One definition, owned by whoever owns the guardrail, with both copies generated from it instead of hand-maintained. Then pick the allowed direction of skew: the pre-merge copy may be weaker, because some rules need information only available at apply time, but it must never be stronger, since a false block before merge stops safe work and is the fastest way to lose trust in the early check. Sequence every change to a rule in the same order: tighten the advisory copy first and let it warn, then tighten the gate. Doing it the other way round manufactures exactly the merged-green-then-denied failure. Finally, measure it: count applies denied whose commit passed the pre-merge check, per team, per week.

go deeper

for a junior

Understand that the same rule usually exists twice, once as an early warning on the change and once as the control that actually refuses the write.

for a middle

Explain why the two copies cannot always be identical, since some rules depend on state that only exists when the object is applied.

for a senior

Show the operating discipline: generate both copies from one definition, keep the early copy a subset, and roll rule changes out to the early copy first.

for a principal

Own the direction of allowed skew and defend it: a false block before merge is worse than a late denial, and the metric you report is post-merge surprises, not gate pass rate.

## Two copies is the normal state, not a defect In a pull-based estate you have a check on the change that gives fast feedback and enforces nothing, and a gate at the write that enforces and gives feedback to nobody the author knows. Both are necessary. That means one rule exists in two places, run by two systems, changed on two cadences, and usually owned by two groups. The lead job is not to eliminate the duplication, it is to decide the relationship between the copies and hold it. ## One definition, two renderings The first move is mechanical: keep a single authoritative statement of the rule and derive both copies from it, rather than writing it twice. Hand-maintained duplicates diverge silently, and every divergence surfaces as a developer being surprised at the worst possible moment. Whoever owns the guardrail owns the definition; the pipeline copy is a build output of it, not a fork of it. This is easier for some rules than others, and it is worth saying so out loud. A rule like *every workload declares requests and limits* is the same predicate over the same fields wherever it runs. A rule that has to compare against a namespace quota, a cluster default or a value another controller fills in cannot be fully expressed against a file on disk. ## Choose the direction the skew is allowed to run This is the actual judgment call, and it has a right answer. **Acceptable:** the pre-merge check passes something the gate later denies. It costs a late failure and a corrective commit, which is annoying and recoverable. **Not acceptable:** the pre-merge check rejects something the gate would have allowed. This is a false block. It stops work that was fine, it is invisible to the person who set the rule, and it teaches an entire organisation that the early check is noise to be worked around. Once teams believe that, you have lost the cheap feedback loop and pushed all your volume onto the late path. So the pre-merge copy is deliberately a subset: everything it rejects, the gate would also reject. Where the rule needs apply-time information the early copy stays silent rather than guessing. ## Sequence every rule change the same way When a rule is tightened, tighten the advisory copy first and let it warn for a period, then tighten the gate. Authors see the new expectation while nothing is enforced, and the ones who fix it early never experience the gate at all. The reverse order is how you manufacture the worst failure available in this architecture: the gate gets stricter, the pre-merge check does not, and teams merge green and are refused afterwards by a rule they were never shown. Every one of those events costs credibility that is expensive to rebuild, and they arrive in a batch because the gate applies to everything at once. ## The number that shows drift Put one metric on the wall: **applies denied whose commit passed the pre-merge check**, broken down by team and week. Every entry is a merged-green-then-blocked event, which is precisely the experience you are trying to make rare. The metric is diagnostic, not just descriptive. A spike straight after a rule change means the copies were sequenced wrong. A steady trickle concentrated in one rule means that rule needs apply-time information and the early copy genuinely cannot cover it, so the fix is better messaging rather than more pre-merge logic. A steady trickle across many rules means the copies have drifted and the generation pipeline is not doing its job. And a number that is persistently high for one team is usually a sign the rule itself does not fit their workload, which is a conversation about the rule, not about compliance. ## Ownership and the conversation with teams Be explicit with the organisation about which control answers which question, because the ambiguity is what makes people angry. The early check answers *will this be accepted*, quickly and approximately, and it is allowed to be wrong in the permissive direction. The gate answers *is this accepted*, authoritatively and late. If a team ever gets a blocked deploy after a green check, that is a bug in your rollout, and you should treat it as one - track it, explain it, fix the sequencing - rather than telling them the gate was right. ## What a weak answer looks like Maintaining the rule twice by hand and calling it fine. Tightening the gate first because it is the one that matters. Treating a pre-merge false block as harmless because the author can always override it. Or assuming the pre-merge copy can be made exactly equivalent to the gate, which is not true for any rule that depends on state the file does not carry.

  • Which copy do you tighten first when a rule gets stricter, and why?
    The advisory pre-merge copy, in warning mode, for a defined period. Authors see the new expectation while nothing is enforced, so most changes are fixed before the gate ever sees them. Tightening the gate first means teams merge green and are refused afterwards by a rule they were never shown, in a batch, which is the most expensive way to introduce anything.
  • When is it correct for the pre-merge copy to be strictly weaker than the gate?
    When the rule needs information that only exists at apply time, such as a namespace quota, a cluster default or a value another controller supplies. Guessing at those in the repository produces false blocks, and a false block costs more trust than a late denial does, so the early copy should stay silent rather than approximate.
  • What single number would you track to show the two copies are drifting?
    Applies denied whose commit passed the pre-merge check, per team per week. A spike after a rule change means the sequencing was wrong, a trickle on one rule means that rule needs apply-time data, and a spread across many rules means the definitions have genuinely diverged.

saying these in an interview costs you the question

  • Hand-maintains the same rule in two places
  • Tightens the cluster gate before the pre-merge check
  • Treats a pre-merge false block as harmless
  • Assumes the repo check can evaluate everything the gate can
  • Tells the blocked team the gate was simply right

context