Your policy engine major and your rule library both need upgrading. Which moves first, and why?
answer
- no order avoids the mismatch
- find the version pair both accept
- move one variable at a time
- name the window, with a clock
- record the gap rather than claim continuity
basics
~20 sNeither, as a big-bang. Find a rule-library version both engine majors accept, ship and bake that, then move the engine, then adopt new-engine-only constructs. Two small windows beat one large one, and you must name the window where enforcement is weakest.
solid answer
~50 sThere is no ordering that avoids a mismatch; the call is which mismatch you can survive and who absorbs it. Engine first means the new engine runs old rules, so anything the major removed breaks — loudly if it fails to load, dangerously if it merely evaluates differently. Rules first means rules written for the new engine may not run on the old one at all. The practical sequence is to get to a pair that is simultaneously valid: rewrite the rule library to the subset both majors accept, ship it and let it bake on the old engine, then upgrade the engine with rules unchanged, then adopt new constructs as a separate third change. Each step changes one variable, so a regression is attributable. Whatever you choose, name the interval in which the guardrail is degraded, decide explicitly whether you accept it or compensate, and record it — the people who rely on that control will ask whether it operated continuously.
go deeper
Understand that the engine and the rules are two separately versioned things, and that a rule library written for one major may not run on another. Know that moving both at once makes a regression hard to attribute.
Explain what breaks in each order and why one-variable-at-a-time is the safer sequence. Be able to describe the intersection step: rewriting rules to the subset both majors accept before touching the engine.
Show the plan: intersection library, bake, engine, then new constructs, with a non-prod-first progression and a decision replay before the window. Be ready to say how you would verify enforcement on the other side.
Own the part nobody volunteers — the interval where the control is degraded, who accepts that risk, what compensates for it, and how it is recorded for whoever later asks whether the control operated continuously.
## The question behind the question An interviewer asking this is not looking for `engine first` or `rules first`. They are checking whether you know that both orders create a period in which the guardrail is not doing what its owners believe it is doing, and whether you would surface that period or let it pass unmentioned. ## What each order costs **Engine first — new engine, old rules.** Everything the major removed or changed now applies to a rule library nobody has revalidated. The good outcome is a loud one: the rules fail to load and the enforcement point is obviously broken. The bad outcome is quiet: they load, they evaluate, and a subset of them now decides differently. You also inherit the enforcement point's behaviour when rules cannot load at all — which for a blocking gate can mean everything stops, and for a non-blocking one can mean nothing is checked. Neither is a state you want to meet unplanned. **Rules first — old engine, new rules.** Rules authored against the new major may use constructs the old engine does not have, so the library will not load and you are back to the same two bad states, only now you have broken a working system with a change that was supposed to be preparation. **Both at once.** Tempting, because it is one change and one window. It is the worst option, because when something decides differently you cannot attribute it: the engine and the rules moved together, and bisecting means unpicking two upgrades under time pressure. ## The sequence that usually wins Aim for a **simultaneously-valid pair** and move one variable at a time: 1. **Rewrite the rule library to the intersection** — the subset of constructs that both engine majors accept. This is a rule-only change, verified on the old engine, and it can ship on a normal schedule with no window at all. 2. **Bake it.** Let the intersection library run in production on the old engine long enough for the estate's real traffic to exercise it. Rule-side regressions surface here, cheaply, with the engine held constant. 3. **Upgrade the engine, rules unchanged.** Now any behaviour change is attributable to the engine alone, and rollback is a single, well-understood step. 4. **Adopt new-engine-only constructs afterwards**, as ordinary rule work. Step 1 does the hard part: it turns an unavoidable mismatch into a period where both sides are valid. When no such intersection exists — the major genuinely removed something with no equivalent — the fallback is to run both pairs concurrently: the old pair keeps enforcing while the new pair evaluates the same inputs and only records, and you cut over when their decisions agree over real traffic. That costs infrastructure and decision latency, so it is for the rules you cannot afford to have off, not for the whole library. Overlay the usual environment progression on whichever order you pick — non-prod, then one willing team's scope, then the estate — so the first evidence of a mismatch arrives somewhere cheap. ## Naming the window, which is the part people skip Every plan here has an interval where enforcement is degraded or absent. The principal-level behaviour is to state it in advance, in writing, with a start and an end: - **Say which rules are affected and how.** `For roughly two hours the deprecated-API deny-list is not evaluated on new workloads` is a sentence someone can make a decision about. `There may be brief disruption` is not. - **Decide: accept or compensate.** Compensations are cheap and unglamorous — freeze the change types the gate protects against, require a named approver for changes in that scope during the window, or keep the old pair enforcing in parallel. Accepting is also legitimate at 3am on a Sunday with a two-hour window; it just has to be a decision with an owner rather than a silence. - **Record it where the control's evidence lives.** Someone will eventually ask whether this control operated for the whole period. A known, bounded, approved gap that you wrote down is a manageable answer; an unrecorded gap discovered later is a much worse conversation, and a claim of continuous operation that the change log contradicts is the worst one. - **Tell the people who depend on it**, including the teams whose changes the gate normally catches — the ones most likely to ship exactly the thing it would have caught, not maliciously but because nothing stopped them. ## What you verify on the other side Before the window, replay the estate's objects through the new pair and diff decisions against the old pair. After it, watch the denial rate: a rule that fired weekly and now fires never is telling you the upgrade broke it. And keep at least one deliberately non-compliant input per rule flowing through the real enforcement path, so that if the new pair matches nothing you learn it from a failing check rather than from an audit.
- There is no rule-library version both engine majors accept. What now?Run the pairs side by side for the overlap. The old pair keeps enforcing; the new pair sees the same inputs and only records its decisions. When the two agree across real traffic for long enough, cut enforcement over and retire the old one. It costs you a second evaluation path and some latency, so scope it to the rules whose absence you cannot accept rather than to the whole library.
- Who signs off on a window in which the gate is not enforcing?The person who owns the risk the gate mitigates, not the platform team doing the work — otherwise the team performing the change is also approving its consequences. Present it as a time-boxed risk acceptance: which rules are off, for how long, what compensates, and what the fallback is if the window overruns. Record the decision alongside the control's other evidence so it is discoverable later rather than reconstructed.
- Why not do the engine and the rules in one change to keep the window short?Because attribution collapses. If a decision changes afterwards, you cannot tell whether the engine or the rules caused it without unpicking both under pressure, and rollback means reverting two upgrades at once. A shorter window bought with an unbisectable change is a bad trade: the failure mode you are optimising for is not window length, it is the regression you find on day three.
It is a bridge replacement, not a swap. You do not lift out the old span and drop in the new one over the same night; you build a shared section both ends can meet on, cross onto it, and then move each end in turn.
saying these in an interview costs you the question
- Insists there is one universally correct order
- Does both upgrades in a single change to shorten the window
- Never names the interval in which enforcement is degraded
- Assumes old rules always run on a new engine major
- Treats a non-enforcing window as invisible to anyone outside the team