skip to content

Before a new hard-block rule acts on live refund requests, how do you find out what it would have done?

level: seniorimportance: should knowfreq 44%

answer

  1. observe before it can act
  2. replay for volume, shadow for fidelity
  3. marginal coverage over existing rules
  4. read the low-band matches by hand
  5. a silent rule may be a broken input

basics

~20 s

Run it in shadow: evaluate it on live traffic and record every match and the action it would have emitted, while the decision in force still stands. Replay over stored records first for cheap volume estimates.

solid answer

~50 s

Two complementary steps. **Replay** the candidate over a recent window of stored request records to get hit volume and overlap with existing rules cheaply - but it is only as good as the fields that were logged at the time. Then run it **in shadow** on the live path: it evaluates on every request and writes what it would have emitted, while the action actually taken comes from the rules already in force. What you read before enabling it is the match rate against the expected population, how much of that another rule already blocks (its marginal coverage), and which matched accounts the model bands as low risk - that disagreement is the interesting part, and a sample of those cases should be looked at by a person. Promote with a version, an owner and a review date, never as an untracked configuration edit.

go deeper

for a junior

Recall that a new blocking rule is observed before it is allowed to act, and that shadow mode means it records what it would have done while the existing decision still stands.

for a middle

Explain the two observation methods and their limits - what a replay over stored records cannot reconstruct, what live shadow adds - and which numbers you read before enabling: match rate, marginal coverage, disagreement with the score band.

for a senior

Show the operating discipline: a sample of matched cases read by a person, versions and owners on every rule, review dates, and a dead-rule audit that distinguishes a solved pattern from an input that silently stopped arriving.

for a principal

Decide what the organisation owes a customer who is wrongly declined, since that is what the observation period buys. Set who may author a blocking rule, what evidence promotion requires, and how the rule set is kept from growing past anyone's ability to state the policy.

## Why a rule gets a rollout at all A rule looks like configuration, which makes it tempting to treat enabling one as reversible. It is not, in the way that matters: the declines it emits while it is wrong have already reached customers, and a customer who was wrongly refused a refund does not un-experience it. So a new hard-block rule in a promotion and refund abuse layer earns the same courtesy as any other change to a live decision - an observation period before it can act. ## Step one: replay over stored records Take a recent window of stored request records and evaluate the candidate against them offline. This is cheap, it runs in minutes, and it answers the first question: **how often would this have matched?** Its two limits are worth stating plainly: - **It can only use fields that were logged.** A condition on a counter whose value at decision time was never stored cannot be replayed faithfully; reconstructing it from today's aggregates gives a different number than the request actually saw. - **It sees only the traffic that arrived under the rules then in force.** Requests an existing terminal rule already blocked are in the record as blocked, so the candidate's apparent coverage is measured on a population that is already filtered. ## Step two: shadow evaluation on the live path The candidate is deployed, evaluated on every live request, and its would-be action written to the decision record - but it is excluded from resolution, so the action emitted is whatever the rules in force produce. Nothing about the live decision changes. For each shadowed request, log at minimum: the candidate's version, whether it matched, the action it would have emitted, the action actually emitted, which rule or band produced that, and the model's score band. | | Replay over stored records | Live shadow | |---|---|---| | Setup cost | Minutes | A deployment | | Input fidelity | Only what was logged | The real request-time values | | Time to a verdict | Immediate | Days to weeks of traffic | | Catches a broken input | No | Yes - a match rate of zero shows up | ## What the numbers must show before it acts 1. **Match rate against the expected population.** Ten times the expected volume usually means the condition is broader than intended - a unit mismatch, a counter over the wrong window, a missing qualifier. 2. **Marginal coverage.** Of the requests it matched, how many were already blocked by an existing rule? A candidate whose matches are a subset of another rule's adds maintenance without adding coverage. 3. **Disagreement with the score band.** Matched requests the model bands as low risk are the cases to read. They are either the pattern the model genuinely misses - which is the rule's whole justification - or the rule's false positives. 4. **A human sample.** Someone reads twenty or so matched cases end to end. No aggregate substitutes for this before a rule is allowed to decline. ## The rest of the lifecycle Shadow is only the entry gate. A rule carries, for as long as it exists: - **A version**, so a decision record can name exactly which text of the rule fired. - **An owner**, a named person or team who can say why it exists. - **A review date**, so the rule is re-justified rather than merely surviving. - **A stated intent**, in prose, which is what makes a later reviewer able to tell a wrong rule from a rule doing its job. And periodically, a **dead-rule audit**: which rules have matched nothing for weeks, and which rules match only requests another rule already resolves. Both are removal candidates - but a zero match count has two very different causes, and they must be told apart. Either the pattern really has gone, or the rule's input stopped being populated: a field renamed upstream, a counter that now returns null, a feature that silently defaults. The second case is a coverage hole wearing the costume of a dead rule, and deleting it makes the hole permanent. Confirm the rule is still evaluated and its inputs still arrive before concluding it has nothing to do.

  • What does a replay over stored records miss that live shadow catches?
    Anything the stored record cannot reconstruct: a value computed at request time and never logged, a velocity counter whose reading at that instant is gone, a payload field added after the window. Replay also only sees traffic that arrived under the rules then in force. Live shadow evaluates the real inputs on the real path, which is also what exposes an input that has quietly stopped arriving.
  • How do you retire a rule that has not matched in months?
    First confirm the zero is real - the rule is still evaluated and its inputs are still populated - because a silent rule is often a broken input rather than a solved problem. Then check whether a broader rule now covers it. Remove it through the same versioned rule-set change as any other edit, keeping the record of what it was and why it existed.

saying these in an interview costs you the question

  • A rule is configuration, so enabling it is instantly reversible
  • If the replay shows matches, the rule is ready to enable
  • Shadow evaluation only tells you the match count
  • A rule that never matches is harmless and can be left in place
  • Rules do not need versions because the condition is readable
  • A rule is safe to enable once it agrees with the model's bands