skip to content

In a refund-abuse decision layer, what changes if the rules run as a cascade before the model instead of entering it as features?

level: middleimportance: must knowfreq 74%

answer

  1. veto against evidence
  2. terminal rule short-circuits the score
  3. a feature's weight is fitted, not promised
  4. the model only sees survivors
  5. max-of keeps strength, loses calibration

basics

~20 s

A cascade makes a matching rule an absolute veto and hides the traffic it decided from the model. Rule hits as features make the rule one weighable piece of evidence the model can discount. One is policy, the other is signal.

solid answer

~50 s

In a **cascade**, ordered rules run first; a terminal rule emits the action and the model is never consulted, so that rule's outcome is guaranteed and the model's later training population contains only the requests rules did not decide. In a **blend**, the rule hit becomes an input - a binary feature the fit weighs, or its own score fused with the model's by a weighted sum or by taking the higher risk - so the rule trades off against everything else and the layer keeps one number to band. The practical split is per rule, not per system: policy and contractual rules belong in the cascade, where the fit cannot discount them; heuristic 'this smells wrong' rules belong in the blend, where their weight is fitted rather than guessed. Max-of fusion keeps a rule's strength but the maximum of two scores is not calibrated.

code

pseudocode · 10 lines
pseudocode
flags = empty_set
for rule in rules_sorted_by_precedence:
    if not rule.matches(request):
        continue
    if rule.terminal:
        return emit(rule.action, reasons = [rule.id], model_called = false)
    flags.add(rule.id)

score = model.score(request, base_features(request) + flags)
return emit(band_of(score), reasons = ["model_band"] + flags, model_called = true)

go deeper

for a junior

Recall the two shapes: rules first and short-circuiting, or rules as inputs to the score. Know that only the first one gives a rule the power to decide on its own.

for a middle

Explain the mechanics both ways - what short-circuiting does to the model's training population, how a fitted weight can shrink a rule to nothing, and why taking the maximum of two scores breaks calibration.

for a senior

Show the per-rule judgment: which conditions are commitments that must stay terminal, which heuristics are safer as weighable features, and why a terminal rule is cheap to add and expensive to retire.

for a principal

Treat it as a trade between guarantees and adaptability. Each terminal rule buys certainty in exchange for a blind spot the organisation must later pay to see into, and that bill arrives years after the rule was written.

## The two compositions, precisely A refund-abuse decision layer holds hand-written rules and a fitted risk score. There are two structurally different ways to put them together, and they are not stylistic variants of each other. **Cascade.** Rules are evaluated in a declared order. A rule marked *terminal* emits the action immediately - decline, allow, challenge, route to human case review - and the model is never called for that request. Non-terminal rules may annotate the request and pass it on. Everything that survives the rule stage is scored, and the score's band produces the action. **Blend.** Every rule is evaluated, but no rule emits an action. Its hit becomes an input: a binary feature the model was fitted with, or its own risk contribution fused with the model's output. Common fusions are a weighted sum of the two scores and *max-of* - taking whichever is higher. | | Cascade | Blend | |---|---|---| | Rule's force | Absolute veto on its own condition | One weighable piece of evidence | | Model's view of traffic | Only the requests rules did not decide | All traffic | | Output | Whichever stage emitted first | One fused score, then a band | | Tuning a rule | Edit the condition and the action | Refit, or re-tune the fusion weight | | Characteristic failure | The model never learns the blocked region | A rule the business promised gets discounted away | ## What a cascade buys, and what it costs It buys a guarantee. If the business has committed that a specific condition always declines, only a terminal rule ahead of the score delivers that; nothing downstream can dilute it. It is also trivially explainable - the winning rule *is* the reason - and it saves a model call on traffic already decided. It costs visibility. Because the model is never asked about requests a terminal rule decided, those requests never enter the model's training population either. The consequence is structural, not cosmetic: the model's behaviour on that region is an extrapolation, and nobody holds an estimate of what would happen if the rule were removed. A cascade is therefore a decision that is easy to add and genuinely hard to reverse. ## What a blend buys, and what it costs It buys weighing. A heuristic rule that was a good guess in April may be a mediocre signal in September; as a feature, the fit discounts it on the next refresh without anybody editing it. The layer also keeps **one** risk number, which is what any downstream banding or expected-loss reasoning wants. It costs the guarantee, twice over: - The fit may assign the rule's feature a small weight - especially when the feature is correlated with several others - so a request the rule flagged still comes out low-risk. A rule expressed as a feature is a suggestion. - If the fusion is **max-of**, the rule's strength survives, but the maximum of two scores is no longer a calibrated probability of abuse. Any downstream arithmetic that reads it as one is then wrong. ## Choosing per rule, not per system The useful decision is made rule by rule: 1. **Is it a commitment?** Contractual, regulatory, or promised-to-finance conditions go in the cascade as terminal rules. They are policy, and policy is not fitted. 2. **Is it a fresh, sharp observation?** A pattern spotted this week has no labelled volume to fit on. Start it in the cascade or as a high-precision flag, and revisit once it has history. 3. **Is it a long-lived heuristic?** Feed it in as a feature. Its weight is then measured rather than asserted, and it decays quietly instead of silently over-blocking. 4. **Does something downstream read the score as a probability?** Then avoid max-of fusion for that path, or band the model score and the rule contribution separately. ## The hybrid that most systems land on In practice the layer is both: a short cascade of terminal policy rules and allowlists at the front, then a model scored on everything that survives, with the non-terminal rule hits passed in as features so the score can use them. That shape keeps the small set of guarantees absolute, keeps the large set of heuristics weighable, and keeps exactly one place - the band or the winning rule - where the emitted action is produced. What it does not solve is the visibility cost: the traffic the front cascade decided is still traffic the model has never scored.

  • What makes a terminal rule hard to remove from the cascade later?
    The model has never scored the traffic that rule decided, so its behaviour there is an extrapolation and nobody holds an estimate of what removing the rule would allow through. Turning it off is therefore a bet, not a config change. Getting a real estimate needs evidence from traffic the rule did not decide, which is a deliberate, separately-owned decision.
  • When is max-of fusion the right blend, and when does it hurt?
    It suits a high-precision heuristic whose strength you want preserved while still emitting one number: the rule can always push the fused score up. It hurts wherever that number is read as a probability of abuse - the maximum of two scores is not calibrated, so an expected-loss calculation sitting on top of it is wrong.
  • Why can a rule expressed as a model feature stop having any effect?
    The fit assigns it whatever weight the data supports, and when the feature is correlated with several others its individual weight can come out near zero while the others carry the signal. The rule then still evaluates and still logs a hit, but no longer moves the emitted action - which looks like a working rule right up until someone checks.

saying these in an interview costs you the question

  • A cascade and a blend produce the same decisions, just organised differently
  • Passing a rule in as a feature guarantees the model respects it
  • A cascade leaves the model's training population unchanged
  • Any fusion of two scores is still a calibrated probability
  • Every rule in the layer must be composed the same way