skip to content

How do you rank adverse-action reason codes when correlated features split the attribution?

level: seniorimportance: nice to knowfreq 30%

answer

  1. the decision is stable, the story is not
  2. three views of one behaviour compete
  3. aggregate before you rank
  4. codes are families, not raw inputs
  5. fixed priority order breaks ties

basics

~20 s

Rank at the level of business reason codes, not raw features: map correlated inputs to one code and sum their contributions inside it. Add a deterministic tie-break so equal contributions produce the same order every run.

solid answer

~50 s

The failure mode is that three correlated inputs — balance, utilization ratio and available headroom — each carry about a third of the decline's contribution, so none individually beats an unrelated input with a concentrated contribution, and the top-four list churns between scoring runs while the decision never moves. The fix sits in the taxonomy layer, not the attribution maths: define reason codes as business-meaningful families up front, map every input to exactly one code, sum contributions within a code, then rank codes. Break exact ties with a fixed documented priority order rather than whatever order the scorer emits, and round to a tolerance so numerical noise cannot reorder anything. Before release, rescore a frozen set of past declines under the candidate model and measure how often the issued code set changes; store the codes actually sent with each decision so the notice is reproducible.

go deeper

for a junior

Recall the core mechanic: correlated inputs each get a slice of the same effect, so no single one looks important. Group them before ranking.

for a middle

Explain the many-to-one mapping from inputs to reason codes, summing contributions within a code, and why deterministic tie-breaking and rounding matter for reproducing a notice.

for a senior

Show you would gate a release on measured explanation churn over a frozen sample of past declines, persist issued codes with the decision, and audit codes that fire on nearly every notice or never at all.

for a principal

Own the taxonomy as a governed artefact: who defines the codes and their wording, how a new input gets a mapping before it ships, and what churn budget you accept against the cost of freezing the model.

## The symptom A lender's decline explanations are generated by attributing each decision to its inputs and reporting the top four. In testing, two things show up. First, an application declined mainly because of credit-line usage gets no usage reason at all — the top four are dominated by unrelated inputs. Second, rescoring the same applications after a small model refresh permutes the reported reasons, even though the approve/decline outcome for nearly every application is unchanged. Both are the same underlying phenomenon: correlated inputs share credit for the same underlying story, and a per-input ranking treats them as competitors. ## Why the split happens A credit file naturally carries several near-redundant views of one fact. Current revolving balance, utilization as a percentage of limit, and available headroom in dollars are three encodings of the same behaviour. Any attribution method that assigns credit per input must divide the effect between them somehow; none of them can receive the full effect, because holding the other two fixed leaves little for the third to explain. So the family carries, say, 30% of the push toward decline, but appears in the ranking as three entries of about 10% each — below a single delinquency input that carries 18% on its own. The list is arithmetically defensible and communicatively wrong: the biggest thing about this applicant does not appear in the explanation they receive. The churn follows from the same structure. When three quantities sit within noise of one another, an insignificant change — a retrain, a slightly different sampling of the training data, a rounding difference — reorders them. Ranking is discontinuous where the underlying quantities are nearly equal. ## Why churn matters even though the decision does not change It is tempting to shrug: the applicant was declined either way. Three reasons not to. - **Reproducibility.** If a notice is challenged, you must be able to show why those four reasons and not four others. `The ordering is unstable` is not a defence. - **Consistency across applicants.** Two applicants with materially identical files, scored a week apart, should not receive different explanations. That is the kind of inconsistency an examiner probes. - **Trust and remediation.** An applicant told to fix the wrong thing does the wrong thing and reapplies to the same decline. ## The remedy: rank codes, not inputs The durable fix is a reason-code taxonomy defined before modelling and owned jointly by the business and compliance. 1. **Define the codes** as customer-meaningful families — credit-line usage, recent delinquency, length of credit history, recent inquiries, income relative to obligations — each with fixed, reviewed wording. 2. **Map every input to exactly one code.** The mapping is a reviewed artefact, versioned with the model. New input, new mapping entry, or it does not ship. 3. **Aggregate then rank.** Sum the per-input contributions within each code, and rank the codes. The usage family now carries its full 30% and appears first, which is what the applicant actually needs to hear. 4. **Make the ranking deterministic.** Round contributions to a stated tolerance and break remaining ties by a fixed documented code priority, so the same application always yields the same list. On points-based scorecards the classical version of this is ranking by points below the maximum achievable for each characteristic — how much the applicant lost relative to the best attainable value — with the same aggregation over related characteristics. The principle carries over unchanged to a more flexible model: rank by shortfall attributable to a code, not by an input's raw value. ## Testing stability before release Treat explanation stability as a release check with a number attached, alongside the usual predictive metrics. - Rescore a frozen sample of past declines under the candidate model and the incumbent. - Measure the fraction whose **top code changes**, and the fraction whose **top-four set** changes (order within the set matters less than membership). - Investigate any code that fires on nearly every decline: a code that appears on 90% of notices is too coarse to be informative and probably needs splitting. - Investigate codes that never fire: dead entries in the taxonomy usually mean a mapping mistake. Set a tolerance, and make a large jump in top-code churn something a human reviews before release rather than something applicants discover. ## What not to do - **Do not drop correlated inputs from the model** to tidy the explanation. That trades predictive quality for reporting convenience, and the surviving input's contribution becomes a proxy for the family anyway. - **Do not rank by raw feature value.** A large balance is not the same as a balance that drove this decision; that is a magnitude, not a contribution. - **Do not let the order be whatever the scoring code returns.** Implicit ordering is unreviewable and changes with unrelated refactors. - **Do not report every input with a non-zero contribution.** A long list is legally weaker and practically useless. ## What the interviewer is checking That you notice the difference between the decision being stable and the explanation being stable, that you locate the fix in the reason-code taxonomy rather than in the attribution maths, and that you treat explanation stability as something you measure and gate on rather than something you hope for.

  • The decision is unchanged between model versions but the reported reasons flip. Is that acceptable?
    No. Two near-identical applicants scored a week apart should receive the same explanation, and a challenged notice has to be defensible. Aggregate contributions into codes so near-ties inside one family stop competing, round to a tolerance, break remaining ties by a documented priority order, and gate the release on measured top-code churn over a frozen decline sample.
  • How would you measure reason-code stability before releasing a retrained model?
    Rescore a frozen sample of past declines under both the incumbent and the candidate, then report the share whose top code changes and the share whose top-four code set changes. Set a tolerance and require human review above it. Also check for codes firing on almost every decline, or never firing — both usually indicate a taxonomy or mapping defect.
  • Would you drop redundant correlated inputs so each reason code has a single feature behind it?
    No. That sacrifices predictive quality for reporting tidiness, and the retained input inherits the family's contribution anyway, so the explanation is no more honest. Keep the inputs and fix the reporting layer: many-to-one mapping from inputs to codes, contributions summed within a code, ranking done on codes.

saying these in an interview costs you the question

  • Ranks reason codes by raw feature magnitude
  • Assumes correlated inputs can be ranked independently
  • Leaves tie-breaking to whatever order the scorer emits
  • Drops predictive features to simplify the explanation
  • Treats a stable decision as proof the explanation is fine
  • Reports every input with non-zero contribution
  • Never measures explanation churn across model versions

context