skip to content

Your interpretable scorecard scores 2 AUC points below a boosted model — how do you decide which to ship?

level: principalimportance: should knowfreq 45%

answer

  1. is the gap even real?
  2. same rows, out of time, paired interval
  3. price it at the operating threshold
  4. who has to sign the logic?
  5. ship one, run the other in shadow

basics

~20 s

Decide from what the decision requires, not the metric gap: ship the readable scorecard when the logic must be signed off, contested or hand-edited, after confirming the 0.02 AUC difference is real out of time.

solid answer

~50 s

I would attack the number before the trade. Two AUC points is 0.02 in the area under the ROC curve, the probability the model ranks a random defaulter above a random non-defaulter, so I want an interval on the paired difference measured on the same rows, and I want it out of time rather than on a random split, because a hand-built model often closes most of that gap once the population shifts. Then I would price the gap at the actual approval threshold: how many more bad accounts caught, in money. Only the residual is traded, and requirements decide it. If a committee must sign the model's logic, or a declined applicant can contest it, or a term must be hand-edited under review, the scorecard wins even at a real cost. I would also run the boosted model in shadow and mine its disagreements for terms to add.

go deeper

for a junior

Know what AUC measures and that a small difference between two models may not be meaningful. Be able to say that readability can be worth a little accuracy when someone has to justify the decision.

for a middle

Explain how to test whether the gap is real: the same held-out rows, a paired comparison, an out-of-time window, and a translation of the ranking difference into outcomes at the deployed threshold.

for a senior

Demonstrate the production reasoning: segment-level gaps, cost of errors at the threshold, monitoring and debugging burden, and the challenger-in-shadow setup that closes the gap without giving up readability.

for a principal

Own it as policy: define which decision classes require an intrinsically readable model, who makes that classification, what evidence moves a use case out of it, and how you keep the argument from being re-fought by every team that fits a stronger model.

## Step 1: interrogate the gap before you trade against it Two AUC points is 0.02 on a scale where 0.5 is coin-flipping and 1.0 is perfect. AUC — the area under the ROC curve — is the probability that the model gives a randomly chosen positive a higher score than a randomly chosen negative. Before treating 0.02 as a fact: - **Is it statistically real?** Compare the two models on the *same* held-out rows and put an interval on the paired difference, not on each model's AUC separately. Small evaluation sets routinely produce differences of this size from sampling alone. - **Does it survive out of time?** A random split flatters the flexible model, which has more capacity to exploit period-specific structure. Score both on a later time window. Gaps of this size frequently shrink or invert. - **Is it stable across segments?** A 0.02 average may be 0.05 on one product and zero on the rest, which changes the conversation from "which model" to "which model where". - **Does it move any decision?** AUC is a ranking summary over all thresholds. Your business operates at one threshold. Translate: at the approval cut-off, how many additional bad accounts are caught and how many good ones lost, in currency? A 0.02 AUC gain can be worth a great deal or nothing at all, and the finance answer is the one that belongs in the discussion. - **Has the readable model actually been given a fair chance?** Much of the gap in tabular problems comes from non-linearity and a few interactions, both of which you can add explicitly: binned or spline-shaped terms, and named interaction terms suggested by where the boosted model disagrees with the scorecard. It is a widely argued position in this literature that on tabular data with well-constructed features, the accuracy gap between a carefully built interpretable model and a strong ensemble is often small or absent — which means "we lose 2 points" should be treated as a provisional measurement of your current scorecard, not a law about model classes. ## Step 2: the requirements that override the metric If, after all that, a genuine gap remains, it is traded against requirements — and the requirements, not the modeller's preference, decide: - **Sign-off.** If a committee or a regulator has to approve the model's logic and be accountable for it, they must be able to read it. An explanation of a model is not the same as the model, and a body that signs one and deploys the other has taken on a risk it cannot see. - **Contestability.** If a person affected by the decision can challenge it, the organisation has to be able to state what the model did and defend it consistently. Post-hoc descriptions can vary between methods and runs, which is a weak position to defend from. - **Editability.** A readable model can be modified under review — a term removed, a direction constrained — with the consequence visible immediately. This matters when domain experts have hard constraints the data cannot be trusted to respect. - **Cost asymmetry of errors.** Where a wrong decision is catastrophic and rare, being able to inspect the reasoning is worth more than a small ranking improvement. - **Debuggability and lifetime cost.** When something drifts at 2am, a twelve-term scorecard is diagnosable; a large ensemble plus an explanation pipeline is a bigger surface with more failure modes and more monitoring to maintain. Total cost of ownership is a legitimate part of the accuracy trade and is usually left out. On the other side: for a high-volume, low-stakes, fully automated decision where nobody will ever be asked to justify a single case, the extra accuracy is straightforwardly worth taking. ## Step 3: refuse the binary where you can Mature answers do not pick a side; they restructure the choice: - **Interpretable in production, black box as challenger.** Ship the scorecard, run the ensemble in shadow, and use the disagreement set as a feature-discovery pipeline. Every recurring disagreement pattern is a candidate term for the scorecard, and the gap closes over successive releases. - **Segment the deployment.** Automate the easy majority with the readable model and route the hard or high-value tail to review, where the ensemble's score is one input among several rather than the decision. - **Constrain the flexible model.** Where the flexible model must be used, impose structure that makes it partly readable — additive form, bounded interactions, enforced effect directions from domain knowledge — buying back auditability at a smaller accuracy cost than full simplification. ## Step 4: own the decision as a policy As a lead, the answer is not a case-by-case instinct but a rule the organisation can apply: define which decision classes require an intrinsically readable model, who classifies a new use case, and what evidence is required to move a use case out of that category. Then the argument happens once, at the right altitude, instead of every time a team fits a stronger model. ## Interview framing Do not answer with a preference. Answer with the sequence: verify the gap, price the gap at the operating threshold, then trade the residual against sign-off, contestability, editability, error cost and maintenance — and offer the challenger-model structure that dissolves most of the dilemma.

  • How would you test whether a 0.02 AUC difference between two models is real?
    Evaluate both on exactly the same held-out rows and form an interval around the paired difference rather than around each AUC separately, since the two scores are correlated. Repeat over several time-based folds and on a later out-of-time window. If the interval for the difference comfortably includes zero, there is no gap to trade against.
  • What would change your answer towards shipping the boosted model?
    A high-volume automated decision with low individual stakes, no regulatory or contractual duty to state the rationale, no one who can contest an outcome, and a gap that prices out to real money at the operating threshold. In that setting the readable model's advantages are largely unused, and the accuracy is worth taking.
  • How does running the ensemble as a shadow challenger help the interpretable model?
    Its disagreements are a map of what the scorecard is missing. Cluster the rows where the two diverge most, look for the structure driving it, and add that as an explicit term — a shaped effect or a named interaction. The gap closes across releases while the deployed model stays readable, and you keep a live measurement of what interpretability is costing.
  • Is the accuracy-interpretability trade always real?
    Less often than it is assumed. On tabular problems with meaningful features, much of a flexible model's advantage comes from non-linearity and a few interactions that can be added explicitly to a readable model. The trade is real for high-dimensional perceptual data; on structured business data it is best treated as a measurement of your current model, not a law.

It is a building-materials choice, not a taste in architecture: the load, the inspection regime and who signs the drawings decide the material, not which one performs best in a lab test.

saying these in an interview costs you the question

  • Takes the 0.02 AUC gap at face value with no interval
  • Never prices the gap at the actual operating threshold
  • Argues purely from personal preference for one model family
  • Assumes a post-hoc explanation satisfies a sign-off requirement
  • Ignores monitoring and maintenance cost in the comparison
  • Treats the accuracy-interpretability trade as a fixed law

context