Your interpretable scorecard scores 2 AUC points below a boosted model — how do you decide which to ship?
answer
- is the gap even real?
- same rows, out of time, paired interval
- price it at the operating threshold
- who has to sign the logic?
- ship one, run the other in shadow
basics
~20 sDecide from what the decision requires, not the metric gap: ship the readable scorecard when the logic must be signed off, contested or hand-edited, after confirming the 0.02 AUC difference is real out of time.
solid answer
~50 sI would attack the number before the trade. Two AUC points is 0.02 in the area under the ROC curve, the probability the model ranks a random defaulter above a random non-defaulter, so I want an interval on the paired difference measured on the same rows, and I want it out of time rather than on a random split, because a hand-built model often closes most of that gap once the population shifts. Then I would price the gap at the actual approval threshold: how many more bad accounts caught, in money. Only the residual is traded, and requirements decide it. If a committee must sign the model's logic, or a declined applicant can contest it, or a term must be hand-edited under review, the scorecard wins even at a real cost. I would also run the boosted model in shadow and mine its disagreements for terms to add.
go deeper
Know what AUC measures and that a small difference between two models may not be meaningful. Be able to say that readability can be worth a little accuracy when someone has to justify the decision.
Explain how to test whether the gap is real: the same held-out rows, a paired comparison, an out-of-time window, and a translation of the ranking difference into outcomes at the deployed threshold.
Demonstrate the production reasoning: segment-level gaps, cost of errors at the threshold, monitoring and debugging burden, and the challenger-in-shadow setup that closes the gap without giving up readability.
Own it as policy: define which decision classes require an intrinsically readable model, who makes that classification, what evidence moves a use case out of it, and how you keep the argument from being re-fought by every team that fits a stronger model.
## Step 1: interrogate the gap before you trade against it Two AUC points is 0.02 on a scale where 0.5 is coin-flipping and 1.0 is perfect. AUC — the area under the ROC curve — is the probability that the model gives a randomly chosen positive a higher score than a randomly chosen negative. Before treating 0.02 as a fact: - **Is it statistically real?** Compare the two models on the *same* held-out rows and put an interval on the paired difference, not on each model's AUC separately. Small evaluation sets routinely produce differences of this size from sampling alone. - **Does it survive out of time?** A random split flatters the flexible model, which has more capacity to exploit period-specific structure. Score both on a later time window. Gaps of this size frequently shrink or invert. - **Is it stable across segments?** A 0.02 average may be 0.05 on one product and zero on the rest, which changes the conversation from "which model" to "which model where". - **Does it move any decision?** AUC is a ranking summary over all thresholds. Your business operates at one threshold. Translate: at the approval cut-off, how many additional bad accounts are caught and how many good ones lost, in currency? A 0.02 AUC gain can be worth a great deal or nothing at all, and the finance answer is the one that belongs in the discussion. - **Has the readable model actually been given a fair chance?** Much of the gap in tabular problems comes from non-linearity and a few interactions, both of which you can add explicitly: binned or spline-shaped terms, and named interaction terms suggested by where the boosted model disagrees with the scorecard. It is a widely argued position in this literature that on tabular data with well-constructed features, the accuracy gap between a carefully built interpretable model and a strong ensemble is often small or absent — which means "we lose 2 points" should be treated as a provisional measurement of your current scorecard, not a law about model classes. ## Step 2: the requirements that override the metric If, after all that, a genuine gap remains, it is traded against requirements — and the requirements, not the modeller's preference, decide: - **Sign-off.** If a committee or a regulator has to approve the model's logic and be accountable for it, they must be able to read it. An explanation of a model is not the same as the model, and a body that signs one and deploys the other has taken on a risk it cannot see. - **Contestability.** If a person affected by the decision can challenge it, the organisation has to be able to state what the model did and defend it consistently. Post-hoc descriptions can vary between methods and runs, which is a weak position to defend from. - **Editability.** A readable model can be modified under review — a term removed, a direction constrained — with the consequence visible immediately. This matters when domain experts have hard constraints the data cannot be trusted to respect. - **Cost asymmetry of errors.** Where a wrong decision is catastrophic and rare, being able to inspect the reasoning is worth more than a small ranking improvement. - **Debuggability and lifetime cost.** When something drifts at 2am, a twelve-term scorecard is diagnosable; a large ensemble plus an explanation pipeline is a bigger surface with more failure modes and more monitoring to maintain. Total cost of ownership is a legitimate part of the accuracy trade and is usually left out. On the other side: for a high-volume, low-stakes, fully automated decision where nobody will ever be asked to justify a single case, the extra accuracy is straightforwardly worth taking. ## Step 3: refuse the binary where you can Mature answers do not pick a side; they restructure the choice: - **Interpretable in production, black box as challenger.** Ship the scorecard, run the ensemble in shadow, and use the disagreement set as a feature-discovery pipeline. Every recurring disagreement pattern is a candidate term for the scorecard, and the gap closes over successive releases. - **Segment the deployment.** Automate the easy majority with the readable model and route the hard or high-value tail to review, where the ensemble's score is one input among several rather than the decision. - **Constrain the flexible model.** Where the flexible model must be used, impose structure that makes it partly readable — additive form, bounded interactions, enforced effect directions from domain knowledge — buying back auditability at a smaller accuracy cost than full simplification. ## Step 4: own the decision as a policy As a lead, the answer is not a case-by-case instinct but a rule the organisation can apply: define which decision classes require an intrinsically readable model, who classifies a new use case, and what evidence is required to move a use case out of that category. Then the argument happens once, at the right altitude, instead of every time a team fits a stronger model. ## Interview framing Do not answer with a preference. Answer with the sequence: verify the gap, price the gap at the operating threshold, then trade the residual against sign-off, contestability, editability, error cost and maintenance — and offer the challenger-model structure that dissolves most of the dilemma.
- How would you test whether a 0.02 AUC difference between two models is real?Evaluate both on exactly the same held-out rows and form an interval around the paired difference rather than around each AUC separately, since the two scores are correlated. Repeat over several time-based folds and on a later out-of-time window. If the interval for the difference comfortably includes zero, there is no gap to trade against.
- What would change your answer towards shipping the boosted model?A high-volume automated decision with low individual stakes, no regulatory or contractual duty to state the rationale, no one who can contest an outcome, and a gap that prices out to real money at the operating threshold. In that setting the readable model's advantages are largely unused, and the accuracy is worth taking.
- How does running the ensemble as a shadow challenger help the interpretable model?Its disagreements are a map of what the scorecard is missing. Cluster the rows where the two diverge most, look for the structure driving it, and add that as an explicit term — a shaped effect or a named interaction. The gap closes across releases while the deployed model stays readable, and you keep a live measurement of what interpretability is costing.
- Is the accuracy-interpretability trade always real?Less often than it is assumed. On tabular problems with meaningful features, much of a flexible model's advantage comes from non-linearity and a few interactions that can be added explicitly to a readable model. The trade is real for high-dimensional perceptual data; on structured business data it is best treated as a measurement of your current model, not a law.
It is a building-materials choice, not a taste in architecture: the load, the inspection regime and who signs the drawings decide the material, not which one performs best in a lab test.
saying these in an interview costs you the question
- Takes the 0.02 AUC gap at face value with no interval
- Never prices the gap at the actual operating threshold
- Argues purely from personal preference for one model family
- Assumes a post-hoc explanation satisfies a sign-off requirement
- Ignores monitoring and maintenance cost in the comparison
- Treats the accuracy-interpretability trade as a fixed law