In credit scoring, what do a KS of 38 and a Gini of 0.52 mean?
answer
- Both summarise ordering, not accuracy of probabilities
- One is a single widest gap
- Bads captured minus goods captured
- The other is a whole-curve summary
- Twice the area measure, minus one
basics
~20 sKS 38 means that at its best cut-off the scorecard separates 38 percentage points more of the bads than of the goods. Gini 0.52 is a whole-curve summary equal to two times AUC minus one, so AUC is 0.76.
solid answer
~50 sBoth are rank-ordering summaries for a binary scorecard, quoted in different dialects. The Kolmogorov-Smirnov statistic is the largest gap between the cumulative share of bads captured and the cumulative share of goods captured as you walk the score from worst to best — the widest vertical distance between TPR and FPR. Credit teams report it times 100, so KS 38 means 0.38, and it also tells you *where* the separation is widest, which is useful context for a policy cut. Gini instead compresses the entire curve into one number: `Gini = 2 * AUC - 1`, so 0.52 back-translates to an AUC of 0.76. Zero means no discrimination and 1 means perfect ordering. The two can disagree: KS is a single point and can be flattered by strong separation in one score band, while Gini averages across all of them. Neither says anything about whether the predicted default probabilities are numerically right.
go deeper
Recall that both numbers describe how well a score separates bads from goods, that higher is better, and that Gini equals two times AUC minus one. Being able to convert 0.52 to 0.76 is the expected minimum.
Explain the mechanics: KS as the maximum gap between the cumulative share of bads and of goods captured, Gini as a whole-curve summary, and why both are unaffected by any monotone rescaling of the score.
Demonstrate judgment about reporting — out-of-time validation, sample size and bad counts behind the number, tracking decay across vintages, and refusing to compare across products or bad definitions.
Own which statistic governs decisions: whether a single-point separation measure or a whole-curve one matches how the score is used, and what movement in it triggers redevelopment rather than a memo.
## Two numbers, one property A consumer-credit application scorecard is usually summarised to a risk committee in a single line such as "KS 38, Gini 0.52". Both numbers measure the same underlying property — how well the score *rank-orders* bads (defaults) above goods — and neither measures whether the score's probabilities are numerically accurate. Understanding what each one is, and where each one lies to you, is standard interview ground for any modelling role in lending, insurance or fraud. ## The KS statistic Sort the population by score. As you sweep a cut-off from one end to the other, track two cumulative quantities: the share of all bads that fall on the reject side (the true-positive rate, if "bad" is the positive class) and the share of all goods that fall there (the false-positive rate). Both start at 0 and end at 1. The Kolmogorov-Smirnov statistic is the maximum gap between them: `KS = max over cut-offs of (share of bads captured - share of goods captured)` Equivalently, it is the largest vertical distance between the two cumulative distributions of the score, one computed on bads and one on goods. Because both curves start and end together, the gap is zero at both ends and peaks somewhere in the middle. A KS of 0 means the two populations score identically; a KS of 1 means some cut-off perfectly splits them. Credit convention multiplies by 100 and drops the decimal point, so "KS 38" is 0.38. Rough industry folklore puts a usable application scorecard somewhere in the 25-45 band, though the achievable range depends heavily on the portfolio and on how "bad" is defined — comparing KS across products or definitions is not meaningful. KS carries one extra piece of information the other summaries do not: the score at which the maximum occurs. That location is often reported alongside the value, because a model whose separation peaks deep in the subprime tail behaves differently in policy from one that peaks near the middle of the book. ## The Gini coefficient Gini for a model is defined off the whole discrimination curve rather than a single point, and it has an exact relationship to the area under the ROC curve: `Gini = 2 * AUC - 1` So AUC 0.5 (no discrimination) maps to Gini 0, and AUC 1.0 maps to Gini 1. The quoted 0.52 back-translates as `AUC = (Gini + 1) / 2 = 0.76`. Learn both directions of that arithmetic; being asked to convert on the spot is common, because model documentation, vendors and regulators do not agree on which one to print. In credit-risk documentation the same quantity also appears as the **accuracy ratio**, computed as the area between the cumulative accuracy profile and the diagonal relative to the perfect model's — it equals `2 * AUC - 1` as well. A caution worth stating explicitly: this Gini has nothing to do with **Gini impurity**, the `1 - sum(p_i^2)` splitting criterion inside a decision tree, and it is only loosely analogous to the economists' Gini of income inequality. Three unrelated meanings share one surname, and mixing them up in an interview is memorable for the wrong reason. ## Where they disagree, and why that matters Because KS is a maximum over cut-offs and Gini is an integral across all of them, two scorecards can trade places depending on which is quoted. A model with a sharp separation in one narrow score band and mediocre ordering elsewhere can post a strong KS and a middling Gini. A model that orders consistently well everywhere but never opens a dramatic single-point gap posts the reverse. Which you prefer depends on use: if policy is a single hard cut-off, the peak matters and KS is informative; if the score drives risk-based pricing across the whole book, the whole-curve summary matters more. Both share the same blind spots. Both are invariant to any monotone transform of the score, so scaling a model into a 300-850 band changes neither. Both are indifferent to calibration: a scorecard whose probabilities are systematically double the truth has identical KS and Gini to a corrected one. And both are point estimates on a sample. On a small validation set, a Gini difference of 0.02 between two candidate models is usually noise, and committees that chase such differences are chasing sampling variation. ## Reporting them honestly Quote them on a holdout, ideally out-of-time as well as out-of-sample, since credit performance decays as the population and economy shift. Show the sample size and the number of bads — a Gini computed on 40 defaults is barely an estimate. Track them over time on the live book, because a falling KS on recent vintages is often the first sign that a scorecard needs redevelopment, well before losses show it.
- Convert a Gini of 0.52 to AUC, and say what AUC of 0.5 implies.AUC = (Gini + 1) / 2 = 0.76. An AUC of 0.5 corresponds to a Gini of 0, meaning the score orders bads and goods no better than shuffling the file: for a randomly chosen bad and a randomly chosen good, the bad is riskier-scored only half the time. That is a scorecard with no discriminatory value.
- Two scorecards have the same Gini but one has a much higher KS. What does that tell you?The higher-KS model concentrates its separation in one score region, opening a wide single-point gap there while ordering no better overall. If policy uses one hard cut-off near that region, it is the better choice; if the score drives pricing or limits across the whole book, the equal Gini says there is nothing to choose between them elsewhere.
- The scorecard's Gini is unchanged but the business says the default probabilities are all too low. What is wrong?Discrimination and calibration are separate properties. Gini depends only on the ordering of scores, so a systematic bias in the level leaves it untouched. The ordering is fine and the mapping from score to probability is not, which is a rescaling problem rather than a reason to rebuild the model.
- Why is comparing KS across two different lending products misleading?KS depends on the portfolio's mix and on the definition of bad — how many days past due, over what window. A near-prime card book with a 90-days-past-due definition and a secured loan book with a different one produce numbers that are not on the same scale. Comparisons are only meaningful within one product and one bad definition.
saying these in an interview costs you the question
- Confuses the model Gini with Gini impurity in trees
- Says Gini equals AUC, or AUC minus 0.5
- Thinks KS is an average gap rather than the maximum
- Claims a high KS means probabilities are accurate
- Compares KS across products with different bad definitions
- Treats a 0.01 Gini difference on a small sample as real