skip to content

Why do credit scorecards replace binned features with their weight of evidence?

level: middleimportance: must knowfreq 68%

answer

  1. one number per bin, not a count
  2. compares two distributions, goods and bads
  3. a logarithm of two class shares
  4. bin log-odds minus population log-odds

basics

~20 s

Weight of evidence replaces a bin with ln(share of non-defaults / share of defaults), which is the bin's log-odds offset from the population. That makes every feature monotone, on one common scale, linear in log-odds, auditable, and lets missing values form their own bin.

solid answer

~50 s

For each bin you compute `WoE = ln(bin's share of all non-defaults / bin's share of all defaults)`. By Bayes, that equals the bin's log-odds of being good minus the population log-odds, so the transform hands a logistic model a feature that is already **linear in log-odds** and expressed in the same units as every other characteristic. Practical wins follow: outliers are absorbed by the extreme bins, missing values become an ordinary bin with their own WoE instead of needing imputation, rare levels are merged into a bin big enough to estimate, and the resulting points table is readable by a reviewer. The price is real: within-bin variation is thrown away, and each WoE is an estimate from the development sample, so thin bins give noisy values. Fix one sign convention for the whole build, because `ln(bad/good)` simply flips every sign.

code

python · 18 lines
python
import math

# (band, non-defaults, defaults) for an applicant-age characteristic
bins = [("18-25", 1200, 180), ("26-35", 3000, 240),
        ("36-50", 3600, 150), ("51+", 2200, 60)]

total_good = sum(g for _, g, _ in bins)
total_bad = sum(b for _, _, b in bins)

iv = 0.0
for band, good, bad in bins:
    p_good = good / total_good      # this bin's share of all non-defaults
    p_bad = bad / total_bad         # this bin's share of all defaults
    woe = math.log(p_good / p_bad)  # positive = safer than the portfolio
    iv += (p_good - p_bad) * woe
    print(band, "WoE=", round(woe, 3))

print("IV=", round(iv, 3))

go deeper

for a junior

Be ready to state the formula and say which direction is safer. Under the goods-over-bads convention, a positive weight of evidence means the bin holds proportionally fewer defaults than the portfolio average.

for a middle

Explain why the transform helps. Weight of evidence is the bin's log-odds shifted by the population log-odds, so a logistic model receives a feature that is already linear in log-odds and on the same scale as every other characteristic.

for a senior

Show you know the cost and the mechanics. Binning discards within-bin variation, WoE values are sample estimates, and the whole table is a fitted object that must be derived on training data and applied unchanged everywhere else.

for a principal

Own the trade-off. A weight-of-evidence scorecard buys monotone, auditable, stable behaviour at some loss of raw accuracy; be able to say for which portfolios and which decisions that price is worth paying, and where you would ship a stronger model instead.

## What weight of evidence is Weight of evidence (WoE) is a supervised transform of a *bin* against a *binary target*. Split a characteristic (applicant age, employment status, bureau balance) into bins, and for each bin count the two outcome classes. In credit language the classes are **goods** (non-defaults, the non-event) and **bads** (defaults, the event). For bin `i`: ``` p_good_i = goods_i / total_goods (the bin's share of ALL goods) p_bad_i = bads_i / total_bads (the bin's share of ALL bads) WoE_i = ln(p_good_i / p_bad_i) ``` Note what is in the denominator: shares of the two *class totals*, not the bin's own size. A bin holding a larger slice of the goods than of the bads gets a positive WoE and is safer than average; a bin over-represented among defaults gets a negative WoE. WoE is exactly 0 when the bin looks like the portfolio as a whole. The opposite convention, `ln(p_bad / p_good)`, is equally common and flips every sign — higher then means riskier. Neither is wrong; mixing them inside one build is. ## Why it linearises the problem By Bayes' rule, `odds(good | bin i) = (p_good_i / p_bad_i) * odds(good)`, so ``` ln odds(good | bin i) = ln odds(good) + WoE_i ``` The WoE of a bin *is* that bin's log-odds, shifted by a constant that is the same for every bin. A logistic model works on the log-odds scale, so feeding it WoE means the relationship between feature and target is already a straight line — no polynomial terms, no splines, no interaction guessing to capture a curved risk shape. The model's job shrinks to weighting characteristics against each other. ## The other practical benefits - **One scale for everything.** Age in years, income in currency and a nominal employment status all arrive as WoE numbers in roughly the same range, so a points table can compare them. - **Missing values are information.** A blank bureau field is often predictive (thin file, new to country). Binning gives it its own bin and its own WoE, so nothing needs imputing and the missingness signal survives. - **Outliers are capped.** A single applicant with an absurd declared income falls into the top bin and gets that bin's WoE; it cannot drag a coefficient around. - **Rare levels are absorbed.** Levels too small to estimate are merged into a bin large enough to give a stable WoE. - **Auditability.** The finished object is a table of bins and points that a reviewer, a credit officer or a regulator can read line by line, and monotone bins let you state the risk story in one sentence per characteristic. ## What it costs Binning is lossy by construction: every applicant inside a bin gets the same value, so within-bin ordering disappears. On raw predictive power a gradient-boosted model on the unbinned features usually wins. The scorecard is chosen when stability, monotonicity and defensibility are worth more than the last few points of separation. The WoE values are also **estimates**. A bin with a handful of defaults produces a WoE with wide sampling error, and a bin with zero defaults produces an infinite one (division by a zero share) — that bin must be merged into a neighbour, or the counts nudged by a small constant such as 0.5 as a stopgap. Because the transform reads the target, the bin cut points and the WoE values are **fitted quantities**. Derive them on the training data only and apply them unchanged to validation, out-of-time and production data; choosing cuts by looking at the whole dataset leaks the outcome and flatters the development results. ## The related summary statistic The WoE table for one characteristic collapses to a single number, the information value, `IV = sum over bins of (p_good_i - p_bad_i) * WoE_i`, which is used to screen characteristics before modelling. Every term has matching signs in its two factors, so IV is never negative. ## What interviewers listen for A strong answer states the formula with the right denominators, says *why* it is a log-odds quantity, names one or two of the practical wins (missing bin, common scale), and volunteers the cost — lost within-bin detail and estimates that are only as good as the bin counts behind them.

  • How does a WoE-binned characteristic handle missing values?
    Missing becomes its own bin with its own WoE, so no imputation is required and the fact of missingness stays in the model — often a genuinely predictive signal, such as a thin credit file. If the missing bin is too small to estimate, merge it into the bin whose WoE it sits closest to, and say in the documentation that you did.
  • Does WoE binning throw information away?
    Yes. Everyone inside a bin receives the same value, so within-bin ordering is discarded, and a model on the raw features will usually separate a little better. The trade is deliberate: you buy a monotone, stable, reviewable relationship and immunity to outliers, and you pay for it in raw discrimination.
  • Both ln(goods/bads) and ln(bads/goods) are used — does the choice matter?
    It flips the sign of every WoE and therefore the direction of the score; it changes nothing about the strength of the relationship, and information value comes out identical either way. Pick one convention, write it down, and keep it across every characteristic and every rebuild, or the points table will contradict itself.
  • Where in the pipeline must the bin cut points and WoE values be computed?
    On the training data only. Binning reads the target, so it is a fitted step: derive cuts and WoE on the development sample, then apply that fixed table to validation, out-of-time and production data. Choosing cuts from the full dataset leaks the outcome and inflates development performance.

It is like re-labelling every price tag in a shop as a percentage off the average price: different goods become directly comparable, and you lose the exact price.

saying these in an interview costs you the question

  • Says weight of evidence is the bin's default rate
  • Divides by the bin's own size rather than class totals
  • Confuses weight of evidence with a frequency or count encoding
  • Computes the bins and WoE on the full dataset
  • Claims binning loses nothing because bins are chosen well

context