A scorecard's age bands break monotonic WoE in one bin — how do you fix the binning?
answer
- is the dip signal or sampling noise?
- look at the default count in that band
- neighbouring bands can be combined
- merge toward the closest weight of evidence
basics
~20 sFirst decide whether the dip is real or sampling noise in a thin band. If it is noise, merge that band with the neighbour it sits closest to in weight of evidence and re-cut under a monotone constraint, accepting a small loss of information value.
solid answer
~50 sStart with the counts behind the offending band. If it holds a few hundred accounts and a dozen defaults, the reversal is almost certainly sampling noise rather than a real risk shape. The standard fix is a **merge**: combine the band with the adjacent bin whose WoE it is nearest, recompute, and repeat until the run across ordered bands is monotone. Re-cutting with a merge-based binning that enforces monotonicity (a chi-squared merge such as ChiMerge, or a shallow tree's cut points followed by monotone merging) does this systematically. Expect information value to fall slightly — that is the price of not fitting noise. Guard rails: keep each bin near 5% or more of accounts with enough defaults to estimate, and never fix a reversal by splitting into *more* bins. A nominal characteristic needs no monotone run at all.
go deeper
Know what monotone binning means: as the ordered bands move in one direction, the weight of evidence should move in one direction too, with no zig-zag along the run.
Explain the merge mechanics: which neighbour you combine the offending band with, how the counts and weight of evidence recompute afterwards, and why adding more bins always inflates information value.
Demonstrate the signal-versus-noise judgement. Quote a minimum bin size, read the default count behind the reversal, name a monotone-constrained binning approach, and say how the decision gets documented for review.
Set the standard the whole build follows: the monotonicity policy, minimum bin sizes and bin counts, how nominal characteristics are treated, who may approve a non-monotone exception, and how the tables are re-validated over time.
## Why monotonicity is a requirement, not a preference In a scorecard built for a consumer-credit decision a regulator or an internal reviewer will audit, each ordered characteristic is expected to tell a one-sentence risk story: *older applicants default less, and each older band is at least as safe as the one below it*. A weight-of-evidence run that rises, dips and rises again breaks that story in three ways: 1. **It is usually noise.** A dip in a middle band is far more often sampling variation than a real reversal in behaviour, and fitting it means the scorecard has memorised the development sample. 2. **It is fragile.** A non-monotone shape rarely reproduces on out-of-time data. When the applicant mix drifts, the dip moves or disappears, and the points table starts contradicting itself. 3. **It is hard to defend.** A points table where an applicant becomes *worse off* by getting older invites the reviewer's next question, and "the data said so" is not an answer that survives it. ## The diagnostic step first Do not reach for the merge before looking at the counts. For the offending band, read off the number of accounts and, more importantly, the number of **defaults**. Weight of evidence is a ratio of class shares, and the share of defaults is estimated from the default count alone — a band holding 12 defaults produces a WoE with a wide confidence interval, easily wide enough to swallow the dip. Compare the reversal's size to that uncertainty. If the band is large and the dip is big, you may be looking at a real effect worth investigating (a product boundary, a policy rule that only applies in that range, a data-quality break) before you smooth it away. ## The fix **Merge, do not split.** Combine the offending band with the adjacent bin whose WoE it is closest to, recompute the collapsed bin's counts and WoE, and check the run again. Repeat until the sequence is monotone. Two things follow from doing it this way: - Bins only get larger, so estimates get more stable, not less. - Information value falls slightly at every merge. That drop is expected and is the correct trade; a rise in IV from re-binning is the warning sign, not the goal. **Do it systematically.** Rather than hand-merging, re-cut with a monotone-constrained binning: a merge-based algorithm such as ChiMerge, which repeatedly combines the adjacent pair of bins whose outcome distributions differ least by a chi-squared test, or a shallow decision tree on the single characteristic to propose candidate cut points, followed by monotone merging. Either way the constraint is applied during binning, not patched afterwards. **Respect the minimum bin size.** Common practice is that no bin holds less than roughly 5% of accounts, and that every bin holds enough defaults for its WoE to mean something. A bin with **zero** defaults is a special case: its share of bads is zero, so its WoE is infinite. Merge it into a neighbour, or add a small constant such as 0.5 to both counts as a temporary measure, and treat the result as provisional. **Aim for few bins.** A continuous characteristic typically ends up with a handful — roughly five to ten. More bins raise IV on the development sample without improving out-of-time performance. ## When monotonicity does not apply Monotonicity only makes sense for a characteristic with an inherent order: age, income, months on book, number of enquiries. A nominal characteristic such as employment status has no order to be monotone in. For those you group levels with similar weights of evidence and enough volume, keep the grouping stable over time, and check that the grouping still makes business sense — not that the values ascend. There are also genuinely non-monotone continuous relationships. If the shape is large, reproducible on a held-out and out-of-time sample, and explainable, an approved exception is legitimate — but it should be a documented decision, not an accident of binning. ## The step people forget Cut points and WoE values are fitted from the target, so derive them on the **training data only** and apply that frozen table to validation, out-of-time and production. Tuning the cuts while watching validation performance leaks the outcome exactly as surely as computing the WoE on all the data does, and the monotone-looking result will not survive contact with next quarter's applications. ## What a strong answer sounds like Counts first, then merge toward the closest neighbour, then re-cut under a monotone constraint, then state the cost in information value and the guard rails (minimum bin size, few bins, training-only fitting) — and note that a nominal characteristic is exempt.
- How many bins should a continuous characteristic end up with, and how large?Usually a handful, roughly five to ten, with each bin holding at least about 5% of accounts and enough defaults for its weight of evidence to be stable. More bins always raise information value on the development sample and rarely improve out-of-time performance, so bin count is a place to be conservative rather than clever.
- One band contains no defaults at all — what happens to its weight of evidence?Its share of defaults is zero, so the log-ratio is infinite and the bin's contribution to information value is undefined. The clean fix is to merge it into the adjacent band. Adding a small constant such as 0.5 to both counts is a workable stopgap, but a zero-default bin is really telling you the bin is too small.
- Where must the cut points and WoE values be computed for this to hold up out of time?On the training sample only, then frozen and applied unchanged to validation, out-of-time and production data. Binning reads the target, so choosing cuts while watching the full dataset or the validation score leaks the outcome; the monotone table then looks excellent in development and degrades on the next quarter's applications.
- Does a nominal characteristic such as employment status need a monotone run?No — there is no natural order for the values to ascend along. You group levels with similar weights of evidence and adequate volume, keep the grouping stable across rebuilds, and sanity-check that it makes business sense. Monotonicity is a constraint for ordered characteristics only.
saying these in an interview costs you the question
- Splits into more bins to make the reversal disappear
- Keeps a zig-zag because it gives higher information value
- Never checks how many defaults sit in the offending bin
- Chooses cut points while watching validation performance
- Forces a monotone run onto a nominal characteristic