skip to content

Why does a k-NN majority vote with k=15 almost never predict a class that is 4% of the data?

level: seniorimportance: nice to knowfreq 32%

answer

  1. count what a majority actually requires
  2. the vote mirrors local proportions
  3. accuracy equals the do-nothing baseline
  4. is the rare class clustered at all?
  5. closeness can outweigh headcount

basics

~20 s

A majority of 15 needs 8 neighbours of the rare class, but in a region where that class is 4% of the data a neighbourhood of 15 typically contains none or one. The vote is therefore won by the common class everywhere, and the rare class is never predicted.

solid answer

~50 s

A plain vote returns whichever class holds the local plurality, so the rare class can only be predicted where it locally outnumbers the common one. With a 4% scrap rate on an assembly line and `k = 15`, a neighbourhood needs 8 defect points to win; unless defects are tightly clustered in feature space, a typical neighbourhood holds zero or one. Even sitting on a genuine defect, if the surrounding readings are mostly normal the vote still says normal. The symptom is a model reporting 96% accuracy with essentially zero recall on defects — it has reproduced the constant-majority baseline. Within the voting rule itself there are two levers: shrink `k` so the neighbourhood can fit inside a small defect cluster, and switch to distance-weighted voting so one defect sitting almost on the query outweighs fourteen normals two or three times further away. Reporting the vote share rather than the hard label also preserves the signal a hard argmax destroys.

go deeper

for a junior

Recall the arithmetic: a majority of 15 needs 8 neighbours of the rare class, and a class that is 4% of the data rarely supplies even one. Know that high accuracy under heavy skew can mean nothing.

for a middle

Explain that the vote reproduces local class proportions, so the rare class is predicted only where it is the local majority, and name the two in-rule levers: a smaller k and distance-weighted voting.

for a senior

Diagnose before fixing: establish whether the rare class is locally clustered by checking neighbour-label agreement around rare points, compare against the constant-majority baseline, and be explicit about the variance you buy by shrinking k.

for a principal

Own the harder call: if the rare class is not locally separable in the current feature space, no neighbourhood rule rescues it, and the right decision is to invest in features or labels rather than to keep tuning a voting scheme that cannot work.

## The mechanism, stated precisely A plain k-NN vote predicts the class with the most representatives among the k retrieved neighbours. That means the prediction for a query reflects the **local class proportions** in its neighbourhood. For the rare class to be predicted anywhere, there must exist regions of feature space where the rare class is locally in the majority. With a 4% scrap rate on an assembly line and `k = 15`, a majority requires 8 of 15 neighbours to be defects. If the defects are scattered through the same region of sensor space as the good units, a neighbourhood of 15 drawn from a region that is roughly 4% defect contains around half a defect on average — zero most of the time, occasionally one. Eight is not remotely reachable. The classifier collapses to predicting "normal" everywhere. Note what the argument does *not* depend on: it is not about the training procedure (there isn't one), the loss function, or optimisation. It is arithmetic on the vote. ## The symptom you will actually see - Accuracy around 96%, which is exactly the rate you get by always predicting "normal". - Recall on the defect class near zero, and precision undefined or reported as zero because no positives were ever predicted. - A confusion matrix with an empty predicted-positive column. Anyone reading only the headline accuracy will call this a strong model. That is why the first diagnostic question on an imbalanced problem is always "how does this compare with the constant-majority baseline?" ## When it does *not* happen The failure is about local density, not global prevalence, so a rare class survives the vote when it is **locally concentrated**. If defects arise from one physical fault mode that produces a tight, characteristic vibration-and-temperature signature, then inside that small region defects really are the local majority, and a k=15 vote will predict them there. A quick check: for each rare-class training point, how many of its own nearest neighbours share its label? If that agreement rate is high, the class is clustered and small-k voting can find it; if it is near the global prevalence, the rare class is scattered and no neighbourhood-majority rule will ever surface it. ## The two levers inside the voting rule **Shrink k.** A smaller neighbourhood can fit inside a small cluster. At `k = 3`, two defects out of three win; at `k = 1` a single stored defect near the query decides. This is the direct fix when the class is clustered, and it buys the usual cost: a more jagged boundary and more sensitivity to mislabelled points — including mislabelled defects, which at k=1 each mint their own false-positive region. **Weight by distance.** Under `1/d` or `1/d^2` weighting, one defect at distance 0.4 contributes far more than fourteen normal units at distances of 2 to 3. The weighted tally can favour the rare class where a count never could, and it does so without abandoning the stability of a larger nominal k. This is often the better first move because it degrades gracefully: where no defect is genuinely close, the weights leave the majority answer intact. **Report the vote share, not just the argmax.** The hard label throws away everything the neighbourhood knew. "2 of 15 neighbours are defects" is a very different signal from "0 of 15", and both are flattened to "normal" by the argmax. Keeping the (weighted) share as a continuous score preserves the ordering, which is what any downstream ranking or review queue actually needs. ## The judgment to bring to it The mistake to avoid is treating this as a defect in k-NN specifically. It is the general behaviour of any decision rule that maximises expected accuracy under skew: predicting the common class everywhere is genuinely the accuracy-optimal answer when the rare class is not locally separable. So the honest diagnosis is usually one of two conclusions. Either the features do not separate the rare class at all — in which case no k, no weighting and no other neighbourhood rule will help, and the fix is better features or better labels — or the class is clustered and simply drowned by a k chosen too large, which small-k or distance-weighted voting resolves directly. Distinguishing those two cases before touching anything else is the senior move here.

  • How would you check whether the rare class is locally clustered at all?
    Measure label agreement in the neighbourhood of the rare points themselves: for each defect in the training data, what fraction of its nearest neighbours are also defects? If that fraction is far above the 4% base rate, the class occupies identifiable regions and shrinking k or weighting will surface it. If it hovers near the base rate, the defects are scattered among normal units and no neighbourhood vote can separate them.
  • The model reports 96% accuracy. What do you say about that number?
    That it carries no information here, because always predicting the majority class scores 96% on a 4% positive rate. It is the baseline, not a result. Judge the model on per-class recall and on how it ranks cases, and compare any candidate against that constant-prediction baseline explicitly before claiming an improvement.
  • Would dropping k from 15 to 1 be a safe fix?
    It is the strongest lever and the riskiest one. At k=1 a single stored defect near the query decides the prediction, which recovers clustered defects — but it also means every mislabelled defect creates its own false-positive region, and the boundary becomes unstable under resampling. A moderate k with distance weighting usually recovers most of the signal for far less variance.

saying these in an interview costs you the question

  • Says raising k will help the rare class
  • Reads 96% accuracy as strong performance
  • Claims k-NN is immune to imbalance because it is non-parametric
  • Blames tie-breaking for the missing rare-class predictions
  • Assumes any fix works even when the rare class is not locally separable

context