In naive Bayes, which correlated features actually flip the predicted class?
answer
- sum of log-likelihood ratios, sign decides
- a cluster of k adds k terms
- direction versus honest margin
- one-sided duplication is the dangerous kind
- measure correlation inside each class
basics
~20 sOnly redundant evidence that favours a class other than the honest winner, and whose repeated counting is large enough to overcome the honest margin. Correlated features that reinforce the class that would have won anyway just inflate confidence without changing the label.
solid answer
~50 sThink in log space: the binary decision is the sign of the log prior ratio plus one log-likelihood-ratio term per feature. A cluster of `k` near-duplicate features enters that sum roughly `k` times instead of once, so it moves the total by about `k` times its true contribution. Whether the label changes depends on direction and margin. If the duplicated cluster points at the class that already had the larger score, you get an inflated posterior and the same prediction. If it points the other way, its multiplied weight can drag the sum across zero and overturn genuine independent evidence for the correct class — that is the flip. Correlated features whose members are split, some favouring each class, partly cancel and are far less dangerous. Two credit-bureau fields like number of open cards and total credit limit are the classic case: they move together, so their agreement is counted as independent corroboration and can outvote everything else in the model.
code
python · 22 linesprior = {'spam': 0.5, 'ham': 0.5}
# P(feature fires | class). 'free' and 'free!!!' are one signal seen twice.
lik = {
'free': {'spam': 0.6, 'ham': 0.2},
'free!!!': {'spam': 0.6, 'ham': 0.2},
'invoice': {'spam': 0.1, 'ham': 0.5},
}
def posterior(features):
score = {}
for c in prior:
p = prior[c]
for f in features:
p *= lik[f][c]
score[c] = p
total = sum(score.values())
return {c: round(v / total, 3) for c, v in score.items()}
print(posterior(['free', 'invoice']))
# {'spam': 0.375, 'ham': 0.625} -> predicts ham
print(posterior(['free', 'free!!!', 'invoice']))
# {'spam': 0.643, 'ham': 0.357} -> predicts spam: the label flippedgo deeper
Know the basic shape: repeated versions of the same signal each get a vote, so the model can be talked into the wrong answer by evidence it heard once but counted twice.
Work the log-space argument: one term per feature, a cluster of near-duplicates adding several nearly identical terms, and a decision made by the sign of the total.
Show the diagnostic and the remedy. Measure association within each class, identify features derived from the same source, and choose between merging, dropping and recalibrating based on which direction the cluster pushes.
Own the trade between feature hygiene and lost signal. Decide when it is cheaper to invest in deduplicating the feature set and when to accept the distortion and correct the scores downstream.
## The decision as a sum, not a product For two classes, naive Bayes predicts class `c1` when ``` log(P(c1)/P(c2)) + sum over features of log(P(x_i|c1)/P(x_i|c2)) > 0 ``` Every feature contributes exactly one term, and that term's size measures how discriminative the feature is: a feature that fires three times as often in `c1` contributes about `log(3)`, and one equally common in both contributes zero. This view makes the effect of dependence easy to reason about, because double counting becomes double *adding* rather than something buried inside a product. ## What a correlated cluster does Suppose `k` features are all near-restatements of a single underlying signal — several spellings of the same spam token, a symptom recorded both as a flag and as a threshold on a measurement, or several bureau fields all driven by how much credit a person has. Conditionally on the class, they are strongly dependent, which is exactly what the model assumes away. The honest contribution of that signal is roughly one term. The model adds about `k` of them. Three cases follow. 1. **The cluster points at the class that would have won anyway.** The sum moves further in the direction it was already going. The sign is unchanged, so the label is unchanged; only the posterior becomes absurd. This is the common case, and it is why the model gets away with the assumption so often. 2. **The cluster points at the losing class, and `k` times its term exceeds the honest margin.** The sign flips and the prediction is wrong in a way that the same features, counted once, would not have produced. This is the failure the assumption actually costs you, and it is worth noticing that the damage scales with *both* how redundant the cluster is and how discriminative each member looks. 3. **The correlated group is split — some members favour one class, some the other.** Their inflated terms partly cancel, and the sum ends up closer to its honest value than the raw dependence would suggest. Dependence per se is not the enemy; systematically one-sided dependence is. A compact illustration: give a spam model the tokens `free` and `free!!!` (one signal, both leaning spam) and one genuinely independent feature that leans ham. Scored with the duplicate removed, ham wins; scored with the duplicate kept, spam wins. Nothing about the truth changed — only how many times one piece of evidence was allowed to vote. ## How to find the dangerous clusters The diagnostic that matters is **correlation within each class**, not overall correlation. Two features can correlate strongly across the dataset simply because both track the label; that is the structure the model is supposed to exploit, and it is harmless. Split the training data by class and measure association among features inside each group — for continuous features a correlation matrix per class, for categorical or binary features a per-class association measure such as mutual information. Pairs that stay strongly associated after conditioning on the class are the double counters. Domain review usually beats statistics here. Features derived from the same raw field, features produced by the same upstream service, a flag and the measurement it was thresholded from, and the same text token under several normalisations are all obvious candidates and can be caught by reading the feature list. ## What to do about them - **Merge the cluster into one feature** that carries the joint signal — a single flag for "any spelling of the token", a single derived ratio instead of two raw bureau fields. This is usually better than deletion because it keeps the information while entering it once. - **Drop all but one member.** Cheap and effective when the members really are restatements, but it loses whatever small independent content the others carried. - **Down-weight the cluster** so its members jointly contribute about one term's worth of evidence. This is a defensible engineering fix, though it moves the model away from being a clean probabilistic story. - **Leave it and recalibrate downstream** when the cluster reliably points at the correct class. You accept the inflated posterior and correct it as a separate calibrated mapping. Legitimate when accuracy is fine and only the numbers are wrong. ## The judgment an interviewer is testing A weak answer is "correlated features break naive Bayes, so remove correlated features". The strong answer distinguishes harmless amplification from an actual flip, names the within-class diagnostic rather than the marginal one, and picks a remedy that fits which of the three cases you are in. It should also concede the cost of over-cleaning: strip every correlated feature and you may throw away real signal to defend an assumption that was never going to hold exactly.
- How would you find the double-counted features before shipping the model?Measure association among features separately inside each class, not across the whole dataset — pairs that stay strongly related after conditioning on the label are the double counters. Pair that with a read of the feature list for items derived from the same raw field, the same upstream service, or the same text token under different normalisations.
- Is dropping one of a correlated pair always the right fix?Not always. Merging the pair into a single feature that carries the joint signal keeps the information while entering it once, which usually beats deletion. And when the redundant cluster reliably points at the correct class, leaving it and recalibrating the score downstream costs you nothing in accuracy.
- Why is overall feature correlation the wrong thing to check?Because much of it is caused by the label itself: two features move together across the dataset precisely because both track the class, and that structure is what the model exploits. Only the association that remains after conditioning on the class violates the assumption, so the marginal correlation matrix flags harmless pairs and misses nothing useful.
saying these in an interview costs you the question
- Says any feature correlation makes naive Bayes unusable
- Checks overall correlation instead of correlation within each class
- Assumes double counting always changes the predicted label
- Thinks adjusting the class prior repairs double-counted evidence
- Removes every correlated feature without weighing the lost signal