skip to content

kNN, Naive Bayes, and SVMs

Three classical classifiers, each built on a different idea: memorise neighbours, multiply likelihoods, maximise margins. Interviewers use them to test assumptions and failure modes.

on this pageshow

explore

questions

page 2 of 2

How do you compute distance for a record mixing salary, tenure, department and a remote flag?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Build a per-feature dissimilarity that already lives on a 0-to-1 scale and average those, which is what Gower's coefficient does: absolute difference over the feature's range for numbers, zero or one for a categorical match or mismatch.

open as a page

In naive Bayes, which correlated features actually flip the predicted class?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Only redundant evidence that favours a class other than the honest winner, and whose repeated counting is large enough to overcome the honest margin. Correlated features that reinforce the class that would have won anyway just inflate confidence without changing the label.

open as a page

What does Mercer's condition require, and why isn't every similarity function a valid kernel?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

A kernel must be symmetric and produce a positive semi-definite matrix of pairwise values on every finite set of points. Only then does it correspond to a genuine inner product in some feature space, which is what the whole method assumes.

open as a page

Your naive Bayes prior comes from a 90/10 label count but deployment runs 50/50 — what do you change?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

Swap the prior, do not retrain. The class prior enters the score as one additive log term, so replacing the training log-prior with the deployment class mix corrects the model, provided the per-class feature distributions themselves have not changed.

open as a page

A hard-margin SVM separates 200 rows of 5,000-feature data perfectly. Why is that unsurprising?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

With far more features than rows, points in general position can be separated by a hyperplane for essentially any labelling, including random ones. Perfect separation therefore carries almost no information; the width of the margin relative to the data's spread is the quantity worth reading.

open as a page

In a one-vs-one SVM router, two classes tie on votes — why not compare margin scores?

level: seniorimportance: nice to knowfreq 24%

basics

~20 s

Each pairwise SVM produces a signed distance measured in its own weight-norm units, learned from its own two classes. Those numbers share no common scale and are not probabilities, so comparing them across pairs is arbitrary.

open as a page

What does condensed nearest neighbour remove from a training set, and what does it risk?

level: seniorimportance: nice to knowfreq 16%

basics

~10 s

Condensed nearest neighbour keeps a subset that still classifies every original training point correctly under 1-NN. It discards interior points and keeps boundary ones. The risk: mislabelled points are exactly what it preserves.

open as a page

Why does a k-NN majority vote with k=15 almost never predict a class that is 4% of the data?

level: seniorimportance: nice to knowfreq 32%

basics

~20 s

A majority of 15 needs 8 neighbours of the rare class, but in a region where that class is 4% of the data a neighbourhood of 15 typically contains none or one. The vote is therefore won by the common class everywhere, and the rare class is never predicted.

open as a page

showing 31–38 of 38