kNN, Naive Bayes, and SVMs
Three classical classifiers, each built on a different idea: memorise neighbours, multiply likelihoods, maximise margins. Interviewers use them to test assumptions and failure modes.
on this pageshowhide
explore
- Nearest Neighbour Methods16 questions
- Neighbour Voting and k4 questions
- Distance Metrics4 questions
- Curse of Dimensionality4 questions
- Exact Neighbour Search4 questions
- Naive Bayes Classifiers7 questions
- Conditional Independence Assumption3 questions
- Likelihood Variants and Smoothing4 questions
- Support Vector Machines15 questions
- Maximum-Margin Classifier4 questions
- Soft Margin and Hinge Loss3 questions
- Kernel Trick4 questions
- Multiclass, SVR and Runtime4 questions
questions
page 2 of 2How do you compute distance for a record mixing salary, tenure, department and a remote flag?
basics
~20 sBuild a per-feature dissimilarity that already lives on a 0-to-1 scale and average those, which is what Gower's coefficient does: absolute difference over the feature's range for numbers, zero or one for a categorical match or mismatch.
In naive Bayes, which correlated features actually flip the predicted class?
basics
~20 sOnly redundant evidence that favours a class other than the honest winner, and whose repeated counting is large enough to overcome the honest margin. Correlated features that reinforce the class that would have won anyway just inflate confidence without changing the label.
What does Mercer's condition require, and why isn't every similarity function a valid kernel?
basics
~20 sA kernel must be symmetric and produce a positive semi-definite matrix of pairwise values on every finite set of points. Only then does it correspond to a genuine inner product in some feature space, which is what the whole method assumes.
Your naive Bayes prior comes from a 90/10 label count but deployment runs 50/50 — what do you change?
basics
~20 sSwap the prior, do not retrain. The class prior enters the score as one additive log term, so replacing the training log-prior with the deployment class mix corrects the model, provided the per-class feature distributions themselves have not changed.
A hard-margin SVM separates 200 rows of 5,000-feature data perfectly. Why is that unsurprising?
basics
~20 sWith far more features than rows, points in general position can be separated by a hyperplane for essentially any labelling, including random ones. Perfect separation therefore carries almost no information; the width of the margin relative to the data's spread is the quantity worth reading.
In a one-vs-one SVM router, two classes tie on votes — why not compare margin scores?
basics
~20 sEach pairwise SVM produces a signed distance measured in its own weight-norm units, learned from its own two classes. Those numbers share no common scale and are not probabilities, so comparing them across pairs is arbitrary.
What does condensed nearest neighbour remove from a training set, and what does it risk?
basics
~10 sCondensed nearest neighbour keeps a subset that still classifies every original training point correctly under 1-NN. It discards interior points and keeps boundary ones. The risk: mislabelled points are exactly what it preserves.
Why does a k-NN majority vote with k=15 almost never predict a class that is 4% of the data?
basics
~20 sA majority of 15 needs 8 neighbours of the rare class, but in a region where that class is 4% of the data a neighbourhood of 15 typically contains none or one. The vote is therefore won by the common class everywhere, and the rare class is never predicted.
showing 31–38 of 38