skip to content

Your team defaults to an RBF-kernel SVM on every tabular problem. When is that default wrong?

level: principalimportance: should knowfreq 37%

answer

  1. the kernel is a prior about similarity
  2. does Euclidean distance mean anything here?
  3. mixed categorical columns have no natural metric
  4. distances concentrate in high dimensions
  5. no per-feature coefficients to hand over

basics

~20 s

A kernel is an assumption about which points count as similar. The RBF default fails when Euclidean distance is not meaningful — heterogeneous or mostly-categorical features, very high dimensions where distances concentrate — or when the problem needs readable per-feature effects.

solid answer

~60 s

Choosing a kernel is choosing a prior, not a default. An RBF kernel says two points are similar when they are close in Euclidean distance, and that claim has to be true of your data for the model to work. It fails on tables of mixed categorical and numeric columns where no natural distance exists, and it weakens in very high dimensions where pairwise distances concentrate and the similarity carries little signal. A polynomial kernel encodes a different prior — interactions up to a fixed order, applied globally — and is the better fit when your domain knowledge really is "products of features matter up to degree three". A linear kernel is right when the classes are close to separable in the original space, and it is the only one that leaves you per-feature coefficients. The alternatives worth naming are a linear model when interpretability matters and a tree ensemble on heterogeneous tables, since trees never use distances at all. Where kernels genuinely win is structured inputs — sequences, graphs, strings — where a domain kernel expresses similarity that no column set captures.

go deeper

for a junior

Recall that the kernel is a choice with consequences, not a default setting. Know that a linear kernel is the sensible baseline and that flexibility bought by an RBF kernel has to be justified by results on held-out data.

for a middle

Explain what each standard kernel assumes: linear says a flat boundary suffices, polynomial says interactions up to a degree matter globally, RBF says similarity decays with Euclidean distance. Be able to name a dataset shape that violates each.

for a senior

Demonstrate the check rather than the preference. Ask whether distance is meaningful over these columns, whether irrelevant features are drowning the signal, and whether the deliverable needs feature-level explanation before you commit to a kernel at all.

for a principal

Own the policy, not the pick. Argue for a baseline-first escalation ladder, a parallel candidate family for heterogeneous tables, and a standing requirement that kernel hyperparameters be searched on this data rather than inherited — then defend that policy against the pull of habit.

## Restate the question before answering it "Which kernel?" is really "what does similarity mean for these objects?" A kernel method has one assumption at its heart: points the kernel scores as similar should tend to share a label. Every kernel is a different, specific claim about that, and the default is wrong exactly when its claim is false of your data. ## What each standard kernel claims - **Linear** (`x . z`): the classes are close to separable by a flat boundary in the original coordinates. Common with wide, sparse data where there are already more features than rows and there is little left for a lift to add. It is also the only option that hands you a weight per feature, which matters when someone will ask why a decision was made. - **Polynomial** (`(x . z + c)^d`): interactions among features up to order `d` matter, globally and uniformly across the input space. Degree `d` is the prior's strength — `d = 2` says pairs of features interact, `d = 3` adds triples. Raising `d` multiplies capacity fast, and high degrees are numerically awkward because kernel values swing across orders of magnitude as the dot product moves either side of one. - **RBF** (`exp(-gamma * ||x - z||^2)`): similarity decays smoothly with Euclidean distance, and the target function is locally smooth. This is the most flexible and the least committal, which is why it is the reflexive default — and why it is also the one whose assumption is most often quietly violated. ## Where the RBF default actually breaks **Heterogeneous tabular data.** Squared Euclidean distance over a table containing an age in years, an income in currency units, three one-hot indicator blocks and an ordinal rating is a sum of incomparable quantities. Any distance you compute is dominated by whichever columns happen to have the widest numeric range, and "nearby" stops meaning anything about the domain. Trees never compute a distance at all; they threshold one feature at a time, which is why gradient-boosted ensembles are usually the stronger reflex on this shape of data. **Very high dimension with many irrelevant features.** As dimension rises, pairwise Euclidean distances between random points concentrate: the nearest and the farthest neighbour become nearly equidistant. Feed that into `exp(-gamma * ||x - z||^2)` and the kernel returns nearly the same value for every pair, so the similarity has little to discriminate with. Irrelevant features make it worse, because each one contributes to the distance without contributing to the label. A kernel has no built-in feature selection — every column enters the distance with equal weight. **Explanation requirements.** A kernel model has no weight vector in the original space to read. If the deliverable includes per-feature effects for a regulator, a clinician or a product owner, you either accept a model-agnostic explanation layer on top or you choose a family that is transparent by construction. **Operational shape.** The trained model is not a self-contained weight vector; it carries its support vectors and evaluates the kernel against them for each prediction. That is a different artefact to ship and version than a linear model's coefficients, and it is worth being deliberate about rather than discovering at deployment. ## Where kernels are the right call Structured inputs are the strongest case: sequences, strings, sets, graphs and other objects that are not naturally a fixed row of numbers. A domain-designed kernel can express "these two protein sequences share motifs" or "these two graphs have similar substructure" directly, without anyone having to invent a column set that captures it. That is the situation where the implicit feature space is genuinely doing work that no explicit table could. Small-to-moderate datasets with dense numeric features on a common scale — sensor arrays, spectra, image descriptors — are the classic tabular case where an RBF kernel earns its keep, because there Euclidean distance really does mean something. ## What to standardise instead of a kernel The defensible team-level policy is not "use kernel X" but "justify the similarity assumption". Concretely: start from a linear baseline so you can see what the lift is worth; escalate to a kernel only when the baseline's residuals show structure it cannot express; keep a tree ensemble as the parallel candidate on any heterogeneous table; and require that a chosen kernel's hyperparameters were searched on held-out data rather than inherited from a previous project, since a kernel width tuned on one feature space is meaningless on another. The failure mode a lead should be most alert to is not picking the wrong kernel — it is picking one by habit and never checking whether its similarity claim survives contact with the data.

  • Why does a tree ensemble often beat a kernel SVM on heterogeneous tabular data?
    Because it never computes a distance. A tree splits one feature at a time on a threshold, so incomparable units, wildly different scales and categorical columns cost it nothing, and irrelevant features are simply not chosen for splits. An RBF kernel, by contrast, sums every column into one squared distance, letting whichever column has the widest range dominate the notion of similarity.
  • When would you prefer a polynomial kernel over an RBF one?
    When your prior is specifically about interactions of a known order rather than local smoothness — you believe products of features up to degree two or three drive the outcome and that the relationship is global, not neighbourhood-shaped. The polynomial kernel encodes exactly that and has one interpretable dial, the degree, instead of a width whose meaning depends on the data's scale.
  • What is the strongest case for a custom kernel over engineering features?
    Structured objects that are not naturally rows of numbers — sequences, strings, sets, graphs. A domain kernel can score "these two share the same substructure" directly, whereas flattening them into columns either loses that structure or explodes into an unusable number of features. Compose the custom kernel from valid pieces so it stays a legitimate kernel.

saying these in an interview costs you the question

  • Treats the RBF kernel as a universally safe default
  • Ignores that Euclidean distance is meaningless over mixed units
  • Assumes more flexible kernels always beat a linear one
  • Forgets that a kernel model offers no per-feature coefficients
  • Reuses a kernel width tuned on a different dataset

context