skip to content

A hard-margin SVM separates 200 rows of 5,000-feature data perfectly. Why is that unsurprising?

level: seniorimportance: nice to knowfreq 28%

answer

  1. counting features against rows
  2. any labelling at all, even random ones
  3. separation stops being evidence
  4. read the width, not the fact
  5. shuffle the labels and refit

basics

~20 s

With far more features than rows, points in general position can be separated by a hyperplane for essentially any labelling, including random ones. Perfect separation therefore carries almost no information; the width of the margin relative to the data's spread is the quantity worth reading.

solid answer

~50 s

Separability gets cheap as dimensions grow. If the points are in general position and the number of features is at least the number of rows minus one, a hyperplane exists that realises *any* labelling of them - including labels drawn at random. With 200 rows and 5,000 features you are far past that threshold, so "the training data is perfectly separated" is a statement about the geometry of the feature space, not about the signal in the data. What still carries information is the width of the achieved margin relative to the radius of the data cloud, since generalisation arguments for margin classifiers depend on that ratio rather than on the feature count. The cheap diagnostic is a label-permutation check: shuffle the labels, refit, record the margin, repeat. If the real margin sits inside the distribution of shuffled margins, the separation is an artefact of dimensionality. And whatever the geometry says, a held-out estimate is what decides.

go deeper

for a junior

Recall the headline: when there are far more features than rows, a separating hyperplane almost always exists, so perfect training separation is not by itself good news.

for a middle

Explain the mechanism - points in general position with at least as many features as rows minus one can be separated under any labelling - and state what to look at instead, namely the margin width and a held-out estimate.

for a senior

Show the diagnostic instinct: build a shuffled-label baseline calibrated to this sample size and feature count, split before feature selection, and report an interval rather than a point estimate on 200 rows.

for a principal

Own the reporting standard. Decide what evidence a wide-data model must present before it influences a decision, and push back on demonstrations whose headline number is a quantity the geometry guarantees.

## The claim that needs deflating "The classifier separates the training data perfectly" sounds like evidence, and in a low-dimensional problem it is. In a wide problem - many more features than rows - it is nearly automatic, and treating it as evidence is one of the classic ways to fool yourself with a margin classifier. ## Why separability is cheap when features outnumber rows Take `n` points in general position in a `d`-dimensional feature space. "General position" means no degenerate coincidences - no more of them lying on a common flat than necessary, which is the typical case for real-valued measurements. If `d >= n - 1`, those points are affinely independent, and an affine hyperplane exists that realises *every possible* assignment of the two labels among them. Not a particular convenient labelling: all of them, including labels generated by coin flips with no relationship to the features whatsoever. With 200 rows and 5,000 features you are far above that threshold. A hard-margin fit that separates the data is doing what the geometry guarantees it can do. The information content of "it separated" is close to zero. This is not an argument that the model must be wrong. Wide problems with real signal exist everywhere - gene expression, text, sensor arrays. It is an argument that separation alone cannot distinguish the two cases. ## What does carry information The margin's **width**, read relative to the scale of the data. Generalisation arguments for large-margin separators are usually stated in terms of the ratio `R^2 / margin^2`, where `R` bounds the radius of the region containing the data. The number of features does not appear explicitly. That is precisely why margin classifiers can be sane in high dimensions at all - but the guarantee is contingent on the margin staying wide relative to `R`, and in a labelling that is separable only because `d` is large, the achievable margin is typically razor-thin. So the question to ask is not "did it separate?" but "how much room did it have?", and "how much room would it have had on data with no signal at all?" ## The permutation baseline The second question has a direct empirical answer. Shuffle the labels so any relationship to the features is destroyed, refit the hard-margin problem, and record the resulting margin. Repeat many times to build the distribution of margins achievable on pure noise with these exact feature dimensions and this exact sample size. Then compare: - If the real-label margin sits comfortably above that whole distribution, the separation reflects structure the features actually carry. - If it sits inside the distribution, you have measured the geometry of your feature space, not a signal. The baseline is valuable because it is calibrated to your `n` and `d` rather than to an asymptotic rule of thumb, and because it is easy to explain to a stakeholder who has just been shown a perfect training separation. ## What to do about it The honest reading of a perfect separation in a wide problem is: unmeasured. Estimate performance on data the model never saw, with the split made before any feature selection - selecting features using all the labels and then splitting is a leak that manufactures separation on the held-out set too. Expect the held-out estimate to be noisy with 200 rows, and report an interval rather than a point. If the effective dimension can be reduced on principled grounds before fitting, do it, but not by peeking at the outcome. Also record the support-vector fraction. A hard-margin fit on wide data frequently makes a large share of the rows support vectors, which says the frontier is crowded and the boundary is held in place by most of the data - the opposite of the compact, well-separated regime the margin argument is comfortable in. ## The general lesson Any statement of the form "the model fits perfectly" is only informative relative to how easy fitting perfectly was. In wide problems it is easy by construction, so the burden shifts entirely onto out-of-sample evidence and onto quantities - margin width, permutation baselines - that are not automatically satisfiable.

  • How would you check whether the perfect separation reflects real signal?
    Build a permutation baseline: shuffle the labels many times, refit each time, and record the margins achievable on pure noise at this sample size and feature count. Compare the real-label margin against that distribution. Then confirm with a held-out estimate, splitting before any feature selection so the selection cannot leak label information.
  • Does theory say a margin classifier must overfit when features outnumber rows?
    No. Margin-based bounds depend on the ratio of the data radius to the margin, not explicitly on the number of features, which is why these models can work in wide problems. The catch is that the bound is only useful while the margin stays wide relative to the data's spread, and separations that exist purely because of dimensionality tend to be thin.
  • What would the support-vector count tell you in this situation?
    If a large share of the 200 rows end up as support vectors, the frontier is crowded and the boundary depends on most of the data rather than a small set of clean extremes. That is a signal that the separation is tight and fragile, and it makes the leave-one-out bound tied to the support-vector fraction useless.

saying these in an interview costs you the question

  • Treats perfect training separation as evidence of signal
  • Says more features always help a margin classifier
  • Cannot connect separability to the feature-count versus row-count comparison
  • Selects features on all the data, then splits
  • Reports a single held-out number from 200 rows without an interval

context