In a soft-margin SVM, why does feature scale change what a given C does?
answer
- the margin is measured in feature units
- dollars sitting beside a ratio in [0,1]
- a wide-range feature needs a tiny weight
- the norm penalty barely restrains it
- same C, different model after scaling
basics
~20 sThe margin is a Euclidean distance in feature space, so a feature in dollars dwarfs one in [0,1]. Rescaling changes the weight vector's length and therefore what a unit of margin costs, so C must be re-tuned after any scaling change.
solid answer
~50 sTake a loan model with applicant income in dollars, roughly 0 to 200,000, beside a debt ratio in `[0,1]`. The margin is `2/||w||`, a Euclidean distance in whatever units the features carry, and the penalty `0.5*||w||^2` charges by weight magnitude. Income needs only a minuscule weight to swing the score across its range, so the penalty barely restrains it; the debt ratio needs a weight roughly a hundred thousand times larger for comparable influence, and the same penalty punishes that hard. The unscaled fit is therefore close to an income-only model with the ratio suppressed — not because income is more predictive, but because of its units. Standardise or min-max the features, computing the statistics on the training portion only, then re-run the search for `C`: a value tuned on raw dollars carries no meaning on standardised inputs.
go deeper
Be ready to say that a soft-margin SVM is distance-based and not scale-invariant, so features must be put on comparable scales before fitting, unlike a tree-based model.
Explain the mechanism: the margin is 2 divided by the norm of the weight vector, the penalty charges squared weights, and a wide-range feature achieves influence with a tiny weight the penalty hardly notices.
Demonstrate operational judgment: re-tune C after any preprocessing change, fit the scaler on training data only, and version the scaler and C together. Diagnose a suspiciously single-feature model by checking the input ranges first.
Own the pipeline contract. Argue for treating preprocessing and hyperparameters as one versioned artefact across training and serving, and for a review habit that treats an untuned C after a preprocessing change as a defect rather than a detail.
## The geometry is measured in feature units A soft-margin SVM minimises `0.5*||w||^2 + C * sum(slack)`. Both terms are expressed in the coordinate system you hand it. The margin width is `2/||w||`, a Euclidean distance in raw feature space, and the penalty term charges the squared length of the weight vector in those same coordinates. Nothing in the formulation normalises anything, so the units you chose for each column silently become part of the model. Concretely: applicant income in dollars ranging over roughly 0 to 200,000, sitting next to a debt ratio confined to `[0,1]`. ## Why the large-range feature wins Ask what weight each feature needs to move the score by one unit across its observed range. - Income spans about 200,000 units, so a weight around `5e-6` already swings the score by one. Squared, that contributes about `2.5e-11` to the penalty — effectively free. - The debt ratio spans one unit, so it needs a weight near `1` for the same swing, contributing about `1` to the penalty — some eleven orders of magnitude more expensive. The optimiser is minimising that penalty, so it will lean on income and shrink the ratio's weight toward zero unless the ratio buys back a great deal of hinge loss. The fitted model is dominated by the feature with the widest numeric range, regardless of which feature actually carries signal. Any kernel built from distances between points inherits the same distortion, because the distance itself is dominated by the large-range coordinate. ## Why C stops meaning anything The two terms are in a fixed exchange rate set by `C`. Rescaling a feature by a factor `s` rescales the corresponding weight by `1/s` at the same decision boundary, which changes `||w||^2` — often by many orders of magnitude — while the hinge term, which depends on `y*f(x)`, stays put for the equivalent boundary. The balance between the two terms shifts, so the optimum shifts. The practical consequence is blunt: **a `C` tuned on raw features is not transferable to scaled features, and vice versa.** If you standardise and keep yesterday's `C`, you have silently chosen a different point on the bias-variance curve. Re-run the search over `C` after any change to the preprocessing, and record the preprocessing and the chosen `C` together — they are one artefact, not two. A second consequence is that the search itself becomes ill-conditioned on unscaled data. Sensible `C` values now depend on the feature units, so the usual log-spaced range may sit entirely in the underfitting or overfitting regime, and you will conclude the model is bad when the units were bad. ## Doing it without leaking Compute the scaling statistics — means and standard deviations, or minima and maxima — on the training portion of each split, then apply those same numbers to the held-out portion. Computing them over the full dataset before splitting lets information from the evaluation data into the fit, which inflates the measured score in a way that will not survive deployment. At serving time, the stored training statistics are applied to each incoming record. ## What scaling does not fix - **Heavy tails.** Standardising income leaves the same skew: a handful of very high earners still sit far from the mass, and at a large `C` they can still dominate the boundary. A monotone transform such as a log applied before standardising often does more good than the standardisation itself. - **Outliers in the scaler.** Min-max scaling is defined by the extremes, so one absurd value squashes every other observation into a narrow band. Standardisation is less brittle but still moves when the tail moves. - **Genuinely irrelevant features.** Putting a noise column on the same scale as a signal column makes it *easier* for the model to use the noise, not harder. ## The interview answer in one line An SVM is a geometric method with an L2 penalty, and geometry plus a norm penalty means units matter. Scale the inputs, tune `C` afterwards, and treat the scaler and `C` as a single fitted object.
- After standardising, do you keep the C you tuned on the raw features?No. Rescaling a feature rescales its weight and therefore the squared norm, which shifts the balance between the penalty and the violation term. The old C now sits at a different point on the bias-variance curve. Re-run the search on the scaled representation and store the scaler and the chosen C together as one artefact.
- Where should the scaling statistics be computed so nothing leaks?On the training portion of each split only, then applied unchanged to the held-out portion and, later, to live records. Computing means and standard deviations over the full dataset before splitting lets the evaluation data influence the fit, which inflates the measured score by an amount you cannot see and will not reproduce in production.
- Does scaling change which points end up violating the margin?Yes, and that is the point. Scaling changes the geometry, so the optimal boundary is genuinely different and a different set of points ends up on or inside the margin. It is not a numerical convenience that leaves the model intact; it changes what model is fitted.
It is like judging which house is nearer on a map whose east axis is in metres and north axis in kilometres. Whichever axis carries the bigger numbers decides every distance you compute.
saying these in an interview costs you the question
- Says scaling only matters for optimiser convergence speed
- Claims an SVM is scale-invariant the way a decision tree is
- Keeps a C tuned on raw features after standardising
- Computes scaling statistics on the whole dataset before splitting
- Assumes the largest-range feature is dominant because it is most predictive