In a soft-margin SVM, what does a large C trade against a small C?
answer
- one knob, two competing terms
- a price tag on each violation
- large C means expensive violations
- narrow margin, memorised noise
- C is inverse regularisation strength
basics
~20 sC sets the price of a margin violation. A large C makes violations expensive, so the boundary narrows its margin and bends around individual points. A small C tolerates violations and buys a wider, more heavily regularised boundary.
solid answer
~50 sA soft-margin SVM minimises `0.5*||w||^2 + C * sum(slack_i)`, where each slack measures how far a training point falls short of sitting a full unit outside the boundary. `C` is the exchange rate between those two terms. A large `C` — say 1000 — makes each violation expensive, so the optimiser shrinks the margin and contorts the boundary around individual points, including mislabelled ones: training error falls, variance rises, test error usually gets worse. A small `C` — say 0.01 — makes violations cheap, so the optimiser keeps `||w||` small and buys a wide margin at the cost of many points sitting inside it; pushed far enough the model underfits toward predicting one class. So `C` is an inverse regularisation strength. It does not say how many errors are allowed; it says what one unit of margin violation costs relative to margin width. Pick it on held-out data.
go deeper
Be ready to state the direction confidently: large C means expensive violations, narrow margin, more overfitting; small C means a wider margin and more tolerated violations. Say that you pick C on held-out data rather than by rule of thumb.
Write the objective and point at both terms, then show that dividing by C turns it into loss plus a penalty of strength 1/(2C). An interviewer expects you to explain why the margin term is an L2 penalty on the weights.
Show the diagnosis: what the training-versus-validation curves look like at each extreme, how boundary stability across resamples changes, and why a C tuned on one preprocessing pipeline or sample size is meaningless on another.
Own the framing that C is not a tuning detail but a stated bias-variance position. Argue when a wide-margin, heavily regularised fit is the right call for a low-signal, noisily labelled domain even though a larger C scores better on training metrics.
## The objective that C lives in A soft-margin support vector machine solves ``` minimise 0.5*||w||^2 + C * sum_i slack_i subject to y_i * (w . x_i + b) >= 1 - slack_i, slack_i >= 0 ``` Here `y_i` is the label coded as `+1` or `-1`, `w . x_i + b` is the raw score the model gives point `i`, `w` is the weight vector, and `slack_i` is how much that point is allowed to fall short of the requirement `y_i * (w . x_i + b) >= 1`. Why the slacks exist at all: think of a loan-default classifier where a few dozen applicants sit squarely in the overlap between repaid and defaulted — same income, same debt ratio, different outcomes. Demand that every point be on the correct side and the problem has no solution; the fit simply does not exist. Slack variables let each point buy its way inside the margin, or even across the boundary, for a price. `C` is that price. ## What the two terms want The two terms pull in opposite directions. - `0.5*||w||^2` wants small weights. The width of the margin is `2/||w||`, so shrinking the weight vector widens the margin. This term is literally an L2 penalty on the weights: a soft-margin SVM is a regularised linear model whose data-fitting loss happens to be hinge loss. - `C * sum(slack_i)` wants every training point at least one full unit of score on its own side of the boundary. Every unit of shortfall costs `C`. One knob balances them, and it is worth writing the objective the other way round. Divide through by `C`: ``` minimise sum_i slack_i + (1/(2C)) * ||w||^2 ``` Now it reads as loss plus penalty with regularisation strength `1/(2C)`. Large `C` means weak regularisation; small `C` means strong regularisation. Candidates who memorise "large C = complex model" and nothing else get caught by any question phrased in terms of regularisation strength. ## Large C versus small C on the same data Take overlapping data and fit it twice. With `C = 1000`, a single unit of violation costs a thousand times what a unit of `0.5*||w||^2` costs, so the optimiser will happily grow `||w||` — narrowing the margin — to rescue a handful of awkward points. The boundary develops kinks around noisy or mislabelled observations. Training accuracy climbs toward perfect. Because the boundary is now determined by the few weirdest points in the sample, a fresh sample moves it a lot: high variance, worse generalisation. With `C = 0.01`, violations are nearly free. The optimiser keeps `||w||` tiny, the margin is very wide, and a large fraction of the training set sits inside it. The boundary is smooth and stable across resamples, but at some point it stops responding to the data at all and collapses toward predicting the majority class everywhere: high bias. A useful side effect: small `C` produces *more* points on or inside the margin, so the solution rests on more points; large `C` concentrates it on fewer. Candidates often state this backwards. ## What C is not - It is **not** a cap on the number of misclassified training points. It prices violations; it does not count them. - It is **not** dimensionless in any practical sense. The loss term sums over `n` points while the penalty term does not, so the same `C` acts more aggressively on a larger training set — which is why some formulations write `C/n`. - It is **not** comparable across different feature scalings. Change the units of a feature and the same `C` gives you a different model. - It is **not** monotone in *accuracy*. Raising `C` weakly reduces the total hinge loss at the optimum, but the raw misclassification count can wobble, because hinge loss is a surrogate for the error rate rather than the error rate itself. ## Choosing it There is no default worth trusting. Sensible `C` values span orders of magnitude, so search on a log scale and judge each value on held-out data, not on training error — training error is monotone-ish in `C` by construction and will always flatter the largest value you try. Report the chosen `C` alongside the preprocessing that produced it, because the two are inseparable.
- How would you spot from validation results that C is set too high?Training accuracy climbs toward perfect while held-out accuracy plateaus and then falls as C grows, and the gap between them widens. The fitted boundary also becomes unstable: refit on a resample of the same data and it moves noticeably. Both are signatures of variance bought by pricing violations too dearly.
- Is the same C comparable across two datasets of very different size?No. The violation term sums over all training points while the margin penalty does not, so doubling the sample size roughly doubles the weight of the loss term at fixed C — the same C is effectively less regularising on the larger set. Re-tune C whenever the training-set size changes materially, and be wary of formulations that scale C by n.
- Does raising C always reduce training error?It weakly reduces the total hinge loss at the optimum, since violations get more expensive relative to margin width. The misclassification count is not guaranteed to fall monotonically, because hinge loss is a convex surrogate rather than the error rate itself, so a step in C can trade one large violation for several small ones.
C is the fine for parking over the line. Set a huge fine and every car squeezes into a narrow bay; set a token fine and the bays stay generously wide while plenty of cars stick out.
saying these in an interview costs you the question
- Says C is the number of misclassifications the model is allowed
- Claims a larger C always generalises better because it fits more points
- Thinks C changes the loss function rather than its weight in the objective
- Says small C yields fewer points on or inside the margin
- Treats C as comparable across datasets and feature units