When should you use LinearSVC or SGDClassifier instead of SVC in scikit-learn?
answer
- libsvm dual versus liblinear primal
- Kernel pairs grow with the square
- A row-count threshold, not a rule of thumb
- cache_size helps, asymptotics do not change
- Approximate the kernel, then fit linear
basics
~20 sSVC's kernel solver has fit time scaling at least quadratically in the number of samples, so it becomes impractical past roughly tens of thousands of rows. For large linear problems reach for LinearSVC, backed by liblinear, or SGDClassifier, whose cost is linear in the sample count.
solid answer
~50 s`SVC` wraps libsvm and solves the dual problem over kernel evaluations, so fit time scales at least quadratically with `n_samples` — scikit-learn's own documentation warns it is impractical beyond tens of thousands of rows, and the `cache_size` knob (200 MB by default) only softens the kernel-matrix thrashing. If your problem is linear, `LinearSVC` wraps liblinear instead, solves the primal directly, and scales far better with both samples and features. For genuinely large or streaming data, `SGDClassifier(loss='hinge')` fits the same linear SVM by stochastic gradient descent, costs one pass per epoch, and supports `partial_fit`. If you need an RBF-like decision boundary at scale, approximate the kernel with `Nystroem` or `RBFSampler` and fit `LinearSVC` on the transformed features. Note the behavioural differences: `LinearSVC` defaults to squared hinge loss, regularizes the intercept, and does multiclass one-vs-rest, while `SVC` fits one-vs-one internally.
code
python · 14 linesfrom sklearn.datasets import make_classification
from sklearn.kernel_approximation import Nystroem
from sklearn.pipeline import make_pipeline
from sklearn.svm import LinearSVC
X, y = make_classification(n_samples=100000, n_features=20, random_state=0)
# RBF-like boundary at linear cost: explicit feature map + linear solver
clf = make_pipeline(
Nystroem(gamma=0.05, n_components=300, random_state=0),
LinearSVC(C=1.0),
).fit(X, y)
print(round(clf.score(X, y), 4))go deeper
Know that SVC struggles on large datasets and that LinearSVC and SGDClassifier are the scalable linear alternatives. A rough row-count threshold is enough at this level.
Explain the mechanism — libsvm solves the dual over pairwise kernel evaluations, so cost grows at least quadratically with samples, while liblinear solves the primal. Name cache_size and what it does and does not fix.
Demonstrate having made the swap: the loss, intercept-regularization and one-vs-one versus one-vs-rest differences that shift metrics, support-vector count as the inference-cost driver, and Nystroem when you still need the nonlinearity.
Own the tradeoff at portfolio level — when a kernel SVM is worth its serving cost versus a linear model or a gradient-boosted tree, and how you keep a prototype that fitted on 20,000 rows from being promoted onto a pipeline that will feed it millions.
## Why SVC stops scaling `SVC` is a Python wrapper over libsvm. It solves the dual formulation of the support-vector problem, which is expressed entirely in terms of kernel evaluations `K(xᵢ, xⱼ)` between pairs of training points. The number of such pairs is quadratic in the sample count, and the SMO-style optimizer visits many of them repeatedly. Scikit-learn's documentation states the practical consequence directly: fit time scales at least quadratically with the number of samples and it is hard to scale beyond a few tens of thousands of samples. `cache_size` (default 200, in MB) controls how much of the kernel matrix is held in memory. Raising it to 1000 or 2000 on a machine with headroom is the cheapest real speedup available for a borderline dataset, because it reduces recomputation of kernel entries. It does not change the asymptotics — a 500,000-row `SVC` fit is not going to finish because you gave it more cache. Prediction cost is proportional to the number of support vectors, available afterwards as `n_support_` and `support_vectors_`. On noisy data with a small `C`, a large fraction of training points end up as support vectors, so inference gets slow too — a detail people miss when they benchmark training only. ## The linear alternatives **`LinearSVC`** wraps liblinear rather than libsvm. Because the kernel is fixed to linear, the problem can be solved in the primal, in terms of the weight vector, so the cost scales roughly with `n_samples * n_features` instead of quadratically. Scikit-learn's docs describe it as having more flexibility in penalties and losses and scaling better to large numbers of samples. It is the right first move whenever a linear boundary is adequate. **`SGDClassifier(loss='hinge')`** fits the same linear SVM objective with stochastic gradient descent: one sample at a time, a fixed number of epochs, cost linear in samples per epoch. It also exposes `partial_fit`, so it can learn from data that never fits in memory at once, and it accepts `loss='log_loss'` if you want a logistic model instead. The catch is that SGD is genuinely sensitive to feature scaling and to the learning-rate schedule; a badly scaled SGD fit can look much worse than the equivalent `LinearSVC` for reasons that have nothing to do with the model class. **Kernel approximation.** If you need nonlinearity but cannot afford `SVC`, `sklearn.kernel_approximation` provides `Nystroem` and `RBFSampler`, which map inputs into an explicit finite-dimensional feature space whose inner products approximate the RBF kernel. Fit `LinearSVC` or `SGDClassifier` on that expansion and you get most of the kernel's expressiveness at linear cost. The `SVC` docstring itself points at this path for large datasets. ## The differences that change results, not just speed Swapping `SVC(kernel='linear')` for `LinearSVC` is not a pure performance change: - **Loss.** `LinearSVC` defaults to `loss='squared_hinge'`; `SVC` uses hinge. Different loss, different fitted boundary. - **Intercept.** liblinear folds the intercept in as a synthetic feature, so it is regularized; `intercept_scaling` exists to blunt that. libsvm's intercept is not penalized. - **Multiclass.** `LinearSVC` does one-vs-rest, fitting one classifier per class. `SVC` fits one-vs-one internally — `n_classes * (n_classes - 1) / 2` binary problems — and then aggregates, with `decision_function_shape='ovr'` controlling only the shape of what `decision_function` hands back, not how the model was fitted. - **Probabilities.** Neither gives probabilities for free; `LinearSVC` has no `predict_proba` at all. - **Dual formulation.** `LinearSVC` has a `dual` parameter, which since scikit-learn 1.5 defaults to `'auto'` and picks the primal or dual solver based on the shape of the data. So a candidate who reports that accuracy shifted after the swap is not necessarily doing something wrong — those are two different estimators that happen to be adjacent. ## Where SVC is still the right answer Small-to-medium datasets — thousands to low tens of thousands of rows — with genuine nonlinearity and a modest feature count. That is a real and common regime, and `SVC(kernel='rbf')` with tuned `C` and `gamma` is a strong baseline there. Two parameter notes: `gamma='scale'` is the default and computes `1 / (n_features * X.var())`, which makes the kernel width adapt to the data's spread; and the whole family is highly sensitive to feature scaling, since the kernel is a distance in input space. ## What the interviewer is listening for The answer that lands is a size threshold plus a mechanism: "past roughly tens of thousands of rows the kernel solver's quadratic-or-worse fit cost dominates, so I move to a primal linear solver, and if I need the nonlinearity I approximate the kernel explicitly first." A candidate who only says "SVMs are slow on big data" has read that somewhere; a candidate who names `LinearSVC`, `SGDClassifier` and `Nystroem` and knows the loss and multiclass differences has actually had to make the swap.
- Does raising SVC's cache_size solve the scaling problem?It helps at the margin and does not change the asymptotics. `cache_size` (200 MB by default) bounds how much of the kernel matrix stays resident; raising it to 1000 or 2000 on a machine with spare RAM cuts recomputation of kernel entries and can be the cheapest real speedup for a borderline dataset. It will not make a several-hundred-thousand-row fit finish.
- Is LinearSVC a drop-in replacement for SVC(kernel='linear')?No. `LinearSVC` defaults to squared hinge loss rather than hinge, regularizes the intercept because liblinear folds it in as a synthetic feature, handles multiclass one-vs-rest where `SVC` fits one-vs-one internally, and has no `predict_proba`. Expect the decision boundary and metrics to shift slightly. It is an adjacent estimator, not the same one running faster.
- SVC prediction is slow in production even though training finished. Why?Inference cost is proportional to the number of support vectors, since each prediction evaluates the kernel against every one of them. Check `n_support_`: on noisy data or with a small `C`, a large share of the training set ends up as support vectors, so the model is effectively carrying most of the training data into serving. Larger `C` or a linear model reduces it.
- How does Nystroem let you keep a nonlinear boundary at large scale?`Nystroem` samples a subset of training points and constructs an explicit finite-dimensional feature map whose inner products approximate the RBF kernel. You transform once, then fit `LinearSVC` or `SGDClassifier` in that space at linear cost. You trade exactness for scalability, tuned by `n_components`. `RBFSampler` offers a random-Fourier-feature alternative to the same end.
saying these in an interview costs you the question
- Says SVMs are simply slow, without the sample-count mechanism
- Expects LinearSVC to reproduce SVC(kernel='linear') exactly
- Thinks cache_size changes the scaling behaviour
- Ignores support-vector count when inference is slow
- Fits SVC on unscaled features and blames the kernel