skip to content

Why does SVC.predict_proba in scikit-learn require probability=True, and what does it cost?

level: middleimportance: should knowfreq 40%

answer

  1. A margin is not a probability
  2. The flag changes fit, not just predict
  3. Sigmoid fitted by internal cross-validation
  4. Fit time multiplies, seed now matters
  5. Argmax may contradict the predicted label

basics

~20 s

An SVM produces a signed distance to the hyperplane, not a probability. Setting probability=True makes fit run Platt scaling — a sigmoid fitted by internal 5-fold cross-validation — which sharply increases fit time and can yield probabilities whose argmax disagrees with predict.

solid answer

~50 s

The support-vector machine has no probabilistic output: `decision_function` returns a signed margin, and nothing in the objective constrains it to look like a probability. Scikit-learn therefore hides `predict_proba` behind `probability=False`, and calling it raises rather than guessing. Setting `probability=True` makes `fit` do extra work: it runs an internal 5-fold cross-validation and fits a sigmoid (Platt scaling) mapping margins to probabilities. That multiplies fit time on an estimator that was already the expensive one, makes results depend on `random_state` because of the internal shuffling, and — as the docs warn — can produce probabilities whose argmax disagrees with what `predict` returns, since `predict` uses the margin directly. If you only need ranking or ROC-AUC, use `decision_function` and skip it. If you need genuinely calibrated probabilities, wrapping the estimator in `CalibratedClassifierCV` gives you the same idea with control over the calibration method and folds.

code

python · 10 lines
python
from sklearn.datasets import make_classification
from sklearn.svm import SVC

X, y = make_classification(n_samples=400, n_features=10, flip_y=0.25, random_state=0)

clf = SVC(probability=True, random_state=0).fit(X, y)

by_predict = clf.predict(X)
by_proba = clf.classes_[clf.predict_proba(X).argmax(axis=1)]
print("disagreements:", int((by_predict != by_proba).sum()))

go deeper

for a junior

Recall that SVC gives you predict_proba only when constructed with probability=True, and that decision_function is the raw score available otherwise. Knowing the flag exists is the baseline here.

for a middle

Explain that the flag changes fit, not predict: Platt scaling fits a sigmoid on an internal 5-fold cross-validation, which multiplies fit time and makes random_state relevant.

for a senior

Bring the failure story — probabilities whose argmax contradicts predict near the boundary — and name the alternatives you would reach for: decision_function for ranking, CalibratedClassifierCV when calibration is a genuine requirement.

for a principal

Own whether the system needs calibrated probabilities at all. If downstream expected-value logic consumes them, that is an architectural requirement that argues for a probabilistic model or an explicit calibration stage, not a constructor flag.

## The model has no probability to report A support-vector machine is fitted by maximizing a margin, not by maximizing a likelihood. What comes out of the fitted model is `decision_function(X)`: a signed distance from the separating hyperplane, positive on one side and negative on the other, in units that depend on the data scaling and on `C`. There is nothing in the training objective that makes that number a probability, and no reason a margin of 2.0 should mean 88% confidence rather than 71%. Scikit-learn handles this honestly. `SVC` constructs with `probability=False`, and calling `predict_proba` on such an estimator raises an `AttributeError` rather than inventing a number. Many candidates meet this parameter for the first time as an exception message. ## What probability=True actually does Setting `probability=True` changes what happens inside `fit`. In addition to fitting the SVM, scikit-learn (via libsvm) performs Platt scaling: it fits a sigmoid `1 / (1 + exp(A * f(x) + B))` that maps decision-function values to probabilities, estimating `A` and `B` on **internal 5-fold cross-validation** of the training data. Three consequences follow directly: 1. **Fit gets much slower.** You are now fitting the SVM several times over, on an estimator whose cost already scales at least quadratically with the number of samples. On a borderline-size dataset this is the difference between a fit that finishes and one that does not. 2. **Results become stochastic.** The internal cross-validation shuffles, so `random_state` now affects `predict_proba` output. The docs note this explicitly. Two fits with the same data and no seed can produce slightly different probabilities. 3. **`predict` and `predict_proba` can disagree.** This is the one that bites. `predict` classifies by the decision function (via the one-vs-one vote), while `predict_proba` runs the separately-fitted sigmoid. Near the boundary, `predict_proba(X).argmax(axis=1)` is not guaranteed to equal `predict(X)`. Scikit-learn's own documentation calls this out as a known inconsistency of the method. Code that computes probabilities, thresholds them at 0.5, and expects to reproduce `predict` will occasionally be wrong, and the discrepancy is small enough to survive testing and surface in production. ## What to do instead, per use case **Ranking, ROC-AUC, precision-recall curves.** These need only a monotone score, and `decision_function` is exactly that. Scikit-learn's scoring machinery already knows this: `scoring='roc_auc'` uses `decision_function` when `predict_proba` is unavailable. Turning on `probability=True` to compute AUC is paying a large fit cost for nothing, and Platt scaling is monotone anyway so the AUC barely moves. **A tunable operating point.** A threshold on `decision_function` works as well as a threshold on a probability. Pick it on validation data. **Genuinely calibrated probabilities**, e.g. because a downstream expected-value calculation consumes them. Wrap the estimator: `CalibratedClassifierCV(SVC(), method='sigmoid', cv=5)` reproduces Platt scaling with your choice of folds, or `method='isotonic'` fits a nonparametric monotone map that is more flexible but needs more data. The wrapper's advantage over the built-in flag is control and consistency — the calibrated wrapper's `predict` is derived from its own probabilities, so the two cannot disagree. ## Related estimators `LinearSVC` has no `probability` parameter at all and no `predict_proba` — it only exposes `decision_function`. If you need probabilities from a linear SVM, `CalibratedClassifierCV` is the route, or you switch to `LogisticRegression`, which is probabilistic by construction and usually the better answer if probabilities are a first-class requirement. ## The diagnostic story The interview-worthy version of this is a bug report, not a definition: "our confidence scores occasionally contradicted the predicted label." The cause is that `predict` and `predict_proba` are, on `SVC`, two different computations stitched together — and the fix is either to derive the label from the probabilities yourself (`argmax`), or to use a properly calibrated wrapper where that inconsistency cannot arise. A candidate who reaches that explanation has used the estimator; one who says "you just set `probability=True`" has read the signature.

  • You only need ROC-AUC. Do you still need probability=True?
    No. ROC-AUC needs a monotone score, and `decision_function` is one. Scikit-learn's scoring machinery falls back to `decision_function` when `predict_proba` is unavailable, so `scoring='roc_auc'` works on a default `SVC`. Platt scaling is monotone anyway, so it would barely change the AUC while multiplying fit time.
  • How would you get calibrated probabilities from a LinearSVC, which has no probability parameter?
    Wrap it: `CalibratedClassifierCV(LinearSVC(), method='sigmoid', cv=5)` fits the base estimator per fold and calibrates its `decision_function` output, giving you `predict_proba`. `method='isotonic'` is the nonparametric alternative when you have enough data. The wrapper also derives `predict` from its own probabilities, so the two cannot contradict each other.
  • Why do two SVC fits with probability=True on identical data give slightly different probabilities?
    Because the Platt sigmoid is estimated on an internal 5-fold cross-validation that shuffles the training data. Without a fixed `random_state`, the folds differ between runs, so the fitted sigmoid parameters differ. Set `random_state` for reproducibility — note that on an `SVC` this parameter only matters at all once `probability=True`.
  • If the probabilities and predict can disagree, which one should downstream code trust?
    Pick one and derive everything from it. If probabilities drive the decision, take `predict_proba(X).argmax(axis=1)` as the label rather than calling `predict` separately. If the label is what matters, use `predict` and expose `decision_function` as the confidence score. Mixing the two is what produces the contradictory outputs in the first place.

saying these in an interview costs you the question

  • Treats the decision function value as a probability
  • Believes probability=True only changes prediction, not fit
  • Assumes predict always matches predict_proba's argmax
  • Enables probability=True just to compute ROC-AUC
  • Expects LinearSVC to expose predict_proba

context