skip to content

In scikit-learn, is LogisticRegression regularized by default, and what does C control?

level: middleimportance: must knowfreq 74%

answer

  1. Not maximum likelihood out of the box
  2. One knob, and it runs backwards
  3. C multiplies the loss, not the penalty
  4. penalty='l2', C=1.0 by default
  5. Smaller C means stronger shrinkage

basics

~20 s

Scikit-learn's LogisticRegression is regularized by default: penalty='l2' with C=1.0. C is the inverse of regularization strength, so a smaller C shrinks coefficients harder and a larger C fits closer to unpenalized. Pass penalty=None to switch it off.

solid answer

~50 s

Yes — unlike the textbook maximum-likelihood fit, `LogisticRegression()` with no arguments already applies an L2 penalty at `C=1.0`, and nothing warns you. The objective it minimizes is `0.5 * wᵀw + C * Σ log(1 + exp(-yᵢ(xᵢᵀw + c)))`: **C multiplies the data-fit term, not the penalty**, which is why it runs backwards from `Ridge`'s and `Lasso`'s `alpha`. Small C means strong shrinkage; large C approaches an unpenalized fit; `penalty=None` gives it exactly. Two consequences bite in practice. First, because the loss is a sum over samples rather than a mean, the effective regularization weakens as the dataset grows, so a C tuned on a 10k-row sample does not transfer to a million rows. Second, a penalty on raw coefficient magnitude makes the fit sensitive to feature scale, so unscaled features get penalized unevenly. Tune C on a log grid.

code

python · 12 lines
python
import numpy as np
from sklearn.datasets import make_classification
from sklearn.linear_model import LogisticRegression

X, y = make_classification(n_samples=500, n_features=20, random_state=0)

for C in (0.001, 1.0, 1000.0):
    clf = LogisticRegression(C=C, max_iter=1000).fit(X, y)
    print(C, round(float(np.abs(clf.coef_).sum()), 3))

unpenalized = LogisticRegression(penalty=None, max_iter=1000).fit(X, y)
print("none", round(float(np.abs(unpenalized.coef_).sum()), 3))

go deeper

for a junior

Remember two facts you can say in one breath: scikit-learn's LogisticRegression regularizes by default, and C is inverse — smaller C, stronger penalty. Knowing penalty=None exists is enough at this level.

for a middle

Be ready to write the objective and point at where C sits, explain why that makes it inverse, and contrast it with Ridge's and Lasso's alpha. Mention that the penalty makes feature scaling necessary.

for a senior

Show you have been burned by it: coefficients that disagree with a statistical fit, a C that stopped being optimal when the dataset grew because the loss is summed not averaged, and the solver/penalty compatibility matrix you check before tuning.

for a principal

Own the policy question — whether the team reports regularized coefficients as explanations at all, and how C is tuned and re-tuned as data volume grows so a hyperparameter frozen at prototype scale does not quietly under-regularize production.

## The default nobody expects In most statistics courses, logistic regression is fitted by unpenalized maximum likelihood: you maximize the likelihood of the observed labels and read the coefficients as effect sizes. Scikit-learn does not do that. `LogisticRegression()` constructed with no arguments uses `penalty='l2'` and `C=1.0`, so every default fit is a *ridge-penalized* logistic regression. There is no warning, no log line, and the API looks identical to an unpenalized fit. This is the single most common source of "my scikit-learn coefficients don't match R's `glm()`" confusion. ## What C actually multiplies For the default L2 penalty, the class minimizes `0.5 * wᵀw + C * Σᵢ log(exp(-yᵢ(xᵢᵀw + c)) + 1)` Read where `C` sits. It scales the **data-fit** (log-loss) term, while the penalty term `0.5 * wᵀw` carries a fixed coefficient. So: - `C → 0`: the data barely matters relative to the penalty, and coefficients collapse toward zero. In the limit the model predicts the base rate. - `C = 1.0` (default): moderate shrinkage. - `C → ∞`: the penalty becomes negligible and you approach the maximum-likelihood solution. Hence the phrase interviewers want: **C is the inverse of regularization strength**. A subtler consequence of the same formula: the loss is a **sum** over samples, not a mean. Double the number of rows and the data-fit term roughly doubles while the penalty stays put, so the *effective* regularization gets weaker as the dataset grows. A `C` value tuned on a 10,000-row sample is not the same amount of regularization on 1,000,000 rows — retune it when the data volume changes materially. ## Why alpha and C point in opposite directions The linear-model estimators use the other convention. `Ridge` minimizes `||y - Xw||²₂ + alpha * ||w||²₂`, and `Lasso` minimizes `(1 / (2 * n_samples)) * ||y - Xw||²₂ + alpha * ||w||₁`. Here `alpha` multiplies the penalty, so *larger* `alpha` means *more* regularization — the exact opposite of `C`. Two traps follow. First, a candidate who reasons "bigger number, stronger penalty" gets `LogisticRegression` backwards. Second, `alpha` is not even comparable between `Ridge` and `Lasso`, because `Lasso`'s objective divides the squared error by `2 * n_samples` and `Ridge`'s does not. The `C` spelling is inherited from the SVM world — `SVC` and `LinearSVC` use the same `C`, and `LogisticRegression`'s liblinear/libsvm heritage is why the convention leaked into a linear model. ## Penalties and their solvers `penalty` accepts `'l2'` (default), `'l1'`, `'elasticnet'`, and `None`. It is not free-choice, because the penalty and the `solver` must be compatible: - `'l2'` works with every solver: `'lbfgs'` (the default), `'liblinear'`, `'newton-cg'`, `'newton-cholesky'`, `'sag'`, `'saga'`. - `'l1'` requires `'liblinear'` or `'saga'`. - `'elasticnet'` requires `'saga'`, and you must also supply `l1_ratio`. - `penalty=None` is not supported by `'liblinear'`. A mismatch raises a `ValueError` at `fit` time rather than silently ignoring the request — which is the good news. A candidate who says "I set `penalty='l1'` and got no sparsity" either changed the solver, or is looking at coefficients that L1 legitimately kept. ## Practical consequences 1. **Feature scale matters.** The penalty is on raw coefficient magnitude, so a feature measured in millimetres and one measured in kilometres are penalized on wildly different terms. Regularized linear models effectively require standardized inputs. 2. **The intercept.** Under `lbfgs`, `newton-cg`, `newton-cholesky`, `sag` and `saga` the intercept is not penalized. `liblinear` *does* regularize it, which is why `intercept_scaling` exists as a blunting knob for that solver only. 3. **Coefficients are shrunk.** `coef_` from a default fit is a biased, regularized estimate. Do not present it as an unbiased effect size or an odds ratio without saying so. 4. **Convergence.** `max_iter` defaults to 100; on unscaled or ill-conditioned data you will see a `ConvergenceWarning`, and the fix is usually scaling, not simply raising `max_iter`. 5. **Tuning.** Sweep `C` over a log grid (e.g. `10**np.linspace(-4, 4, 9)`). `LogisticRegressionCV` bakes this in through its `Cs` argument if you want the estimator to do it itself. ## The debugging signature "My coefficients are smaller than the statsmodels fit" → the default penalty. "Adding more data changed which C won" → the sum-not-mean objective. "Turning C up made it overfit" → correct, and the candidate who expects the reverse has the convention inverted.

  • Which solvers do you need if you want an L1 or elastic-net penalty here?
    `penalty='l1'` requires `solver='liblinear'` or `solver='saga'`; `penalty='elasticnet'` requires `saga` and an explicit `l1_ratio`. The default `lbfgs`, along with `newton-cg`, `newton-cholesky` and `sag`, only handle `'l2'` or no penalty, and a mismatch raises a `ValueError` at fit time. `liblinear` is the odd one out in the other direction: it cannot take `penalty=None`.
  • You want the coefficients to match an unpenalized statistical fit. What do you change?
    Pass `penalty=None` and use a solver that supports it — `lbfgs`, `newton-cg`, `newton-cholesky`, `sag` or `saga`, but not `liblinear`. Expect slower or shakier convergence on separable or collinear data, since the penalty was also acting as a numerical stabilizer; raise `max_iter` and standardize the features if you hit a `ConvergenceWarning`.
  • Why does the best C shift when you train on ten times more data?
    Because the objective sums the log-loss over samples instead of averaging it, while the penalty term is fixed. More rows means the data-fit term grows relative to `0.5 * wᵀw`, so the same numeric `C` regularizes less. Treat `C` as tied to the training-set size and retune it whenever the volume changes by an order of magnitude.
  • Does the intercept get regularized too?
    Not under `lbfgs`, `newton-cg`, `newton-cholesky`, `sag` or `saga` — the penalty applies to `coef_` only. `liblinear` is the exception: it folds a synthetic constant feature into the design matrix, so the intercept is penalized, and `intercept_scaling` exists to reduce that effect by scaling that synthetic feature up.

saying these in an interview costs you the question

  • Says LogisticRegression is unpenalized unless you ask
  • Thinks larger C means stronger regularization
  • Confuses C with a learning rate or step size
  • Reads default coef_ as an unbiased effect size
  • Sets penalty='l1' without changing the solver

context