skip to content

Linear and Logistic Regression

You will learn the workhorse baselines: the loss each of linear and logistic regression minimises, why the logistic one is a classifier, and how gradient descent fits both when no closed form helps.

on this pageshow

explore

questions

30

If you add a squared term to a linear regression, is it still a linear model?

level: juniorimportance: must knowfreq 62%

answer

  1. linear in what, exactly?
  2. the coefficients, not the input
  3. x squared is just another column
  4. same least-squares solve, wider design matrix
  5. contrast: a parameter inside an exponential

basics

~10 s

Yes. Linear here means linear in the coefficients, not in the inputs. A squared term is just another column, so the same least-squares fit still applies, even though the curve in the input bends.

solid answer

~40 s

Yes. In `y = b0 + b1*x + b2*x^2`, the model is still linear in the parameters `b0`, `b1` and `b2` — each one enters as a plain multiplier on a column of known numbers, and the terms are summed. That is the only structure least squares needs, so you fit it by treating `x^2` as an ordinary extra feature. What is nonlinear is the shape in `x`, which is the whole point: crop yield against nitrogen dose rises and then falls, and no straight line can bend like that. A model that is genuinely nonlinear in its parameters looks like `y = b0 * exp(b1*x)`, where `b1` sits inside a nonlinear function; that one needs iterative optimisation, not a single linear solve.

go deeper

for a junior

Be ready to say that linear regression is linear in its coefficients, and that a squared input is simply one more column handed to the same fitting procedure.

for a middle

Explain it through the design matrix: expansion adds columns, the objective and the solve are unchanged, and the strong correlation between an input and its square is what degrades conditioning.

for a senior

Show where the trick stops being enough — when the shape you need has an unknown rate or asymptote inside a nonlinear function, so the fit becomes an iterative search with starting values and convergence checks.

for a principal

Own the tradeoff of staying linear in the parameters: a deterministic fit, readable uncertainty and an auditable form, bought at the cost of specifying the shape by hand instead of letting a flexible learner discover it.

## The word 'linear' describes the parameters A linear regression model has the form `y = b0 + b1*z1 + b2*z2 + ... + bm*zm + error` where `z1 ... zm` are columns of known numbers computed from the data and `b0 ... bm` are the coefficients to be estimated. The word *linear* describes how the coefficients enter: each multiplies one column, and the pieces are added together. It says nothing about how those columns relate to the raw measurement they were built from. So take a single measured input `x` and construct two columns, `z1 = x` and `z2 = x^2`. The model `y = b0 + b1*x + b2*x^2` is a linear model in exactly the technical sense that matters, because once `x^2` has been computed it is just a number sitting in a column. The fitted relationship between `x` and `y` is a parabola — visibly not a straight line — and there is no contradiction. ## The design-matrix view The practical way to see it is to picture the design matrix: one row per observation, one column per term, with a leading column of ones for the intercept. For an observation with `x = 3`, the row of a quadratic model is `[1, 3, 9]`. Squaring happened before the model saw anything. Least squares then chooses the coefficients that minimise the sum of squared residuals, and that minimisation problem has exactly the same structure — a quadratic in the coefficients with a single minimum — whether the columns hold raw measurements, squares, cubes, products of two features, or spline basis functions. This is why the technique is called **basis expansion**. You replace the single input `x` with a set of basis functions evaluated at `x`, say `f1(x), f2(x), ..., fm(x)`, and fit a linear combination of them. Polynomial terms are one choice of basis; piecewise-cubic spline basis functions defined by knots are another; the product of two features is another. The model class stays linear-in-parameters throughout, and everything that follows from that — a convex objective, a unique solution for a full-rank design, closed-form standard errors — is preserved. ## Why anyone bothers Crop yield against nitrogen dose is the standard picture. Yield climbs steeply with the first units of nitrogen, the gain flattens, and past some dose yield actually falls as the excess damages the crop. A model with only a nitrogen column can fit one slope, so it will report either a positive or a negative average effect and be wrong over most of the range. Adding a squared nitrogen column gives the fit the freedom to rise and then turn over, while every property of the fitting procedure is unchanged. ## What nonlinear in the parameters looks like Contrast `y = b0 * exp(b1*x)`. Here `b1` is inside an exponential, so the objective as a function of `b1` is no longer a simple quadratic. There is no single linear solve that returns the answer; you need an iterative optimiser, sensible starting values, and a check that it converged to a good optimum rather than a poor local one. The same is true of a logistic growth curve with an unknown rate and asymptote. The distinction is not academic: it decides whether the fit is a deterministic one-shot computation or a numerical search. ## The costs that come with the extra columns Two caveats belong with the answer, and mentioning them separates a memorised definition from understanding. First, the added columns are usually strongly correlated with the ones they came from. Over a range of positive doses, `x` and `x^2` move together almost in lockstep, which makes the design matrix poorly conditioned: individual coefficients become unstable and hard to read even though the fitted curve is fine. Rescaling the input, or using an orthogonalised polynomial basis, addresses the numerical side of this. Second, every column added is another parameter estimated from the same data. Flexibility bought this way is paid for in the stability of the fit, and the payment grows quickly with the degree. ## The one-line answer Linear regression is linear in its coefficients. Squares, cubes, splines and interaction products are all just columns you compute first, so the model remains a linear model while the relationship it expresses can be as curved as the basis allows.

  • What would make a regression model nonlinear in its parameters?
    A parameter appearing inside a nonlinear function rather than as a plain multiplier — an exponential growth form `b0 * exp(b1*x)`, or a saturating curve with an unknown asymptote and rate. The objective is then no longer a simple quadratic in the coefficients, so there is no one-shot linear solve: you need iterative optimisation, starting values, and a check that the optimiser did not settle in a poor local optimum.
  • Does adding polynomial terms change how the model is fitted?
    No. The objective is still the sum of squared residuals and the solution is still the usual linear solve; the design matrix simply gains columns. What does change is conditioning and stability: the new columns correlate strongly with their parents, so coefficient estimates get noisier even where the fitted curve looks reasonable.
  • Why is the term 'basis expansion' used for this?
    Because you are replacing the raw input with a set of basis functions evaluated at it — `f1(x), f2(x), ...` — and fitting a linear combination of them. Polynomial powers, spline basis functions defined by knots, and products of two features are all choices of basis. The model class does not change; only the columns do.

The recipe is linear in the ingredient amounts even if one ingredient is already a sauce made by cooking another one down. What matters is that the amounts are simply multiplied and added.

saying these in an interview costs you the question

  • Says any curved fit means it is no longer linear regression
  • Claims a different optimiser is needed once a squared term appears
  • Confuses linearity in the parameters with a straight-line relationship
  • Cannot give an example of a model nonlinear in its parameters
  • Thinks squaring the feature changes the fitting objective

context

open as a page

In gradient-descent fitting, how do full-batch, stochastic and mini-batch updates differ per epoch?

level: juniorimportance: must knowfreq 78%

basics

~20 s

They differ in how many rows each gradient step averages. Full-batch uses all n rows for one update per epoch; stochastic uses one row, giving n updates; mini-batch averages a chunk of B rows, giving about n/B updates.

open as a page

Why does gradient descent on a logistic regression's log loss reach the same fit from any starting weights?

level: juniorimportance: must knowfreq 62%

basics

~20 s

Log loss is convex in a logistic model's weights, so the error surface has one global minimum and no local minima. Every run that converges lands on the same coefficients, which is why random restarts add nothing.

open as a page

In logistic regression, what does the sigmoid do to the linear score w*x + b?

level: juniorimportance: must knowfreq 82%

basics

~20 s

The sigmoid squashes any linear score into the open range 0 to 1, monotonically: large positive scores approach 1, large negative scores approach 0, and a score of 0 maps to 0.5. The result is read as the event probability.

open as a page

Why is a logistic regression's decision boundary a straight line in feature space?

level: juniorimportance: must knowfreq 74%

basics

~20 s

Logistic regression scores each point with one weighted sum. The sigmoid is monotone, so a 0.5 probability cut is exactly the rule score above zero, and the set where the score equals zero is flat: a line, plane or hyperplane.

open as a page

What does the softmax function do to the class scores of a multiclass linear model?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Softmax exponentiates each of the K class scores and divides by their total, turning them into K positive numbers that add up to one. The largest score keeps the largest probability, so the predicted class is unchanged.

open as a page

Why does a degree-15 polynomial fitted to 20 observations swing wildly between the data points?

level: middleimportance: must knowfreq 58%

basics

~20 s

Sixteen free coefficients against twenty noisy points leave the fit almost enough freedom to pass through every observation, so it chases noise. Polynomial columns are global and highly correlated, so small data changes swing the curve enormously.

open as a page

Why do a logistic regression's weights run to infinity when a feature perfectly separates the label?

level: middleimportance: must knowfreq 55%

basics

~20 s

Because a larger weight always lowers the loss. When a feature splits the classes with no overlap, scaling that weight up pushes predictions toward 0 and 1 and log loss toward zero, so no finite maximum-likelihood estimate exists.

open as a page

Why is logistic regression trained with cross-entropy loss rather than squared error?

level: middleimportance: must knowfreq 76%

basics

~20 s

Cross-entropy is the Bernoulli negative log-likelihood, and its gradient on the linear score is simply predicted minus actual. Squared error through a sigmoid is non-convex in the weights and its gradient nearly vanishes exactly where the model is confidently wrong.

open as a page

Why can a linear classifier not separate an XOR pattern over two binary features?

level: middleimportance: must knowfreq 61%

basics

~20 s

A linear classifier cuts feature space with one flat boundary and labels each side uniformly. XOR puts its two positive cells on one diagonal and its two negative cells on the other, and crossing diagonals cannot be split by a straight line.

open as a page

How does minimising pinball loss at tau = 0.9 produce a 90th-percentile prediction?

level: middleimportance: must knowfreq 52%

basics

~20 s

Pinball loss at tau = 0.9 charges 0.9 per unit of under-prediction and 0.1 per unit of over-prediction. That nine-to-one asymmetry is minimised by the value the target falls below 90 percent of the time.

open as a page

Why does one extreme target value distort a squared-error regression fit more than an absolute-error fit?

level: middleimportance: must knowfreq 62%

basics

~20 s

Squared error penalises a residual by its square, so its gradient grows with the residual: a point ten times further off pulls ten times harder. Absolute error's gradient has fixed size, so extreme points get no extra vote.

open as a page

In linear regression, why is the fit chosen to minimise squared errors rather than absolute errors?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Squared error is smooth everywhere, so setting its derivative to zero gives one exact formula for the best-fitting weights. Absolute error has a kink at zero, needs iterative fitting, and squaring also makes large misses far more expensive.

open as a page

Why must a linear model be handed an explicit product column to capture a discount-by-loyalty-tier interaction?

level: middleimportance: should knowfreq 50%

basics

~20 s

A linear model is additive: every feature contributes its own amount, and nothing lets one feature change another's contribution. If discount depth works differently for each loyalty tier, that column has to be built by hand.

open as a page

Why shuffle the training rows before each epoch of stochastic or mini-batch gradient descent?

level: middleimportance: should knowfreq 45%

basics

~20 s

Shuffling makes each batch a random sample of the training set, so its averaged gradient fairly estimates the full-data gradient. Without it, a sorted file gives correlated batches, a sawtooth loss curve, and weights biased toward the pass's last rows.

open as a page

Why does a closed-form least-squares solve scale cubically in features but linearly in rows?

level: middleimportance: should knowfreq 45%

basics

~20 s

Rows are touched once, in a streaming pass building a p-by-p system. Solving that system costs about p^3 and stores p^2 numbers, so total cost is roughly n*p^2 + p^3: wide designs break the exact solve, not tall ones.

open as a page

What does the delta parameter control in Huber loss for a regression fit?

level: middleimportance: should knowfreq 46%

basics

~20 s

Delta is the residual size at which Huber switches from squared to absolute behaviour. Below delta a row's pull grows with its error; above delta the pull is capped at delta, so extreme rows stop dominating the fit.

open as a page

How do you read a multiclass log loss of 1.2 from a 20-language classifier?

level: middleimportance: should knowfreq 44%

basics

~20 s

Multiclass log loss averages the negative log of the probability given to the true class. Uniform guessing over 20 languages scores log 20, about 3.0, so 1.2 beats chance and means an average true-class probability near 0.30.

open as a page

Would you fit one multinomial softmax model or 40 one-vs-rest models for 40 categories?

level: middleimportance: should knowfreq 52%

basics

~20 s

One multinomial fit trains all 40 classes jointly through a shared normaliser, so its probabilities already sum to one. Forty one-vs-rest fits are independent, so their scores need renormalising, while one-vs-one would need 780 models.

open as a page

Your training loss becomes NaN by epoch three of a gradient-descent fit — how do you diagnose it?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Log the loss per update, not per epoch. A loss climbing geometrically with weights flipping sign means the step size is too large: cut the learning rate tenfold. A NaN on the very first update points to bad inputs or a broken loss.

open as a page

Your logistic fit's loss drops 1e-9 per epoch while the weight norm doubles — what is happening?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Those two symptoms together are the signature of separated data, not slow learning. The loss is creeping toward zero along a direction with no finite optimum, so the weights grow without bound and extra training will not help.

open as a page

A clinic no-show model weights 'previous no-shows' negatively - how do you investigate before launch?

level: seniorimportance: should knowfreq 44%

basics

~20 s

A wrong-sign weight tilts the decision boundary against domain sense, so treat it as a suspected defect: check how the label and the feature are coded and computed, whether a correlated predictor absorbs the effect, and whether the sign is stable.

open as a page

What happens to a closed-form least-squares fit when two features are nearly identical?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Near-duplicate columns leave the cross-product matrix almost singular, so the two coefficients come out enormous, opposite in sign and wildly unstable across refits. Predictions inside the training range stay fine; the individual coefficients are not interpretable.

open as a page

Your demand model's stockouts cost four times what excess inventory does — how do you reflect that in the loss?

level: principalimportance: should knowfreq 34%

basics

~20 s

A four-to-one cost ratio is pinball loss with tau = 4/(4+1) = 0.8, so forecast the 80th percentile of demand rather than the mean. Confirm the ratio is real, linear and stable before baking it into training.

open as a page

Why is squared-error loss on a logistic model's sigmoid output not convex in the weights?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Squaring the gap between a label and a sigmoid composes a bowl with an S-curve, producing flat saturated regions and stationary points that are not the global best. Cross-entropy instead collapses to a convex function of the linear score.

open as a page

Why does a cubic dose term give an absurd prediction for a dose above anything in the training data?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Polynomial terms grow without bound, so once you leave the fitted range the highest power dominates and the prediction diverges. Least squares constrains the curve only where observations existed, and nothing pins down the tails.

open as a page

What stopping criteria end a gradient-descent fit when you watch only the training loss?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

Three, used together: the full-data gradient norm falling below a relative tolerance, the relative improvement in the epoch-average loss falling below something like 1e-6, and a maximum-epoch budget as a backstop. Record which one fired.

open as a page

Your logistic uptake model's scores will be reused as propensity scores - what changes in how you fit it?

level: seniorimportance: nice to knowfreq 31%

basics

~20 s

The label becomes the treatment indicator and the probability level matters more than the ranking. Keep confounders even when they add no accuracy, exclude anything measured after treatment, and do not rebalance the classes or over-penalise the weights.

open as a page

What can go wrong when you build a P10-P90 band from two separately fitted quantile models?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Two independent fits can cross, putting the P10 above the P90 for some inputs, and their 80 percent width is only nominal. Tail quantiles rest on few effective rows, so verify empirical coverage per segment.

open as a page

Why does adding 5 to every logit of a softmax classifier leave its probabilities unchanged?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Softmax divides by the sum of the exponentials, so a constant added to every score contributes the same factor to numerator and denominator and cancels. Only differences between class scores are identified, which leaves softmax over-parameterised.

open as a page