skip to content

Squared-Error Regression

A weighted sum of features predicts a continuous target, fit in closed form or iteratively, under squared error or a robust alternative. Interviewers use it as the baseline you must justify leaving.

on this pageshow

explore

questions

12

If you add a squared term to a linear regression, is it still a linear model?

level: juniorimportance: must knowfreq 62%

answer

  1. linear in what, exactly?
  2. the coefficients, not the input
  3. x squared is just another column
  4. same least-squares solve, wider design matrix
  5. contrast: a parameter inside an exponential

basics

~10 s

Yes. Linear here means linear in the coefficients, not in the inputs. A squared term is just another column, so the same least-squares fit still applies, even though the curve in the input bends.

solid answer

~40 s

Yes. In `y = b0 + b1*x + b2*x^2`, the model is still linear in the parameters `b0`, `b1` and `b2` — each one enters as a plain multiplier on a column of known numbers, and the terms are summed. That is the only structure least squares needs, so you fit it by treating `x^2` as an ordinary extra feature. What is nonlinear is the shape in `x`, which is the whole point: crop yield against nitrogen dose rises and then falls, and no straight line can bend like that. A model that is genuinely nonlinear in its parameters looks like `y = b0 * exp(b1*x)`, where `b1` sits inside a nonlinear function; that one needs iterative optimisation, not a single linear solve.

go deeper

for a junior

Be ready to say that linear regression is linear in its coefficients, and that a squared input is simply one more column handed to the same fitting procedure.

for a middle

Explain it through the design matrix: expansion adds columns, the objective and the solve are unchanged, and the strong correlation between an input and its square is what degrades conditioning.

for a senior

Show where the trick stops being enough — when the shape you need has an unknown rate or asymptote inside a nonlinear function, so the fit becomes an iterative search with starting values and convergence checks.

for a principal

Own the tradeoff of staying linear in the parameters: a deterministic fit, readable uncertainty and an auditable form, bought at the cost of specifying the shape by hand instead of letting a flexible learner discover it.

## The word 'linear' describes the parameters A linear regression model has the form `y = b0 + b1*z1 + b2*z2 + ... + bm*zm + error` where `z1 ... zm` are columns of known numbers computed from the data and `b0 ... bm` are the coefficients to be estimated. The word *linear* describes how the coefficients enter: each multiplies one column, and the pieces are added together. It says nothing about how those columns relate to the raw measurement they were built from. So take a single measured input `x` and construct two columns, `z1 = x` and `z2 = x^2`. The model `y = b0 + b1*x + b2*x^2` is a linear model in exactly the technical sense that matters, because once `x^2` has been computed it is just a number sitting in a column. The fitted relationship between `x` and `y` is a parabola — visibly not a straight line — and there is no contradiction. ## The design-matrix view The practical way to see it is to picture the design matrix: one row per observation, one column per term, with a leading column of ones for the intercept. For an observation with `x = 3`, the row of a quadratic model is `[1, 3, 9]`. Squaring happened before the model saw anything. Least squares then chooses the coefficients that minimise the sum of squared residuals, and that minimisation problem has exactly the same structure — a quadratic in the coefficients with a single minimum — whether the columns hold raw measurements, squares, cubes, products of two features, or spline basis functions. This is why the technique is called **basis expansion**. You replace the single input `x` with a set of basis functions evaluated at `x`, say `f1(x), f2(x), ..., fm(x)`, and fit a linear combination of them. Polynomial terms are one choice of basis; piecewise-cubic spline basis functions defined by knots are another; the product of two features is another. The model class stays linear-in-parameters throughout, and everything that follows from that — a convex objective, a unique solution for a full-rank design, closed-form standard errors — is preserved. ## Why anyone bothers Crop yield against nitrogen dose is the standard picture. Yield climbs steeply with the first units of nitrogen, the gain flattens, and past some dose yield actually falls as the excess damages the crop. A model with only a nitrogen column can fit one slope, so it will report either a positive or a negative average effect and be wrong over most of the range. Adding a squared nitrogen column gives the fit the freedom to rise and then turn over, while every property of the fitting procedure is unchanged. ## What nonlinear in the parameters looks like Contrast `y = b0 * exp(b1*x)`. Here `b1` is inside an exponential, so the objective as a function of `b1` is no longer a simple quadratic. There is no single linear solve that returns the answer; you need an iterative optimiser, sensible starting values, and a check that it converged to a good optimum rather than a poor local one. The same is true of a logistic growth curve with an unknown rate and asymptote. The distinction is not academic: it decides whether the fit is a deterministic one-shot computation or a numerical search. ## The costs that come with the extra columns Two caveats belong with the answer, and mentioning them separates a memorised definition from understanding. First, the added columns are usually strongly correlated with the ones they came from. Over a range of positive doses, `x` and `x^2` move together almost in lockstep, which makes the design matrix poorly conditioned: individual coefficients become unstable and hard to read even though the fitted curve is fine. Rescaling the input, or using an orthogonalised polynomial basis, addresses the numerical side of this. Second, every column added is another parameter estimated from the same data. Flexibility bought this way is paid for in the stability of the fit, and the payment grows quickly with the degree. ## The one-line answer Linear regression is linear in its coefficients. Squares, cubes, splines and interaction products are all just columns you compute first, so the model remains a linear model while the relationship it expresses can be as curved as the basis allows.

  • What would make a regression model nonlinear in its parameters?
    A parameter appearing inside a nonlinear function rather than as a plain multiplier — an exponential growth form `b0 * exp(b1*x)`, or a saturating curve with an unknown asymptote and rate. The objective is then no longer a simple quadratic in the coefficients, so there is no one-shot linear solve: you need iterative optimisation, starting values, and a check that the optimiser did not settle in a poor local optimum.
  • Does adding polynomial terms change how the model is fitted?
    No. The objective is still the sum of squared residuals and the solution is still the usual linear solve; the design matrix simply gains columns. What does change is conditioning and stability: the new columns correlate strongly with their parents, so coefficient estimates get noisier even where the fitted curve looks reasonable.
  • Why is the term 'basis expansion' used for this?
    Because you are replacing the raw input with a set of basis functions evaluated at it — `f1(x), f2(x), ...` — and fitting a linear combination of them. Polynomial powers, spline basis functions defined by knots, and products of two features are all choices of basis. The model class does not change; only the columns do.

The recipe is linear in the ingredient amounts even if one ingredient is already a sauce made by cooking another one down. What matters is that the amounts are simply multiplied and added.

saying these in an interview costs you the question

  • Says any curved fit means it is no longer linear regression
  • Claims a different optimiser is needed once a squared term appears
  • Confuses linearity in the parameters with a straight-line relationship
  • Cannot give an example of a model nonlinear in its parameters
  • Thinks squaring the feature changes the fitting objective

context

open as a page

Why does a degree-15 polynomial fitted to 20 observations swing wildly between the data points?

level: middleimportance: must knowfreq 58%

basics

~20 s

Sixteen free coefficients against twenty noisy points leave the fit almost enough freedom to pass through every observation, so it chases noise. Polynomial columns are global and highly correlated, so small data changes swing the curve enormously.

open as a page

How does minimising pinball loss at tau = 0.9 produce a 90th-percentile prediction?

level: middleimportance: must knowfreq 52%

basics

~20 s

Pinball loss at tau = 0.9 charges 0.9 per unit of under-prediction and 0.1 per unit of over-prediction. That nine-to-one asymmetry is minimised by the value the target falls below 90 percent of the time.

open as a page

Why does one extreme target value distort a squared-error regression fit more than an absolute-error fit?

level: middleimportance: must knowfreq 62%

basics

~20 s

Squared error penalises a residual by its square, so its gradient grows with the residual: a point ten times further off pulls ten times harder. Absolute error's gradient has fixed size, so extreme points get no extra vote.

open as a page

In linear regression, why is the fit chosen to minimise squared errors rather than absolute errors?

level: juniorimportance: should knowfreq 58%

basics

~20 s

Squared error is smooth everywhere, so setting its derivative to zero gives one exact formula for the best-fitting weights. Absolute error has a kink at zero, needs iterative fitting, and squaring also makes large misses far more expensive.

open as a page

Why must a linear model be handed an explicit product column to capture a discount-by-loyalty-tier interaction?

level: middleimportance: should knowfreq 50%

basics

~20 s

A linear model is additive: every feature contributes its own amount, and nothing lets one feature change another's contribution. If discount depth works differently for each loyalty tier, that column has to be built by hand.

open as a page

Why does a closed-form least-squares solve scale cubically in features but linearly in rows?

level: middleimportance: should knowfreq 45%

basics

~20 s

Rows are touched once, in a streaming pass building a p-by-p system. Solving that system costs about p^3 and stores p^2 numbers, so total cost is roughly n*p^2 + p^3: wide designs break the exact solve, not tall ones.

open as a page

What does the delta parameter control in Huber loss for a regression fit?

level: middleimportance: should knowfreq 46%

basics

~20 s

Delta is the residual size at which Huber switches from squared to absolute behaviour. Below delta a row's pull grows with its error; above delta the pull is capped at delta, so extreme rows stop dominating the fit.

open as a page

What happens to a closed-form least-squares fit when two features are nearly identical?

level: seniorimportance: should knowfreq 58%

basics

~20 s

Near-duplicate columns leave the cross-product matrix almost singular, so the two coefficients come out enormous, opposite in sign and wildly unstable across refits. Predictions inside the training range stay fine; the individual coefficients are not interpretable.

open as a page

Your demand model's stockouts cost four times what excess inventory does — how do you reflect that in the loss?

level: principalimportance: should knowfreq 34%

basics

~20 s

A four-to-one cost ratio is pinball loss with tau = 4/(4+1) = 0.8, so forecast the 80th percentile of demand rather than the mean. Confirm the ratio is real, linear and stable before baking it into training.

open as a page

Why does a cubic dose term give an absurd prediction for a dose above anything in the training data?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Polynomial terms grow without bound, so once you leave the fitted range the highest power dominates and the prediction diverges. Least squares constrains the curve only where observations existed, and nothing pins down the tails.

open as a page

What can go wrong when you build a P10-P90 band from two separately fitted quantile models?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Two independent fits can cross, putting the P10 above the P90 for some inputs, and their 80 percent width is only nominal. Tail quantiles rest on few effective rows, so verify empirical coverage per segment.

open as a page