skip to content

Parametric vs Non-Parametric

A parametric model fixes its parameter count in advance; a non-parametric one grows with the data, as kNN keeps every point. Interviewers use the split to probe memory and inference cost.

on this pageshow

questions

3

What distinguishes a parametric model from a non-parametric one in machine learning?

level: juniorimportance: must knowfreq 64%

answer

  1. Ask what happens at ten times the data
  2. About capacity, not about zero parameters
  3. Does the trained object grow with n?
  4. Fixed summary versus keeping the rows

basics

~20 s

A parametric model summarises the training data with a fixed number of parameters, decided before fitting, so its size does not change as data grows. A non-parametric model's effective parameter count grows with the training set.

solid answer

~40 s

The line is about capacity, not about whether parameters exist. A parametric model commits to a fixed functional form with a fixed-length parameter vector — a linear or logistic regression with `p` features has `p + 1` coefficients whether you train it on 1,000 rows or 10 million. A non-parametric model lets its complexity grow with the data: k-nearest neighbours keeps the training rows themselves, a decision tree adds splits as more structure becomes supportable, a kernel-based support vector machine retains a subset of training points whose count can grow with `n`. Practically, parametric buys a small fixed memory footprint, fast scoring and stable behaviour on little data, at the price of bias if the true relationship does not match the assumed form. Non-parametric buys flexibility and pays in data, memory and prediction cost.

go deeper

for a junior

Be ready to state the difference in one sentence — fixed parameter count versus a count that grows with the data — and to name one example on each side without hesitating.

for a middle

Expect to be pushed past the definition into borderline cases: classify a decision tree, a size-capped tree, and a model with 500 engineered features, and explain why capacity rather than parameter existence is the criterion.

for a senior

Show that the label maps to consequences you have operated: fixed footprint and fixed scoring cost against a store and a query bill that grow with retained data, plus how much data each family needs before it pays off.

for a principal

Own the framing that this is a capacity-and-cost decision rather than a taxonomy quiz. Be able to argue when a knowingly biased fixed-size model is the right organisational call for the deployment simplicity it buys.

## The word "non-parametric" is misleading The most common mistake is reading "non-parametric" as "has no parameters". Every usable model has quantities that are set by looking at the data. The distinction is whether the **number** of those quantities is fixed in advance, or whether it is allowed to grow as the training set grows. - **Parametric**: you choose a functional form up front, and that choice fixes the size of the parameter vector. Fitting means picking values for a fixed-length vector. - **Non-parametric**: the model's effective complexity is a function of the data. More data can mean more stored examples, more splits, more retained points, more local structure. A compact test: *if I gave this method ten times the data, would the trained object get bigger?* If yes, it is non-parametric. ## Worked examples **Parametric.** Linear regression with `p` features: the fitted object is `p + 1` numbers (`p` coefficients plus an intercept), full stop. Logistic regression: the same, with the coefficients feeding a log-odds. A mixture model with a fixed number of components. In each case the model is a *lossy fixed-size summary* of the training data — once fitted, the training rows can be thrown away and the predictions are unchanged. **Non-parametric.** k-nearest neighbours: nothing is compressed at all; the "model" is the training set. A decision tree: with 200 rows you can only support a handful of splits, with 2 million rows the same growth procedure can support thousands, so the number of leaves — each with its own fitted value — grows with the data. Kernel density estimation: one kernel per observation. Gaussian process regression: the prediction depends on the whole training set through a kernel matrix that grows with `n`. A kernel support vector machine: the decision function is written in terms of retained training points (support vectors), and their count typically grows with the dataset. **Borderline cases are where interviewers push.** A decision tree with an unrestricted growth budget is non-parametric; the *same* algorithm with a hard cap on its size is bounded above and behaves much more like a parametric model, because no amount of extra data can make it bigger. "Non-parametric" is best read as *capacity not bounded in advance*, and capping capacity is exactly what moves a method toward the parametric end. ## Why the distinction earns its keep It predicts four things you care about in practice. 1. **Memory.** A parametric model has a footprint you can state on day one and it never moves. A non-parametric model's footprint is a function of how much data you have kept; the storage question never goes away. 2. **Prediction cost.** A fixed-size model scores a request with a fixed amount of arithmetic. A model that consults stored data does more work as the store grows. 3. **Sample efficiency.** Parametric models trade bias for variance: the assumed form is a strong prior, so they behave sensibly on small samples. Non-parametric models make weak assumptions, so they need enough data to see the shape they are supposed to discover — flexibility is paid for in rows. 4. **Asymptotics.** Given enough data, a well-chosen non-parametric method can approach the true relationship arbitrarily closely, because its capacity keeps growing. A parametric model cannot: if the truth is not in the family you chose, the gap remains no matter how much data arrives. This is *approximation error* or *model bias*, and no dataset size removes it. ## Non-parametric is not assumption-free A second common confusion is that non-parametric means "makes no assumptions". It does not. k-nearest neighbours assumes that points close under your chosen distance have similar targets, and that the distance itself is meaningful — which quietly assumes your features are on comparable scales. Kernel methods assume a notion of similarity. Trees assume the target is well approximated by axis-aligned regions. What is dropped is the commitment to a *global* parametric form, not the assumptions themselves. ## The parametric/non-parametric label is not the same as lazy/eager These are two different axes and interviewers enjoy the crossover. *Lazy* (instance-based) methods defer nearly all work to prediction time — k-nearest neighbours does no meaningful fitting at all. *Eager* methods do the work up front and produce a standalone object. A decision tree is **non-parametric but eager**: it grows with data, yet all the computation happens during fitting and scoring is a cheap walk down the tree. Knowing that a model is non-parametric tells you its capacity story; knowing it is lazy tells you where the compute and memory bill lands. ## How to answer in an interview Give the definition in terms of parameter count versus dataset size, name one example on each side, then immediately convert it into a consequence: fixed footprint and low data need versus unbounded capacity, higher data need and a storage/latency bill. That last move is what separates a memorised definition from an engineer who has shipped both kinds.

  • Is a decision tree parametric or non-parametric?
    Non-parametric. The number of leaves — and therefore the number of fitted values — grows as more data supports more splits, so its capacity is not bounded before you see the data. The caveat is that a tree grown under a hard size cap is bounded above and behaves much more like a parametric model, because extra data can no longer buy extra capacity.
  • Does non-parametric mean the model makes no assumptions?
    No. It means no fixed global functional form, not no assumptions. k-nearest neighbours assumes that points nearby under your distance metric have similar targets, and that the distance is meaningful, which quietly assumes comparable feature scales. Kernel methods assume a similarity function; trees assume axis-aligned regions approximate the target. Weak assumptions still fail when they are wrong.
  • Is logistic regression still parametric if you engineer 500 extra features?
    Yes, for a fixed feature set: you have 501 coefficients before and after training, and that count does not move with the number of rows. It stops being parametric only if the feature construction itself is driven by dataset size — for example one basis function per observation — because then capacity grows with the data.
  • Which family reaches the true relationship given unlimited data?
    The non-parametric one, in principle: its capacity keeps growing, so it can approximate the target arbitrarily well. A parametric model is capped by its family — if the truth is not representable by a straight line in your features, a linear model retains that bias forever. Unlimited data removes variance, never a wrong functional form.

A parametric model is a recipe card: fixed size no matter how many meals you have cooked. A non-parametric model is the photo album of every meal — it keeps growing, and you flip through it each time you cook.

saying these in an interview costs you the question

  • Says non-parametric means the model has no parameters
  • Calls k-nearest neighbours parametric because k is a number you set
  • Claims non-parametric models make no assumptions at all
  • Equates parametric with linear and non-parametric with nonlinear
  • Thinks more parameters automatically means better accuracy

context

open as a page

Why do non-parametric models often lose to simple parametric ones on a 250-row dataset?

level: middleimportance: should knowfreq 44%

basics

~20 s

Flexibility is paid for in rows. With 250 examples a model whose capacity grows with the data builds each local estimate from a handful of points, so it mostly tracks noise. A fixed functional form acts as a stabilising prior.

open as a page

How would you serve a lazy instance-based model under a 10 ms budget as its store keeps growing?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

A lazy model does its work at request time against everything it has stored, so memory and per-query cost both scale with the store. Bound what is stored, or distil the same data into a fixed-size eager model.

open as a page