skip to content

Model Families and Regimes

How a method stores what it learned and when it updates it - a fixed parameter vector, the training set itself, a joint distribution - and which family that argues for on your data.

on this pageshow

explore

questions

13

How do you choose a model family for a new tabular supervised learning problem?

level: juniorimportance: must knowfreq 72%

answer

  1. cheapest thing that could work
  2. get a number to beat first
  3. mixed column types, scaling irrelevant
  4. boosted trees are the tabular default
  5. row count and constraints override defaults

basics

~10 s

Start with a regularized linear or logistic baseline for a cheap number to beat, then reach for gradient-boosted trees, the default on mixed numeric-and-categorical tabular data. Data shape and hard constraints override that default.

solid answer

~40 s

My default order on tabular data is baseline first, then boosted trees. A regularized linear or logistic model fits in seconds, needs almost no tuning, and its score becomes the bar every later model has to clear; a suspiciously perfect baseline usually means leakage rather than an easy problem. If the baseline is not good enough, gradient-boosted trees are the standard next step: they take mixed numeric and categorical columns, are unaffected by feature scaling or any monotone transform, pick up interactions and non-linear thresholds I never specified, and shrug off outliers in the inputs. Then I let the situation override the default - very few rows, far more columns than rows, sparse text, or a hard interpretability or latency requirement all push me back toward the linear family.

go deeper

for a junior

Be ready to name your first two reaches out loud - a regularized linear or logistic baseline, then gradient-boosted trees - and to say in one sentence why each is a sensible starting point on tabular data.

for a middle

Explain the properties behind the default: trees handle mixed column types, ignore feature scaling and monotone transforms, and find interactions nobody wrote down. Say what the baseline buys you beyond a score.

for a senior

Show that you let the data shape overrule the heuristic. Talk through the row-count, column-count and sparsity checks you run before choosing, and how an implausibly strong baseline exposes leakage.

for a principal

Own the cost side: a default the whole team applies consistently beats a per-project artisanal choice, and every escalation from baseline to heavy model should be justified by a measured gap rather than by preference.

### The decision you are actually making Choosing a model family is a search-narrowing decision, not a search. On a new tabular supervised problem there are only a handful of families worth considering at all: regularized linear models, tree ensembles, distance-based methods, kernel methods, simple probabilistic classifiers, and neural networks. The useful skill is eliminating most of them in about a minute from the shape of the data and the constraints around the model, so the time you have goes into features, labels and evaluation rather than into a bake-off nobody asked for. ### Step one: the linear or logistic baseline Fit a regularized linear model (linear regression for a continuous target, logistic regression for a binary one) before anything else. It costs seconds, has almost nothing to tune, and pays for itself four ways. - It produces the number every later model must beat. Without it, a boosted model scoring 0.91 AUC is a number with no meaning; with it, you know whether 0.91 is a triumph over 0.72 or a rounding error over 0.90. - It is a leakage detector. A baseline that scores near-perfectly on a problem experts find hard almost always means a column encodes the answer, a row was duplicated across the split, or the label was computed from a feature. - It exercises the whole pipeline end to end, so when a complex model later behaves oddly you already know the splits, the metric and the feature build work. - Its coefficient signs are a cheap sanity check. A feature whose weight points the opposite way from every domain expectation is worth a look before you hide it inside an ensemble. ### Step two: gradient-boosted trees as the tabular default If the baseline is not good enough, the standard next reach on mixed-type tabular data is a gradient-boosted tree ensemble. The reasons are properties of the family, not brand loyalty. - Mixed column types are native. Numeric, ordinal and categorical columns coexist without a common scale, and a heavy categorical such as a postcode with tens of thousands of levels can be split on directly; some boosting algorithms (CatBoost and LightGBM among them) handle categorical splits inside the algorithm itself. - Splits depend only on the order of a feature's values, so scaling, standardising or taking logs of an input changes nothing. Any monotone transform of a feature leaves the fitted model identical. - Interactions and non-linear thresholds come for free. A tree of depth d can express an interaction among d features that you never wrote down, and a step at any cut point captures effects such as risk changing sharply above a certain age. - Outliers in the inputs are harmless: an extreme value is simply on the far side of a split rather than a large number multiplying a weight. - Many boosting algorithms have a defined rule for routing missing values down one branch, so a sparsely populated column does not force an imputation decision up front. What the family cannot do matters just as much. Trees do not extrapolate: outside the range of the training data the prediction is flat, so a model that must forecast beyond observed values is a bad fit. A smooth, near-linear relationship needs many splits to approximate as a staircase, which is one reason the linear baseline sometimes simply wins. ### What overrides the default The heuristic is a starting point, and four situations override it before you look at any score. - Very few rows. With a few hundred rows a boosted ensemble has ample capacity to fit noise, and you do not have enough data to tune it or to measure the difference reliably. - Far more columns than rows. A flexible family cannot be supported by the data, and a strongly regularized linear model is the honest choice. - Wide sparse data such as bag-of-n-gram text, where the signal is spread additively over tens of thousands of mostly-zero columns. - A hard external constraint: a scoring rule a person must be able to read, a strict latency or memory budget at serving time, or a model that must be updated continuously rather than refit. Constraints are filters, not tie-breakers. Apply them first and let accuracy decide only among the families that survive. ### Running the comparison honestly Shortlist two or three families, not seven. Fix the evaluation protocol and the metric before you fit anything, give each candidate a comparable tuning budget, and compare once. The common failure is not picking the wrong family; it is tuning the challenger against the same held-out data that produced the baseline's score, so the winner is decided by how many times you looked.

  • Your dataset is 500 rows from an agricultural yield trial. Does the boosted-tree default still hold?
    No. With 500 rows a boosted ensemble has plenty of capacity to fit noise, and I have too little data both to tune its hyperparameters and to measure the difference reliably - the resampled score is itself noisy. A regularized linear model, or an additive model with a smooth term per factor, is usually within noise of the boosted version and far more stable across resamples. I would escalate only if the baseline residuals show clear interaction structure.
  • A 2-million-row insurance pricing table has a postcode column with 40,000 levels. Which family does that push you toward?
    The tree family. Trees split on groups of categories, so the column can be used close to as it stands, and boosting algorithms with native categorical handling - CatBoost and LightGBM among them - absorb it inside the split search. A linear or distance-based model needs that column turned into numbers before it can see it at all, which makes the encoding, rather than the model, the main design decision.
  • What do you gain from the linear baseline if you already expect boosted trees to win?
    A calibrated expectation and a working pipeline. The baseline tells me what the problem is worth before I spend a day tuning: if it reaches 0.90 AUC, a boosted model reaching 0.91 is not a story. It also exercises the whole path - splits, features, metric - so later surprises are model problems rather than plumbing problems, and its coefficient signs flag any feature pointing the wrong way.

It is the order a mechanic works in: check the cheap, obvious things first, because the answer is often there and because you now know the car runs at all before you start pulling the engine apart.

saying these in an interview costs you the question

  • Jumps straight to the most complex model available
  • Claims one algorithm is always best for tabular data
  • Standardises features before fitting trees, expecting a gain
  • Treats the baseline as a formality rather than a bar to clear
  • Ignores row count when choosing model capacity
  • Lets a constraint like a latency budget decide only after the bake-off

context

open as a page

What is the difference between a generative and a discriminative classifier?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A discriminative classifier models p(y|x) - the label given the features - directly. A generative classifier models the joint p(x,y), normally as a class prior p(y) times a per-class feature distribution p(x|y), and inverts it to score labels.

open as a page

What distinguishes a parametric model from a non-parametric one in machine learning?

level: juniorimportance: must knowfreq 64%

basics

~20 s

A parametric model summarises the training data with a fixed number of parameters, decided before fitting, so its size does not change as data grows. A non-parametric model's effective parameter count grows with the training set.

open as a page

How does refitting a model in batch differ from updating it incrementally as new data arrives?

level: middleimportance: must knowfreq 58%

basics

~20 s

Batch refitting discards the current model and fits a new one over the whole dataset. Incremental (online) learning keeps the existing parameters and nudges them with each new example or chunk, without revisiting the older rows.

open as a page

Which model family suits 20,000 gene-expression columns and only 180 patients?

level: middleimportance: should knowfreq 46%

basics

~20 s

With far more columns than rows, use a strongly regularized linear family - ridge, lasso or elastic net, or a linear SVM - possibly after dimension reduction. Flexible families such as deep trees have far too little data per decision.

open as a page

Why does a model that is warm-started and then updated only on recent data forget older patterns?

level: middleimportance: should knowfreq 40%

basics

~20 s

A warm start carries old data in only as a starting point. Once every later update uses recent rows alone, each step moves the parameters further from that start, so an old example's influence decays away.

open as a page

Why can naive Bayes beat logistic regression on a small training set but lose on a large one?

level: middleimportance: should knowfreq 42%

basics

~20 s

Naive Bayes trades bias for variance: its generative assumptions are usually wrong, putting a floor under its error, but its per-class estimates stabilise after very few labels. Logistic regression has a lower floor and needs more data to reach it.

open as a page

Why do non-parametric models often lose to simple parametric ones on a 250-row dataset?

level: middleimportance: should knowfreq 44%

basics

~20 s

Flexibility is paid for in rows. With 250 examples a model whose capacity grows with the data builds each local estimate from a handful of points, so it mostly tracks noise. A fixed functional form acts as a stabilising prior.

open as a page

How do you train a model on a 400 GB file that never fits in memory?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Fit out-of-core: stream the file in fixed-size chunks and update a learner that has a per-chunk update rule, so only one chunk sits in memory. First check whether a sample or fewer columns removes the problem.

open as a page

Why does a linear model usually beat boosted trees on sparse n-gram text features?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

An n-gram matrix has tens of thousands of mostly-zero columns holding many weak additive cues. A linear model sums all their weights at once; a tree tests one feature per split, so it needs thousands of trees.

open as a page

What can a model of the joint p(x, y) do that a model of p(y|x) alone cannot?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Modelling the joint gives you a distribution over the features themselves, so you can draw synthetic records, integrate out a feature missing at scoring time, and judge how unusual an input is. A p(y|x) model answers only the labelling question.

open as a page

How would you serve a lazy instance-based model under a 10 ms budget as its store keeps growing?

level: seniorimportance: nice to knowfreq 36%

basics

~20 s

A lazy model does its work at request time against everything it has stored, so memory and per-query cost both scale with the store. Bound what is stored, or distil the same data into a fixed-size eager model.

open as a page

When is a deliberately frozen model the right choice over one that keeps updating itself?

level: principalimportance: nice to knowfreq 26%

basics

~20 s

Freeze the model when validating a new version costs more than staleness does, when a past decision must be explainable against a specific artifact, or when its outputs shape future labels. Freezing trades silent decay for reviewable behaviour.

open as a page