skip to content

Algorithm Choice Heuristics

Which family to reach for given the data you have: a linear baseline first, gradient-boosted trees as the tabular default, kNN only in low dimension. Interviewers use it as a first-model question.

on this pageshow

questions

3

How do you choose a model family for a new tabular supervised learning problem?

level: juniorimportance: must knowfreq 72%

answer

  1. cheapest thing that could work
  2. get a number to beat first
  3. mixed column types, scaling irrelevant
  4. boosted trees are the tabular default
  5. row count and constraints override defaults

basics

~10 s

Start with a regularized linear or logistic baseline for a cheap number to beat, then reach for gradient-boosted trees, the default on mixed numeric-and-categorical tabular data. Data shape and hard constraints override that default.

solid answer

~40 s

My default order on tabular data is baseline first, then boosted trees. A regularized linear or logistic model fits in seconds, needs almost no tuning, and its score becomes the bar every later model has to clear; a suspiciously perfect baseline usually means leakage rather than an easy problem. If the baseline is not good enough, gradient-boosted trees are the standard next step: they take mixed numeric and categorical columns, are unaffected by feature scaling or any monotone transform, pick up interactions and non-linear thresholds I never specified, and shrug off outliers in the inputs. Then I let the situation override the default - very few rows, far more columns than rows, sparse text, or a hard interpretability or latency requirement all push me back toward the linear family.

go deeper

for a junior

Be ready to name your first two reaches out loud - a regularized linear or logistic baseline, then gradient-boosted trees - and to say in one sentence why each is a sensible starting point on tabular data.

for a middle

Explain the properties behind the default: trees handle mixed column types, ignore feature scaling and monotone transforms, and find interactions nobody wrote down. Say what the baseline buys you beyond a score.

for a senior

Show that you let the data shape overrule the heuristic. Talk through the row-count, column-count and sparsity checks you run before choosing, and how an implausibly strong baseline exposes leakage.

for a principal

Own the cost side: a default the whole team applies consistently beats a per-project artisanal choice, and every escalation from baseline to heavy model should be justified by a measured gap rather than by preference.

### The decision you are actually making Choosing a model family is a search-narrowing decision, not a search. On a new tabular supervised problem there are only a handful of families worth considering at all: regularized linear models, tree ensembles, distance-based methods, kernel methods, simple probabilistic classifiers, and neural networks. The useful skill is eliminating most of them in about a minute from the shape of the data and the constraints around the model, so the time you have goes into features, labels and evaluation rather than into a bake-off nobody asked for. ### Step one: the linear or logistic baseline Fit a regularized linear model (linear regression for a continuous target, logistic regression for a binary one) before anything else. It costs seconds, has almost nothing to tune, and pays for itself four ways. - It produces the number every later model must beat. Without it, a boosted model scoring 0.91 AUC is a number with no meaning; with it, you know whether 0.91 is a triumph over 0.72 or a rounding error over 0.90. - It is a leakage detector. A baseline that scores near-perfectly on a problem experts find hard almost always means a column encodes the answer, a row was duplicated across the split, or the label was computed from a feature. - It exercises the whole pipeline end to end, so when a complex model later behaves oddly you already know the splits, the metric and the feature build work. - Its coefficient signs are a cheap sanity check. A feature whose weight points the opposite way from every domain expectation is worth a look before you hide it inside an ensemble. ### Step two: gradient-boosted trees as the tabular default If the baseline is not good enough, the standard next reach on mixed-type tabular data is a gradient-boosted tree ensemble. The reasons are properties of the family, not brand loyalty. - Mixed column types are native. Numeric, ordinal and categorical columns coexist without a common scale, and a heavy categorical such as a postcode with tens of thousands of levels can be split on directly; some boosting algorithms (CatBoost and LightGBM among them) handle categorical splits inside the algorithm itself. - Splits depend only on the order of a feature's values, so scaling, standardising or taking logs of an input changes nothing. Any monotone transform of a feature leaves the fitted model identical. - Interactions and non-linear thresholds come for free. A tree of depth d can express an interaction among d features that you never wrote down, and a step at any cut point captures effects such as risk changing sharply above a certain age. - Outliers in the inputs are harmless: an extreme value is simply on the far side of a split rather than a large number multiplying a weight. - Many boosting algorithms have a defined rule for routing missing values down one branch, so a sparsely populated column does not force an imputation decision up front. What the family cannot do matters just as much. Trees do not extrapolate: outside the range of the training data the prediction is flat, so a model that must forecast beyond observed values is a bad fit. A smooth, near-linear relationship needs many splits to approximate as a staircase, which is one reason the linear baseline sometimes simply wins. ### What overrides the default The heuristic is a starting point, and four situations override it before you look at any score. - Very few rows. With a few hundred rows a boosted ensemble has ample capacity to fit noise, and you do not have enough data to tune it or to measure the difference reliably. - Far more columns than rows. A flexible family cannot be supported by the data, and a strongly regularized linear model is the honest choice. - Wide sparse data such as bag-of-n-gram text, where the signal is spread additively over tens of thousands of mostly-zero columns. - A hard external constraint: a scoring rule a person must be able to read, a strict latency or memory budget at serving time, or a model that must be updated continuously rather than refit. Constraints are filters, not tie-breakers. Apply them first and let accuracy decide only among the families that survive. ### Running the comparison honestly Shortlist two or three families, not seven. Fix the evaluation protocol and the metric before you fit anything, give each candidate a comparable tuning budget, and compare once. The common failure is not picking the wrong family; it is tuning the challenger against the same held-out data that produced the baseline's score, so the winner is decided by how many times you looked.

  • Your dataset is 500 rows from an agricultural yield trial. Does the boosted-tree default still hold?
    No. With 500 rows a boosted ensemble has plenty of capacity to fit noise, and I have too little data both to tune its hyperparameters and to measure the difference reliably - the resampled score is itself noisy. A regularized linear model, or an additive model with a smooth term per factor, is usually within noise of the boosted version and far more stable across resamples. I would escalate only if the baseline residuals show clear interaction structure.
  • A 2-million-row insurance pricing table has a postcode column with 40,000 levels. Which family does that push you toward?
    The tree family. Trees split on groups of categories, so the column can be used close to as it stands, and boosting algorithms with native categorical handling - CatBoost and LightGBM among them - absorb it inside the split search. A linear or distance-based model needs that column turned into numbers before it can see it at all, which makes the encoding, rather than the model, the main design decision.
  • What do you gain from the linear baseline if you already expect boosted trees to win?
    A calibrated expectation and a working pipeline. The baseline tells me what the problem is worth before I spend a day tuning: if it reaches 0.90 AUC, a boosted model reaching 0.91 is not a story. It also exercises the whole path - splits, features, metric - so later surprises are model problems rather than plumbing problems, and its coefficient signs flag any feature pointing the wrong way.

It is the order a mechanic works in: check the cheap, obvious things first, because the answer is often there and because you now know the car runs at all before you start pulling the engine apart.

saying these in an interview costs you the question

  • Jumps straight to the most complex model available
  • Claims one algorithm is always best for tabular data
  • Standardises features before fitting trees, expecting a gain
  • Treats the baseline as a formality rather than a bar to clear
  • Ignores row count when choosing model capacity
  • Lets a constraint like a latency budget decide only after the bake-off

context

open as a page

Which model family suits 20,000 gene-expression columns and only 180 patients?

level: middleimportance: should knowfreq 46%

basics

~20 s

With far more columns than rows, use a strongly regularized linear family - ridge, lasso or elastic net, or a linear SVM - possibly after dimension reduction. Flexible families such as deep trees have far too little data per decision.

open as a page

Why does a linear model usually beat boosted trees on sparse n-gram text features?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

An n-gram matrix has tens of thousands of mostly-zero columns holding many weak additive cues. A linear model sums all their weights at once; a tree tests one feature per split, so it needs thousands of trees.

open as a page