skip to content

Learning Paradigms and Problem Framing

You will learn the map of machine learning: which family of methods fits which problem, and how to frame a business question as classification, regression, ranking, or clustering. Interviewers open with this to check you can scope an ML problem before touching an algorithm.

on this pageshow

explore

questions

page 1 of 2

How do you choose a model family for a new tabular supervised learning problem?

level: juniorimportance: must knowfreq 72%

answer

  1. cheapest thing that could work
  2. get a number to beat first
  3. mixed column types, scaling irrelevant
  4. boosted trees are the tabular default
  5. row count and constraints override defaults

basics

~10 s

Start with a regularized linear or logistic baseline for a cheap number to beat, then reach for gradient-boosted trees, the default on mixed numeric-and-categorical tabular data. Data shape and hard constraints override that default.

solid answer

~40 s

My default order on tabular data is baseline first, then boosted trees. A regularized linear or logistic model fits in seconds, needs almost no tuning, and its score becomes the bar every later model has to clear; a suspiciously perfect baseline usually means leakage rather than an easy problem. If the baseline is not good enough, gradient-boosted trees are the standard next step: they take mixed numeric and categorical columns, are unaffected by feature scaling or any monotone transform, pick up interactions and non-linear thresholds I never specified, and shrug off outliers in the inputs. Then I let the situation override the default - very few rows, far more columns than rows, sparse text, or a hard interpretability or latency requirement all push me back toward the linear family.

go deeper

for a junior

Be ready to name your first two reaches out loud - a regularized linear or logistic baseline, then gradient-boosted trees - and to say in one sentence why each is a sensible starting point on tabular data.

for a middle

Explain the properties behind the default: trees handle mixed column types, ignore feature scaling and monotone transforms, and find interactions nobody wrote down. Say what the baseline buys you beyond a score.

for a senior

Show that you let the data shape overrule the heuristic. Talk through the row-count, column-count and sparsity checks you run before choosing, and how an implausibly strong baseline exposes leakage.

for a principal

Own the cost side: a default the whole team applies consistently beats a per-project artisanal choice, and every escalation from baseline to heavy model should be justified by a measured gap rather than by preference.

### The decision you are actually making Choosing a model family is a search-narrowing decision, not a search. On a new tabular supervised problem there are only a handful of families worth considering at all: regularized linear models, tree ensembles, distance-based methods, kernel methods, simple probabilistic classifiers, and neural networks. The useful skill is eliminating most of them in about a minute from the shape of the data and the constraints around the model, so the time you have goes into features, labels and evaluation rather than into a bake-off nobody asked for. ### Step one: the linear or logistic baseline Fit a regularized linear model (linear regression for a continuous target, logistic regression for a binary one) before anything else. It costs seconds, has almost nothing to tune, and pays for itself four ways. - It produces the number every later model must beat. Without it, a boosted model scoring 0.91 AUC is a number with no meaning; with it, you know whether 0.91 is a triumph over 0.72 or a rounding error over 0.90. - It is a leakage detector. A baseline that scores near-perfectly on a problem experts find hard almost always means a column encodes the answer, a row was duplicated across the split, or the label was computed from a feature. - It exercises the whole pipeline end to end, so when a complex model later behaves oddly you already know the splits, the metric and the feature build work. - Its coefficient signs are a cheap sanity check. A feature whose weight points the opposite way from every domain expectation is worth a look before you hide it inside an ensemble. ### Step two: gradient-boosted trees as the tabular default If the baseline is not good enough, the standard next reach on mixed-type tabular data is a gradient-boosted tree ensemble. The reasons are properties of the family, not brand loyalty. - Mixed column types are native. Numeric, ordinal and categorical columns coexist without a common scale, and a heavy categorical such as a postcode with tens of thousands of levels can be split on directly; some boosting algorithms (CatBoost and LightGBM among them) handle categorical splits inside the algorithm itself. - Splits depend only on the order of a feature's values, so scaling, standardising or taking logs of an input changes nothing. Any monotone transform of a feature leaves the fitted model identical. - Interactions and non-linear thresholds come for free. A tree of depth d can express an interaction among d features that you never wrote down, and a step at any cut point captures effects such as risk changing sharply above a certain age. - Outliers in the inputs are harmless: an extreme value is simply on the far side of a split rather than a large number multiplying a weight. - Many boosting algorithms have a defined rule for routing missing values down one branch, so a sparsely populated column does not force an imputation decision up front. What the family cannot do matters just as much. Trees do not extrapolate: outside the range of the training data the prediction is flat, so a model that must forecast beyond observed values is a bad fit. A smooth, near-linear relationship needs many splits to approximate as a staircase, which is one reason the linear baseline sometimes simply wins. ### What overrides the default The heuristic is a starting point, and four situations override it before you look at any score. - Very few rows. With a few hundred rows a boosted ensemble has ample capacity to fit noise, and you do not have enough data to tune it or to measure the difference reliably. - Far more columns than rows. A flexible family cannot be supported by the data, and a strongly regularized linear model is the honest choice. - Wide sparse data such as bag-of-n-gram text, where the signal is spread additively over tens of thousands of mostly-zero columns. - A hard external constraint: a scoring rule a person must be able to read, a strict latency or memory budget at serving time, or a model that must be updated continuously rather than refit. Constraints are filters, not tie-breakers. Apply them first and let accuracy decide only among the families that survive. ### Running the comparison honestly Shortlist two or three families, not seven. Fix the evaluation protocol and the metric before you fit anything, give each candidate a comparable tuning budget, and compare once. The common failure is not picking the wrong family; it is tuning the challenger against the same held-out data that produced the baseline's score, so the winner is decided by how many times you looked.

  • Your dataset is 500 rows from an agricultural yield trial. Does the boosted-tree default still hold?
    No. With 500 rows a boosted ensemble has plenty of capacity to fit noise, and I have too little data both to tune its hyperparameters and to measure the difference reliably - the resampled score is itself noisy. A regularized linear model, or an additive model with a smooth term per factor, is usually within noise of the boosted version and far more stable across resamples. I would escalate only if the baseline residuals show clear interaction structure.
  • A 2-million-row insurance pricing table has a postcode column with 40,000 levels. Which family does that push you toward?
    The tree family. Trees split on groups of categories, so the column can be used close to as it stands, and boosting algorithms with native categorical handling - CatBoost and LightGBM among them - absorb it inside the split search. A linear or distance-based model needs that column turned into numbers before it can see it at all, which makes the encoding, rather than the model, the main design decision.
  • What do you gain from the linear baseline if you already expect boosted trees to win?
    A calibrated expectation and a working pipeline. The baseline tells me what the problem is worth before I spend a day tuning: if it reaches 0.90 AUC, a boosted model reaching 0.91 is not a story. It also exercises the whole path - splits, features, metric - so later surprises are model problems rather than plumbing problems, and its coefficient signs flag any feature pointing the wrong way.

It is the order a mechanic works in: check the cheap, obvious things first, because the answer is often there and because you now know the car runs at all before you start pulling the engine apart.

saying these in an interview costs you the question

  • Jumps straight to the most complex model available
  • Claims one algorithm is always best for tabular data
  • Standardises features before fitting trees, expecting a gain
  • Treats the baseline as a formality rather than a bar to clear
  • Ignores row count when choosing model capacity
  • Lets a constraint like a latency budget decide only after the bake-off

context

open as a page

Is a 1-to-5 star satisfaction rating a classification or a regression target?

level: juniorimportance: must knowfreq 72%

basics

~20 s

A 1-to-5 star rating is an ordinal target: the values are ordered, but the gaps between them are not guaranteed equal. Regression uses the order and assumes equal spacing; five-class classification keeps the classes but discards the order.

open as a page

How do you turn a raw time series into a supervised table of lag and rolling-window features?

level: juniorimportance: must knowfreq 78%

basics

~20 s

Build one row per time step whose features use only past values — lags such as 1, 7 and 14 steps back, plus rolling means or spreads over trailing windows — and whose label is the value at the forecast horizon.

open as a page

Why is seasonal-naive the first baseline you fit for a 48-hour electricity-load forecast?

level: juniorimportance: must knowfreq 68%

basics

~20 s

Because it sets the bar almost for free. Seasonal-naive predicts each future hour with the load observed at the same hour one week earlier, capturing the daily and weekly cycles without any training; a model that cannot beat it is adding nothing.

open as a page

What is the difference between a generative and a discriminative classifier?

level: juniorimportance: must knowfreq 62%

basics

~20 s

A discriminative classifier models p(y|x) - the label given the features - directly. A generative classifier models the joint p(x,y), normally as a class prior p(y) times a per-class feature distribution p(x|y), and inverts it to score labels.

open as a page

Why can unlabelled data improve a classifier in semi-supervised learning?

level: juniorimportance: must knowfreq 55%

basics

~20 s

Unlabelled data helps only when its shape reveals where the classes separate. Semi-supervised methods assume points in the same dense cluster share a label, so the boundary belongs in a low-density gap. When that assumption is false, unlabelled data hurts.

open as a page

What distinguishes a parametric model from a non-parametric one in machine learning?

level: juniorimportance: must knowfreq 64%

basics

~20 s

A parametric model summarises the training data with a fixed number of parameters, decided before fitting, so its size does not change as data grows. A non-parametric model's effective parameter count grows with the training set.

open as a page

How do supervised and unsupervised learning differ in what the training data must contain?

level: juniorimportance: must knowfreq 86%

basics

~20 s

Supervised learning trains on examples that each carry a known target value and learns to predict it. Unsupervised learning has no target; it describes structure in the inputs alone - groups, directions of variation, or where the data is dense.

open as a page

How does refitting a model in batch differ from updating it incrementally as new data arrives?

level: middleimportance: must knowfreq 58%

basics

~20 s

Batch refitting discards the current model and fits a new one over the whole dataset. Incremental (online) learning keeps the existing parameters and nudges them with each new example or chunk, without revisiting the older rows.

open as a page

An article can carry any of 20 topic tags at once - how should you frame that target?

level: middleimportance: must knowfreq 66%

basics

~20 s

That is a multilabel target: the tags are not mutually exclusive, so an article can carry zero, one or several. Frame it as 20 independent yes/no decisions with a probability per tag, not one distribution over 20 tags summing to one.

open as a page

How does bandit feedback differ from the labels a supervised learner sees?

level: middleimportance: must knowfreq 55%

basics

~20 s

Bandit feedback reveals the outcome only for the action actually taken; a supervised label states the correct answer regardless of what the model did. Showing one of six headlines teaches you nothing about the other five.

open as a page

In self-training, how do a model's own pseudo-labels reinforce its errors?

level: middleimportance: must knowfreq 58%

basics

~20 s

Self-training adds the model's own confident predictions to its training set as ground truth. Some are wrong, and training on them raises confidence in the same mistakes, so more of them clear the threshold each round.

open as a page

When do you pose a triage task as ranking a queue rather than classifying each case?

level: middleimportance: must knowfreq 62%

basics

~20 s

Pose it as ranking when a fixed review capacity, not a decision rule, consumes the output. Only the relative order near the top then changes outcomes, so absolute score values and a class boundary buy nothing.

open as a page

When do you pose anomaly detection as one-class modelling of normal data instead of supervised classification?

level: seniorimportance: must knowfreq 56%

basics

~20 s

Choose one-class modelling when the anomalies you will face are not represented by the anomalies you have labelled, so no boundary learned from past examples transfers. Choose supervised classification when labelled anomalies are plentiful and representative of what comes next.

open as a page

Which model family suits 20,000 gene-expression columns and only 180 patients?

level: middleimportance: should knowfreq 46%

basics

~20 s

With far more columns than rows, use a strongly regularized linear family - ridge, lasso or elastic net, or a linear SVM - possibly after dimension reduction. Flexible families such as deep trees have far too little data per decision.

open as a page

Why does a model that is warm-started and then updated only on recent data forget older patterns?

level: middleimportance: should knowfreq 40%

basics

~20 s

A warm start carries old data in only as a starting point. Once every later update uses recent rows alone, each step moves the parameters further from that start, so an old example's influence decays away.

open as a page

When would you use direct multi-step forecasting instead of feeding predictions back in recursively?

level: middleimportance: should knowfreq 52%

basics

~20 s

Recursive forecasting iterates one one-step model, feeding its predictions back as inputs, so errors compound. Direct forecasting fits a separate model per horizon on real observed history: no compounding, but many models and thinner data.

open as a page

Why can naive Bayes beat logistic regression on a small training set but lose on a large one?

level: middleimportance: should knowfreq 42%

basics

~20 s

Naive Bayes trades bias for variance: its generative assumptions are usually wrong, putting a floor under its error, but its per-class estimates stabilise after very few labels. Logistic regression has a lower floor and needs more data to reach it.

open as a page

How does uncertainty sampling in active learning choose the next rows to label?

level: middleimportance: should knowfreq 45%

basics

~20 s

Uncertainty sampling scores every unlabelled row by how unsure the current model is — one minus the top predicted probability, the top-two gap, or entropy — and buys human labels for the least certain rows, then retrains and repeats.

open as a page

Why do non-parametric models often lose to simple parametric ones on a 250-row dataset?

level: middleimportance: should knowfreq 44%

basics

~20 s

Flexibility is paid for in rows. With 250 examples a model whose capacity grows with the data builds each local estimate from a handful of points, so it mostly tracks noise. A fixed functional form acts as a stabilising prior.

open as a page

How do you judge whether an unsupervised grouping of unlabelled data is any good?

level: middleimportance: should knowfreq 55%

basics

~20 s

An unsupervised grouping has no accuracy score, because there is no ground truth. Judge it on stability under resampling, on whether domain experts recognise the groups, and on whether acting on it improves a measurable downstream outcome.

open as a page

How do you train a model on a 400 GB file that never fits in memory?

level: seniorimportance: should knowfreq 44%

basics

~20 s

Fit out-of-core: stream the file in fixed-size chunks and update a learner that has a per-chunk update rule, so only one chunk sits in memory. First check whether a sample or fewer columns removes the problem.

open as a page

Should you predict delivery ETA in minutes and threshold it, or classify 'late by 30+ minutes' directly?

level: seniorimportance: should knowfreq 52%

basics

~20 s

Model the number when you need the ETA itself or the cutoff may move; classify the flag directly when the cutoff is fixed and only the decision matters. A regressor optimises accuracy everywhere, not around the 30-minute boundary.

open as a page

Your 06:00 taxi-demand forecast lacks the last three hours of counts — how do you frame the task?

level: seniorimportance: should knowfreq 42%

basics

~20 s

Define features against the last landed observation, not the clock. With a three-hour delay the shortest usable lag is four hours, and training rows must be rebuilt at that cutoff so the model never relies on values serving lacks.

open as a page

Why can't a loan-approval model learn anything from the applicants it rejected?

level: seniorimportance: should knowfreq 45%

basics

~20 s

Repayment is observed only for approved applicants, so rejected ones carry no outcome to learn from. The training set is the slice the previous policy approved, and nothing in it says whether a rejection would have repaid.

open as a page

A stakeholder asks for a model that flags 'high performers' but nobody has defined the term — how do you proceed?

level: seniorimportance: should knowfreq 42%

basics

~20 s

An undefined target is a definition problem, not a modelling problem. Find a recorded outcome that can stand in for 'high performer', name the bias that proxy encodes, test whether two reviewers agree, price the labelling, and consider a descriptive deliverable.

open as a page

As reviewed labels accumulate from an anomaly-alert queue, should the task be re-posed as supervised classification?

level: principalimportance: should knowfreq 28%

basics

~20 s

Only with a correction for how those labels were produced. Reviewers judged only what the incumbent detector surfaced, so a classifier trained on them learns to imitate that detector and stays blind wherever it never looked.

open as a page

Hourly bike-dock rental counts are your target - what breaks if you treat them as plain continuous regression?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

A count is a non-negative integer, usually right-skewed with many zeros and a spread that grows with its own level. Plain squared-error regression predicts unbounded numbers and assumes constant spread, so it emits negative counts and lets busy hours dominate.

open as a page

How does graded relevance change a search-ranking task compared with binary relevant-or-not labels?

level: middleimportance: nice to knowfreq 34%

basics

~20 s

Graded labels make the target an ordering among relevant items rather than a set of hits. With perfect, good, marginal and bad grades, ranking a marginal result above a perfect one is an error binary labels cannot express.

open as a page

Why does a linear model usually beat boosted trees on sparse n-gram text features?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

An n-gram matrix has tens of thousands of mostly-zero columns holding many weak additive cues. A linear model sums all their weights at once; a tree tests one feature per split, so it needs thousands of trees.

open as a page

showing 1–30 of 36