skip to content

Machine Learning Fundamentals

You will learn how models learn from data, how to keep them from overfitting, how to evaluate them honestly, and when to reach for which classical algorithm. This is the 'explain bias-variance', 'precision or recall here?' core of every DS/MLE screen, probed independent of any framework.

on this pageshow

explore

questions

533 · 13 sections

How do you choose a model family for a new tabular supervised learning problem?

level: juniorimportance: must knowfreq 72%
basics
~10 s

Start with a regularized linear or logistic baseline for a cheap number to beat, then reach for gradient-boosted trees, the default on mixed numeric-and-categorical tabular data. Data shape and hard constraints override that default.

open as a page

Is a 1-to-5 star satisfaction rating a classification or a regression target?

level: juniorimportance: must knowfreq 72%
basics
~20 s

A 1-to-5 star rating is an ordinal target: the values are ordered, but the gaps between them are not guaranteed equal. Regression uses the order and assumes equal spacing; five-class classification keeps the classes but discards the order.

open as a page

How do you turn a raw time series into a supervised table of lag and rolling-window features?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Build one row per time step whose features use only past values — lags such as 1, 7 and 14 steps back, plus rolling means or spreads over trailing windows — and whose label is the value at the forecast horizon.

open as a page

Why is seasonal-naive the first baseline you fit for a 48-hour electricity-load forecast?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Because it sets the bar almost for free. Seasonal-naive predicts each future hour with the load observed at the same hour one week earlier, capturing the daily and weekly cycles without any training; a model that cannot beat it is adding nothing.

open as a page

What is the difference between a generative and a discriminative classifier?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A discriminative classifier models p(y|x) - the label given the features - directly. A generative classifier models the joint p(x,y), normally as a class prior p(y) times a per-class feature distribution p(x|y), and inverts it to score labels.

open as a page

What is model capacity, and why do underfitting and overfitting sit at its two ends?

level: juniorimportance: must knowfreq 88%
basics
~20 s

Capacity is how large and flexible a model's space of possible fitted functions is. Too little capacity and the model cannot represent the real pattern, which is underfitting. Too much and it reproduces the training sample's noise, which is overfitting.

open as a page

Why does a held-out score only estimate future performance if the data is i.i.d.?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A held-out score is an average of errors over a sample of rows. It estimates future error only if future rows are independent draws from that same distribution. When the population or the relationship moves, the number stops applying.

open as a page

In the bias-variance decomposition of prediction error, what do bias and variance mean?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Bias is the gap between the true value and the average prediction a model makes when it is retrained on different training samples. Variance is how much that prediction moves from one training sample to the next.

open as a page

Why is a model's error on its own training rows an optimistically biased estimate of future error?

level: juniorimportance: must knowfreq 85%
basics
~20 s

Training error is optimistic because the fitting procedure tuned the model to those exact rows, absorbing their random noise as if it were signal. On fresh rows the noise does not repeat, so error rises.

open as a page

What does a validation curve plot, and how do you read it to pick a hyperparameter value?

level: juniorimportance: must knowfreq 62%
basics
~20 s

A validation curve plots training score and held-out validation score against one hyperparameter that controls model capacity. Pick the value where validation peaks; a widening gap between the two lines past that point is the overfitting signal.

open as a page

In a lasso coefficient path drawn with lambda decreasing left to right, what do the two ends show?

level: juniorimportance: must knowfreq 55%
basics
~20 s

At the far left the penalty is largest and every coefficient is exactly zero, so the model predicts a constant. Moving right, predictors enter one by one, and the far right is the unpenalised least-squares fit.

open as a page

What is early stopping in an iteratively fitted model, and why does it act as regularization?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Early stopping halts an iterative fit once a held-out validation score stops improving, and keeps the best-scoring iterate. Each extra iteration lets the model absorb finer detail from the training rows, so stopping sooner limits effective capacity and cuts variance.

open as a page

What penalty does elastic net add to a linear model, and why mix L1 with L2?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Elastic net penalises the sum of absolute coefficients (L1) and the sum of squared coefficients (L2) together. The L1 part drives weak predictors to exactly zero; the L2 part keeps correlated predictors together and makes the fit stable.

open as a page

In ridge regression, what happens to the coefficients as lambda goes to 0 and as lambda grows very large?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Ridge with lambda = 0 reproduces the ordinary least squares fit. As lambda grows, every slope coefficient shrinks smoothly toward zero and the model flattens toward a constant, trading variance for bias. No coefficient reaches exactly zero at finite lambda.

open as a page

Why must predictors be standardised before fitting a ridge or lasso model?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An L1 or L2 penalty adds up coefficient sizes, and a coefficient's size depends on the units of its predictor. Standardising puts every predictor on a common spread so one lambda penalises them all comparably.

open as a page

How do you check whether a classifier's predicted probabilities are calibrated?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Bucket the predictions by score, then compare each bucket's mean predicted probability with the fraction of rows in it that were actually positive. Plotted, that is a reliability diagram: a calibrated model sits on the diagonal.

open as a page

In a binary classifier's confusion matrix, what are the four cells and how is accuracy computed?

level: juniorimportance: must knowfreq 84%
basics
~20 s

The four cells count true positives, false positives, true negatives and false negatives - one per pairing of predicted and actual label. Accuracy is (TP + TN) divided by all four counts; the error rate is 1 minus accuracy.

open as a page

Why is 99.7% accuracy uninformative for a defect detector when 0.3% of units are defective?

level: juniorimportance: must knowfreq 88%
basics
~20 s

Because a model that never predicts 'defective' already scores 99.7% on that data. The majority class supplies almost every row, so accuracy measures performance on good units and says nothing about whether a single defect was ever caught.

open as a page

How do you read per-class precision and recall off a 10-class confusion matrix?

level: juniorimportance: must knowfreq 62%
basics
~20 s

With true classes as rows and predicted classes as columns, a class's recall is its diagonal cell divided by its row total, and its precision is that same diagonal cell divided by its column total.

open as a page

What is the difference between precision and recall, and why do they trade off?

level: juniorimportance: must knowfreq 88%
basics
~20 s

Precision is the share of predicted positives that are truly positive; recall is the share of actual positives the model catches. Loosening the decision rule labels more examples positive, which raises recall and usually lowers precision.

open as a page

How does a bootstrap resample estimate a model's prediction error, and which rows are held out?

level: juniorimportance: must knowfreq 54%
basics
~20 s

Draw n rows with replacement from the sample of n and train on that draw. Roughly 37% of the original rows are never drawn; score the model on those. Repeat a few hundred times and average the errors.

open as a page

Why shuffle the rows before splitting them into k-fold cross-validation folds?

level: juniorimportance: must knowfreq 64%
basics
~20 s

Exported files usually arrive sorted — by label, by date, by customer. Cutting such a file into contiguous blocks gives folds that are not representative, sometimes single-class, so every fold score misleads. Shuffling row order first breaks that ordering.

open as a page

Why is it wrong to fit a scaler on the full dataset before splitting into train and test?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Fitting a scaler on all rows lets the test rows' mean and standard deviation shape the transform, so the test set is no longer unseen and the score is optimistic. Fit on training rows only, then apply it to test.

open as a page

Five-fold CV returns accuracies 0.73, 0.78, 0.81, 0.84 and 0.89 - what do you report?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Report the mean, 0.81, together with the spread: the fold standard deviation is about 0.06, giving a standard error of roughly 0.03. Write it as 0.81 plus or minus 0.03 over five folds, never as the best fold's 0.89.

open as a page

What does stratified k-fold cross-validation preserve, and when does plain k-fold fail?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Stratified k-fold gives every fold roughly the same class proportions as the whole dataset. Plain k-fold fails when a class is rare: some folds get very few or zero minority rows, so fold scores swing or go undefined.

open as a page

If you add a squared term to a linear regression, is it still a linear model?

level: juniorimportance: must knowfreq 62%
basics
~10 s

Yes. Linear here means linear in the coefficients, not in the inputs. A squared term is just another column, so the same least-squares fit still applies, even though the curve in the input bends.

open as a page

In gradient-descent fitting, how do full-batch, stochastic and mini-batch updates differ per epoch?

level: juniorimportance: must knowfreq 78%
basics
~20 s

They differ in how many rows each gradient step averages. Full-batch uses all n rows for one update per epoch; stochastic uses one row, giving n updates; mini-batch averages a chunk of B rows, giving about n/B updates.

open as a page

Why does gradient descent on a logistic regression's log loss reach the same fit from any starting weights?

level: juniorimportance: must knowfreq 62%
basics
~20 s

Log loss is convex in a logistic model's weights, so the error surface has one global minimum and no local minima. Every run that converges lands on the same coefficients, which is why random restarts add nothing.

open as a page

In logistic regression, what does the sigmoid do to the linear score w*x + b?

level: juniorimportance: must knowfreq 82%
basics
~20 s

The sigmoid squashes any linear score into the open range 0 to 1, monotonically: large positive scores approach 1, large negative scores approach 0, and a score of 0 maps to 0.5. The result is read as the event probability.

open as a page

Why is a logistic regression's decision boundary a straight line in feature space?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Logistic regression scores each point with one weighted sum. The sigmoid is monotone, so a 0.5 probability cut is exactly the rule score above zero, and the set where the score equals zero is flat: a line, plane or hyperplane.

open as a page

In k-NN, why must features be standardised before any distance is computed?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Distance adds up per-feature gaps, so a feature measured in large units, like income in dollars, swamps one measured in small units, like age in years. Standardising puts every feature on a comparable scale so each can contribute.

open as a page

What does naive Bayes' conditional independence assumption claim about the features?

level: juniorimportance: must knowfreq 80%
basics
~20 s

Naive Bayes assumes that once the class label is known, the features carry no further information about each other. That licenses replacing the whole likelihood P(features | class) with a product of one-feature terms P(x_i | class).

open as a page

How does Laplace smoothing stop one unseen word from zeroing a class score in naive Bayes?

level: juniorimportance: must knowfreq 76%
basics
~20 s

Laplace smoothing adds one pseudo-count to every word-class pair before dividing, so no estimated likelihood is exactly zero. Without it, a word never seen in a class drives that class's whole product to zero whatever the other words say.

open as a page

Why is a hard-margin SVM unchanged when you delete a training point that is not a support vector?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Only the training points lying exactly on the margin boundary - the support vectors - hold the hyperplane in place. Every other point sits strictly further away, so removing it leaves the same widest street and the same fitted boundary.

open as a page

Why must a 40-class SVM product router be built from many binary SVMs?

level: juniorimportance: must knowfreq 58%
basics
~20 s

A support vector machine optimises one hyperplane, which has exactly two sides, so its objective can only express a positive and a negative class. Forty categories are covered by training one SVM per class pair: 780 models.

open as a page

Why does a decision tree approximate a diagonal decision boundary with a staircase?

level: juniorimportance: must knowfreq 72%
basics
~20 s

Every split tests one feature against one threshold, so each cut is a line perpendicular to one axis. A boundary that depends on two features jointly can only be tiled by many small axis-parallel steps.

open as a page

What is bagging, and why does averaging bootstrap-trained models cut variance but not bias?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Bagging trains many copies of one model on bootstrap resamples of the training rows, then averages their predictions. Averaging cancels the errors that differ from copy to copy, cutting variance; the systematic error every copy shares survives, so bias stays.

open as a page

How does gradient boosting build a regression model stage by stage?

level: juniorimportance: must knowfreq 78%
basics
~10 s

Gradient boosting starts from a single constant prediction, then repeatedly fits a shallow tree to the errors the model still makes and adds that tree to the running total. Earlier trees are never refitted.

open as a page

What is the difference between leaf-wise and level-wise tree growth in gradient boosting?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Level-wise growth splits every leaf at the current depth before going deeper, producing balanced trees. Leaf-wise growth repeatedly splits whichever leaf anywhere in the tree promises the largest loss reduction, producing deeper, lopsided trees that fit harder and overfit sooner.

open as a page

Which growth-stopping limits keep a decision tree from splitting until every leaf is pure?

level: juniorimportance: must knowfreq 74%
basics
~20 s

Four pre-pruning limits: a maximum tree depth, a minimum node size to split, a minimum number of samples per resulting leaf, and a minimum impurity decrease per split. Any split breaking one of them is refused.

open as a page

In DBSCAN, what makes a point a core, border, or noise point?

level: juniorimportance: must knowfreq 68%
basics
~20 s

DBSCAN calls a point core when at least minPts points lie within radius eps of it. A non-core point that sits inside some core point's eps-radius is a border point. Anything else is noise and gets no cluster.

open as a page

In a Gaussian mixture model, what does it mean for a point to have responsibilities 0.6 and 0.4?

level: juniorimportance: must knowfreq 64%
basics
~20 s

The point is split softly between the two components: given the fitted model, there is a 60 percent posterior chance the first component generated it and 40 percent the second. It is an ambiguity score, not a label.

open as a page

What does each iteration of Lloyd's algorithm for k-means do?

level: juniorimportance: must knowfreq 84%
basics
~10 s

Each Lloyd iteration does two things: assign every point to its nearest centroid, then move each centroid to the mean of the points assigned to it. The loop repeats until assignments stop changing.

open as a page

What does a dendrogram show, and how does agglomerative clustering build one?

level: juniorimportance: must knowfreq 68%
basics
~20 s

A dendrogram is a tree recording the order and distance at which clusters merge. Agglomerative clustering starts with every observation as its own cluster and repeatedly merges the two closest ones, drawing each merge at the distance it happened.

open as a page

What does PCA maximise when it picks the first principal component?

level: juniorimportance: must knowfreq 78%
basics
~20 s

PCA picks the first principal component as the unit-length direction in the mean-centred data along which the projected points have the largest variance. Each later component maximises the remaining variance while staying orthogonal to all earlier ones.

open as a page

For an electricity-demand model, which features would you extract from a raw timestamp column?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Split the timestamp into the calendar parts demand depends on: hour of day, day of week, an is-weekend flag, month or quarter, and a public-holiday flag. The raw timestamp alone only lets a model learn a trend.

open as a page

How do you turn a raw event log into one row per customer for a churn model?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Choose the entity and a reference date, keep only that entity's events inside a window ending at the reference date, then collapse them into one row: counts, sums, means, maxima, distinct counts, and days since the last event.

open as a page

20% of survey respondents skipped the income question — do you drop those rows, drop the column, or impute?

level: juniorimportance: must knowfreq 82%
basics
~20 s

Impute in most cases. Deleting the rows discards a fifth of the data and has no equivalent at scoring time, where every request must return a prediction. Drop the column only if it is nearly empty or adds no lift.

open as a page

Why is coding education level as 0-3 acceptable but coding job family the same way risky?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Education level has a real order - high-school below bachelor below master below PhD - so integer codes carry that information. Job family has no order, so codes 0-3 invent a ranking the model takes literally. Encode unordered categories as one-hot indicators instead.

open as a page

What is the difference between z-score standardisation and min-max scaling?

level: juniorimportance: must knowfreq 84%
basics
~20 s

Z-score standardisation subtracts the mean and divides by the standard deviation, producing mean 0, standard deviation 1, and no fixed bounds. Min-max scaling rescales values into a fixed range such as 0 to 1 using the training minimum and maximum.

open as a page

Why does removing race and gender columns from the training data fail to make a model fair?

level: juniorimportance: must knowfreq 70%
basics
~20 s

Other features reconstruct the removed attribute. ZIP code, school, job history and purchase patterns correlate with it, so the model relearns it indirectly. Dropping the column removes your ability to measure the disparity, not the disparity itself.

open as a page

A partial dependence curve of predicted electricity demand against outdoor temperature is U-shaped — what does that tell you?

level: juniorimportance: must knowfreq 42%
basics
~20 s

The model has learned a non-monotone relationship: average predicted demand is highest at cold temperatures and at hot temperatures, and lowest in between. The curve summarises the model's average behaviour over the whole dataset, not any single building.

open as a page

What makes a model interpretable by design rather than explained after the fact?

level: juniorimportance: must knowfreq 62%
basics
~20 s

An interpretable-by-design model is one whose fitted structure is itself the explanation: a sparse weighted sum, a depth-3 tree, a short rule list. Post-hoc methods leave the model opaque and build a separate, approximate account of it.

open as a page

How is permutation importance computed for a single feature, and what does the number mean?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Permutation importance shuffles one column's values across held-out rows, breaking that column's link to the label, then re-scores the already-trained model. The drop from the baseline score, averaged over several shuffles, is the feature's importance.

open as a page

How do demographic parity and equalized odds differ as group fairness criteria?

level: middleimportance: must knowfreq 60%
basics
~20 s

Demographic parity requires the same positive-prediction rate in every group, ignoring the true outcome. Equalized odds requires equal true-positive and false-positive rates in every group, comparing errors only among people who share a true label.

open as a page

What do support, confidence and lift measure for a basket association rule?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Support is the share of all baskets containing the rule's items. Confidence is the share of baskets holding the left-hand side that also hold the right-hand side. Lift divides confidence by the right-hand side's own support.

open as a page

How does a content-based recommender score an item using only that user's own history?

level: juniorimportance: must knowfreq 70%
basics
~20 s

It describes each item by its own features, builds a profile vector from the features of items that user liked, and ranks candidates by similarity between profile and item vector. No other user's data is involved.

open as a page

How do implicit feedback signals like plays and skips differ from explicit star ratings?

level: juniorimportance: must knowfreq 78%
basics
~20 s

Explicit ratings are stated preferences on a scale and can express dislike. Implicit signals only record actions, so they are abundant and cheap but noisy, one-class, and count-valued: no interaction is not a stated dislike.

open as a page

An install-probability ranker's predicted numbers are wildly off but ordered correctly — does ranking suffer?

level: juniorimportance: must knowfreq 50%
basics
~20 s

No. A ranked list depends only on the order the scores induce, and any strictly increasing transform of every score leaves that order identical. The wrong numbers cost you elsewhere: thresholds, expected-value arithmetic, anything reading the score as a probability.

open as a page

How does matrix factorization predict a missing rating in a users-by-films matrix?

level: juniorimportance: must knowfreq 68%
basics
~20 s

Matrix factorization learns a short vector of latent factors for every user and every film, then predicts a rating as the dot product of the two vectors plus bias terms. Every empty cell can be filled that way.

open as a page

Why does an epsilon-greedy reinforcement learning agent deliberately take non-greedy actions?

level: juniorimportance: must knowfreq 78%
basics
~20 s

An agent's action values are only estimates. With probability epsilon it acts randomly so it keeps sampling actions its estimates currently undervalue, which corrects an unlucky early estimate instead of locking onto the first action that looked good.

open as a page

What are the five components of a Markov decision process?

level: juniorimportance: must knowfreq 72%
basics
~20 s

An MDP is written as states, actions, a transition function giving the probability of the next state, a reward function, and a discount factor gamma. Together they define a sequential decision problem an agent solves by choosing actions.

open as a page

Value iteration needs a known model — what exactly must that model give you?

level: juniorimportance: must knowfreq 50%
basics
~20 s

For every state and action, the model must give the probability of each possible next state and the expected reward. Value iteration averages over all successors in every update, so it plans inside the model rather than learning from experience.

open as a page

Why do policy gradient methods suit continuous action spaces better than value-based control?

level: juniorimportance: must knowfreq 66%
basics
~20 s

Policy gradient methods tune the policy's parameters directly, so an action is sampled from a distribution. Value-based control has to maximise a value estimate over all actions at every step, which is intractable when the action is a real-valued vector.

open as a page

What is the difference between V(s) and Q(s,a) in reinforcement learning?

level: juniorimportance: must knowfreq 80%
basics
~20 s

V(s) is the expected return from state s when the agent follows its policy from there on. Q(s,a) is the expected return from taking action a in state s first, then following the same policy.

open as a page