skip to content

Linear Algebra for ML

The vector and matrix machinery ML is written in: multiplication, rank, eigendecomposition, SVD, norms and projections. Interviewers probe it to see whether you know what embeddings and PCA do.

on this pageshow

explore

questions

page 2 of 2

How does the Moore-Penrose pseudoinverse solve a rank-deficient least-squares problem?

level: seniorimportance: should knowfreq 40%

basics

~20 s

Build A^+ = V S^+ U^T from the SVD by inverting each nonzero singular value and leaving zeros as zeros. Then x = A^+ b minimizes ||Ax - b|| and, among the infinitely many minimizers, is the one with smallest norm.

open as a page

When does Mahalanobis distance flag a point that Euclidean distance calls ordinary?

level: seniorimportance: should knowfreq 38%

basics

~20 s

Whenever features are correlated and a point breaks the correlation. Mahalanobis distance divides out the covariance, so a point that is unremarkable on each feature alone but sits off the joint ridge gets a large distance. Euclidean distance sees nothing unusual.

open as a page

Why does one vector in R^2 have different coordinates in the standard basis and a rotated basis?

level: middleimportance: nice to knowfreq 32%

basics

~20 s

Coordinates are the weights that rebuild a vector from a chosen basis, so they belong to the basis, not to the vector. Rotate the basis and the arrow is unchanged while the list of weights describing it changes.

open as a page

Why does squared Euclidean distance violate the triangle inequality?

level: middleimportance: nice to knowfreq 26%

basics

~20 s

Squaring grows faster than adding. For points 0, 1 and 2 on a line, squared distance gives 4 end to end but 1 plus 1 for the legs, so a detour beats the direct route. It is a divergence, not a metric.

open as a page

Why does Gram-Schmidt lose orthogonality when the input vectors are nearly parallel?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Each step subtracts projections from a vector, and for nearly parallel inputs that subtraction cancels two almost identical vectors. The tiny survivor is mostly rounding error, and normalising it magnifies that error, so the resulting directions are not truly perpendicular.

open as a page

How does power iteration find the dominant eigenvector of a large matrix?

level: seniorimportance: nice to knowfreq 30%

basics

~20 s

Start from a random nonzero vector, repeatedly multiply by the matrix, and renormalize each step. Components along smaller eigenvalues shrink relative to the largest one, so the vector converges to the dominant eigenvector at rate |lambda2 / lambda1| per iteration.

open as a page

Why is a tiny determinant not by itself evidence that a matrix is nearly singular?

level: seniorimportance: nice to knowfreq 28%

basics

~10 s

Because the determinant depends on scale: for an n x n matrix, det(cA) = c^n det(A). Shrinking every entry drives the determinant toward zero without making the matrix any harder to invert.

open as a page

How does associativity of matrix multiplication let you cut the cost of a three-matrix chain?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

Associativity means (AB)C and A(BC) give identical results, so the grouping is yours to choose. An (m x n) by (n x p) product costs about mnp multiply-adds, so pick the grouping whose intermediate is smallest.

open as a page

Why solve a least-squares system by QR factorization rather than forming A transpose A?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Forming the cross-product matrix squares the condition number, so you lose twice as many digits and can turn a full-rank problem into a numerically singular one. QR instead factors the original matrix with length-preserving transformations.

open as a page

How do you pick the truncated-SVD rank k for a users-by-items ratings matrix?

level: principalimportance: nice to knowfreq 28%

basics

~20 s

Treat k as a budgeted tradeoff, not a formula. Let the flattening of the singular-value spectrum bracket a range, then choose inside it by held-out reconstruction error and the downstream metric, subject to storage and serving limits.

open as a page

showing 31–40 of 40