Linear Algebra for ML
The vector and matrix machinery ML is written in: multiplication, rank, eigendecomposition, SVD, norms and projections. Interviewers probe it to see whether you know what embeddings and PCA do.
on this pageshowhide
explore
- Vectors and Matrix Operations10 questions
- Matrix Multiplication5 questions
- Inverse, Determinant and Trace5 questions
- Rank and Linear Systems10 questions
- Span, Basis and Nullspace5 questions
- Solving Ax = b5 questions
- Eigendecomposition and SVD11 questions
- Eigenvalues and Eigenvectors6 questions
- Singular Value Decomposition5 questions
- Norms, Distance and Projections9 questions
- Vector and Matrix Norms5 questions
- Cosine Similarity and Projections4 questions
questions
page 2 of 2How does the Moore-Penrose pseudoinverse solve a rank-deficient least-squares problem?
basics
~20 sBuild A^+ = V S^+ U^T from the SVD by inverting each nonzero singular value and leaving zeros as zeros. Then x = A^+ b minimizes ||Ax - b|| and, among the infinitely many minimizers, is the one with smallest norm.
When does Mahalanobis distance flag a point that Euclidean distance calls ordinary?
basics
~20 sWhenever features are correlated and a point breaks the correlation. Mahalanobis distance divides out the covariance, so a point that is unremarkable on each feature alone but sits off the joint ridge gets a large distance. Euclidean distance sees nothing unusual.
Why does one vector in R^2 have different coordinates in the standard basis and a rotated basis?
basics
~20 sCoordinates are the weights that rebuild a vector from a chosen basis, so they belong to the basis, not to the vector. Rotate the basis and the arrow is unchanged while the list of weights describing it changes.
Why does squared Euclidean distance violate the triangle inequality?
basics
~20 sSquaring grows faster than adding. For points 0, 1 and 2 on a line, squared distance gives 4 end to end but 1 plus 1 for the legs, so a detour beats the direct route. It is a divergence, not a metric.
Why does Gram-Schmidt lose orthogonality when the input vectors are nearly parallel?
basics
~20 sEach step subtracts projections from a vector, and for nearly parallel inputs that subtraction cancels two almost identical vectors. The tiny survivor is mostly rounding error, and normalising it magnifies that error, so the resulting directions are not truly perpendicular.
How does power iteration find the dominant eigenvector of a large matrix?
basics
~20 sStart from a random nonzero vector, repeatedly multiply by the matrix, and renormalize each step. Components along smaller eigenvalues shrink relative to the largest one, so the vector converges to the dominant eigenvector at rate |lambda2 / lambda1| per iteration.
Why is a tiny determinant not by itself evidence that a matrix is nearly singular?
basics
~10 sBecause the determinant depends on scale: for an n x n matrix, det(cA) = c^n det(A). Shrinking every entry drives the determinant toward zero without making the matrix any harder to invert.
How does associativity of matrix multiplication let you cut the cost of a three-matrix chain?
basics
~20 sAssociativity means (AB)C and A(BC) give identical results, so the grouping is yours to choose. An (m x n) by (n x p) product costs about mnp multiply-adds, so pick the grouping whose intermediate is smallest.
Why solve a least-squares system by QR factorization rather than forming A transpose A?
basics
~20 sForming the cross-product matrix squares the condition number, so you lose twice as many digits and can turn a full-rank problem into a numerically singular one. QR instead factors the original matrix with length-preserving transformations.
How do you pick the truncated-SVD rank k for a users-by-items ratings matrix?
basics
~20 sTreat k as a budgeted tradeoff, not a formula. Let the flattening of the singular-value spectrum bracket a range, then choose inside it by held-out reconstruction error and the downstream metric, subject to storage and serving limits.
showing 31–40 of 40