Calculus & Optimization Basics
You will learn gradients, the chain rule, convexity, and why gradient descent finds minima — the calculus that makes 'training a model' a meaningful sentence. Interviewers probe it to see whether you can explain what a loss surface is and what happens when optimization goes wrong.
on this pageshowhide
explore
- Derivatives & Gradients11 questions
- Derivative Rules6 questions
- Gradient Vector5 questions
- Chain Rule & Curvature9 questions
- Chain Rule & Jacobians5 questions
- Hessians & Taylor Expansion4 questions
- Convexity & Optimality10 questions
- Convex Sets & Functions5 questions
- Critical Points & Saddles5 questions
- Descent & Constraints15 questions
- Gradient Descent Dynamics5 questions
- Lagrange Multipliers & KKT5 questions
- Newton & Quasi-Newton Methods5 questions
questions
page 2 of 2What is the subgradient set of the convex function f(x) = |x| at the point x = 0?
basics
~20 sThe subdifferential of |x| at 0 is the closed interval from -1 to 1: every slope of magnitude at most 1 gives a line supporting the V from below. Zero lies inside, which certifies 0 as the minimiser.
Why does the loss fall smoothly with full-batch gradient steps but rattle with stochastic ones?
basics
~20 sA full-batch step uses the exact gradient, so with a small enough step size the loss decreases every iteration. A stochastic step uses a sampled gradient that is correct only on average, so individual steps can go uphill and the iterates settle into a noise ball rather than a point.
When is the local quadratic model f(x0) + g^T d + 0.5 d^T H d a trustworthy stand-in for a loss surface?
basics
~20 sOnly near the base point, where the loss is smooth and curvature changes slowly. The model's error grows like the cube of the displacement, so it is exact only when the loss is itself quadratic.
How does iteratively reweighted least squares fit a logistic regression model?
basics
~10 sIRLS runs Newton's method on the log-likelihood. Each iteration computes fitted probabilities, forms weights p(1-p), and solves a weighted least-squares problem on a working response, repeating until the coefficients stop moving.
How does L-BFGS make Newton-style optimization affordable for a 10,000-parameter model?
basics
~10 sL-BFGS never forms an inverse Hessian. It stores only m recent pairs of parameter and gradient differences and rebuilds the step from them, cutting memory from about d squared entries to m times d.
When should a requirement be a hard constraint rather than a penalty in the objective?
basics
~20 sUse a hard constraint when any violation is unacceptable and a feasible solution is guaranteed to exist. Use a penalty when the requirement is a tradeoff you are willing to price. Constraints give guarantees; penalties keep the problem solvable.
What is the Jacobian determinant of the polar map (r, theta) -> (r cos theta, r sin theta)?
basics
~20 sThe determinant is r. The Jacobian is [[cos theta, minus r sin theta], [sin theta, r cos theta]], whose determinant is r cos^2 theta plus r sin^2 theta, which equals r. That factor is the map's local area scaling.
How does logarithmic differentiation give the derivative of f(x) = x^x?
basics
~20 sTake logs first: ln y = x ln x for y = x^x, which differentiates to y'/y = ln x + 1, giving y' = x^x times (ln x + 1). Base and exponent both vary, so neither standard rule applies.
How do you read the curvature of a surface along a direction v from its Hessian H?
basics
~10 sCompute the scalar v^T H v for a unit vector v. That number is the second derivative of the function along the line in direction v: positive bends upward, near zero is locally flat.
In nonlinear least squares, why does Gauss-Newton approximate the Hessian by J^T J?
basics
~20 sFor a sum of squared residuals the exact Hessian is J^T J plus a term weighted by the residuals themselves. Gauss-Newton drops that second term, so it needs only first derivatives and always gets a positive semidefinite curvature matrix.
Why is the pointwise maximum of two linear functions convex despite its kink?
basics
~20 sConvexity is defined by the chord inequality, which never mentions derivatives. Each linear piece satisfies that inequality, and taking the maximum of two bounds preserves it, so the max is convex. The corner where the two lines cross affects smoothness only.
How do you distinguish a long flat plateau from a true stationary point of an objective?
basics
~20 sAt a true stationary point the gradient is exactly zero. On a plateau it is merely small and still consistently signed, so the objective keeps creeping down and a long enough step leaves the flat region.
What are the gradients of w^T x and of x^T A x with respect to the vector x?
basics
~20 sThe gradient of the linear form w^T x with respect to x is w. The gradient of the quadratic form x^T A x is (A + A^T) x, which becomes 2Ax when A is symmetric.
How does the constraint ||w|| <= t relate to minimizing loss plus lambda times ||w||?
basics
~20 sThe penalty weight lambda is the KKT multiplier of the constraint ||w|| <= t. For a convex loss the two forms trace the same family of solutions, with a larger budget t matching a smaller lambda.
For a least-squares fit with 10^7 rows and 10^5 features, do you solve in closed form or iterate?
basics
~20 sIterate. The closed-form route needs a 100,000 by 100,000 cross-product matrix — around 80 GB dense — costing on the order of 10^17 operations to form and 10^15 to factorise, while gradient steps cost one pass over the data each and store only the parameter vector.
showing 31–45 of 45