skip to content

Calculus & Optimization Basics

You will learn gradients, the chain rule, convexity, and why gradient descent finds minima — the calculus that makes 'training a model' a meaningful sentence. Interviewers probe it to see whether you can explain what a loss surface is and what happens when optimization goes wrong.

on this pageshow

explore

questions

page 2 of 2

What is the subgradient set of the convex function f(x) = |x| at the point x = 0?

level: seniorimportance: should knowfreq 36%

basics

~20 s

The subdifferential of |x| at 0 is the closed interval from -1 to 1: every slope of magnitude at most 1 gives a line supporting the V from below. Zero lies inside, which certifies 0 as the minimiser.

open as a page

Why does the loss fall smoothly with full-batch gradient steps but rattle with stochastic ones?

level: seniorimportance: should knowfreq 62%

basics

~20 s

A full-batch step uses the exact gradient, so with a small enough step size the loss decreases every iteration. A stochastic step uses a sampled gradient that is correct only on average, so individual steps can go uphill and the iterates settle into a noise ball rather than a point.

open as a page

When is the local quadratic model f(x0) + g^T d + 0.5 d^T H d a trustworthy stand-in for a loss surface?

level: seniorimportance: should knowfreq 41%

basics

~20 s

Only near the base point, where the loss is smooth and curvature changes slowly. The model's error grows like the cube of the displacement, so it is exact only when the loss is itself quadratic.

open as a page

How does iteratively reweighted least squares fit a logistic regression model?

level: seniorimportance: should knowfreq 34%

basics

~10 s

IRLS runs Newton's method on the log-likelihood. Each iteration computes fitted probabilities, forms weights p(1-p), and solves a weighted least-squares problem on a working response, repeating until the coefficients stop moving.

open as a page

How does L-BFGS make Newton-style optimization affordable for a 10,000-parameter model?

level: seniorimportance: should knowfreq 46%

basics

~10 s

L-BFGS never forms an inverse Hessian. It stores only m recent pairs of parameter and gradient differences and rebuilds the step from them, cutting memory from about d squared entries to m times d.

open as a page

When should a requirement be a hard constraint rather than a penalty in the objective?

level: principalimportance: should knowfreq 32%

basics

~20 s

Use a hard constraint when any violation is unacceptable and a feasible solution is guaranteed to exist. Use a penalty when the requirement is a tradeoff you are willing to price. Constraints give guarantees; penalties keep the problem solvable.

open as a page

What is the Jacobian determinant of the polar map (r, theta) -> (r cos theta, r sin theta)?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

The determinant is r. The Jacobian is [[cos theta, minus r sin theta], [sin theta, r cos theta]], whose determinant is r cos^2 theta plus r sin^2 theta, which equals r. That factor is the map's local area scaling.

open as a page

How does logarithmic differentiation give the derivative of f(x) = x^x?

level: middleimportance: nice to knowfreq 24%

basics

~20 s

Take logs first: ln y = x ln x for y = x^x, which differentiates to y'/y = ln x + 1, giving y' = x^x times (ln x + 1). Base and exponent both vary, so neither standard rule applies.

open as a page

How do you read the curvature of a surface along a direction v from its Hessian H?

level: middleimportance: nice to knowfreq 33%

basics

~10 s

Compute the scalar v^T H v for a unit vector v. That number is the second derivative of the function along the line in direction v: positive bends upward, near zero is locally flat.

open as a page

In nonlinear least squares, why does Gauss-Newton approximate the Hessian by J^T J?

level: middleimportance: nice to knowfreq 28%

basics

~20 s

For a sum of squared residuals the exact Hessian is J^T J plus a term weighted by the residuals themselves. Gauss-Newton drops that second term, so it needs only first derivatives and always gets a positive semidefinite curvature matrix.

open as a page

Why is the pointwise maximum of two linear functions convex despite its kink?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

Convexity is defined by the chord inequality, which never mentions derivatives. Each linear piece satisfies that inequality, and taking the maximum of two bounds preserves it, so the max is convex. The corner where the two lines cross affects smoothness only.

open as a page

How do you distinguish a long flat plateau from a true stationary point of an objective?

level: seniorimportance: nice to knowfreq 33%

basics

~20 s

At a true stationary point the gradient is exactly zero. On a plateau it is merely small and still consistently signed, so the objective keeps creeping down and a long enough step leaves the flat region.

open as a page

What are the gradients of w^T x and of x^T A x with respect to the vector x?

level: seniorimportance: nice to knowfreq 38%

basics

~20 s

The gradient of the linear form w^T x with respect to x is w. The gradient of the quadratic form x^T A x is (A + A^T) x, which becomes 2Ax when A is symmetric.

open as a page

How does the constraint ||w|| <= t relate to minimizing loss plus lambda times ||w||?

level: seniorimportance: nice to knowfreq 28%

basics

~20 s

The penalty weight lambda is the KKT multiplier of the constraint ||w|| <= t. For a convex loss the two forms trace the same family of solutions, with a larger budget t matching a smaller lambda.

open as a page

For a least-squares fit with 10^7 rows and 10^5 features, do you solve in closed form or iterate?

level: principalimportance: nice to knowfreq 40%

basics

~20 s

Iterate. The closed-form route needs a 100,000 by 100,000 cross-product matrix — around 80 GB dense — costing on the order of 10^17 operations to form and 10^15 to factorise, while gradient steps cost one pass over the data each and store only the parameter vector.

open as a page

showing 31–45 of 45