What are the gradients of w^T x and of x^T A x with respect to the vector x?
answer
- expand into index notation and differentiate one entry
- a linear form has a constant gradient
- x appears twice in the quadratic form
- one contribution from each index position
- the two collapse into one factor when A^T = A
basics
~20 sThe gradient of the linear form w^T x with respect to x is w. The gradient of the quadratic form x^T A x is (A + A^T) x, which becomes 2Ax when A is symmetric.
solid answer
~40 sFor the linear form, `w^T x = w1*x1 + w2*x2 + ...`, so the partial with respect to `xk` is `wk` and the gradient is `w` itself — the vector analogue of differentiating `w*x` to get `w`. For the quadratic form, write `x^T A x = sum over i, j of xi * A_ij * xj`. Differentiating with respect to `xk` picks up two contributions, one where `k` sits in the row index and one where it sits in the column index, giving `(Ax)_k + (A^T x)_k`. So `grad = (A + A^T) x`, and when `A` is symmetric that collapses to `2Ax`. Two guards: state the symmetry assumption rather than writing `2Ax` reflexively, and check shapes — the gradient is a vector the same size as `x`, never a matrix.
go deeper
Recognise the notation first: w^T x and x^T A x are single numbers built from a vector, not matrices, and their gradients are vectors of the same length as x.
Derive both results by expanding into index notation and differentiating one entry, and be able to point at where the factor of two in the quadratic case comes from.
Handle the general case cleanly: give (A + A^T)x first, state the symmetry condition that reduces it to 2Ax, and run the shape and one-dimensional checks aloud.
Own why the symmetric form is the one that matters: only the symmetric part of A affects the quadratic form's value, so symmetrising is a free normalisation rather than an assumption that costs generality.
## Why these two forms A huge share of the objectives that appear in modelling and optimisation are built from exactly two pieces: a **linear form** `w^T x`, a weighted sum of the entries of `x`, and a **quadratic form** `x^T A x`, a weighted sum of all pairwise products of entries. Knowing their gradients cold means you can differentiate most objectives without dropping to index notation each time. Interviewers ask it to see whether you understand vector differentiation or merely memorised two formulas. ## Conventions first Take `x` to be a column vector with `n` entries, `w` a constant column vector with `n` entries, and `A` a constant `n`-by-`n` matrix. Both `w^T x` and `x^T A x` evaluate to a single number. The gradient of a scalar with respect to `x` is the vector of partial derivatives, one per entry of `x`, so it has the same shape as `x`. That shape rule is the fastest error check available: if your candidate answer is a matrix or a scalar, it is wrong before you check any arithmetic. ## The linear form Expanded, `w^T x = w1*x1 + w2*x2 + ... + wn*xn`. Differentiate with respect to a single entry `xk`. Every term except the `k`-th is constant in `xk` and contributes zero; the `k`-th term is `wk * xk`, contributing `wk`. So the `k`-th component of the gradient is `wk`, for every `k`, and therefore `grad_x (w^T x) = w` The scalar analogue is `d/dx (w*x) = w`. The gradient is constant — the same vector at every point — which is exactly what 'linear' means: the rate of change does not depend on where you are. ## The quadratic form Expanded with indices, `x^T A x = sum_i sum_j xi * A_ij * xj`. Differentiate with respect to `xk` and ask which terms survive. 1. Terms where `i = k` and `j` is anything other than `k`: each is `xk * A_kj * xj`, contributing `A_kj * xj`. Summing over `j` gives the `k`-th entry of `A x`. 2. Terms where `j = k` and `i` is anything other than `k`: each is `xi * A_ik * xk`, contributing `A_ik * xi`. Summing over `i` gives the `k`-th entry of `A^T x`. 3. The single term where `i = j = k` is `A_kk * xk^2`, whose derivative is `2 * A_kk * xk` — which is precisely one contribution from each of the two counts above, so it is already included. Adding the pieces, the `k`-th component of the gradient is `(Ax)_k + (A^T x)_k`, hence `grad_x (x^T A x) = (A + A^T) x` ## The symmetric case If `A` is symmetric, meaning `A^T = A`, the two contributions coincide and `grad_x (x^T A x) = 2 A x` This is the version most often quoted, because the matrices appearing in quadratic objectives are very frequently symmetric by construction. But quoting `2Ax` for a general `A` is a genuine error, and interviewers plant asymmetric matrices to catch it. The safe habit is to derive `(A + A^T)x` and then say 'which is `2Ax` because `A` here is symmetric'. A related fact worth knowing: only the symmetric part of `A` affects the value of `x^T A x` at all, because the antisymmetric part contributes zero for every `x`. That is the deeper reason the gradient depends on `A` only through `A + A^T`. ## Special cases and checks **Sum of squares.** Take `A` to be the identity matrix. Then `x^T x = x1^2 + x2^2 + ...`, and the formula gives `grad = 2x`. That matches the direct computation: the partial with respect to `xk` is `2*xk`. **One dimension.** With `n = 1`, `w^T x` is `w*x` with derivative `w`, and `x^T A x` is `a*x^2` with derivative `2*a*x`. Both formulas reduce correctly, which makes the scalar case a reliable memory aid. **Shape check.** `A` is `n`-by-`n` and `x` is `n`-by-1, so `(A + A^T)x` is `n`-by-1, matching `x`. Any answer of the wrong shape — say, `A` alone, or `x^T A` — fails immediately. **A trap.** The reflex 'power rule' answer `A x` for the quadratic form drops one of the two ways `xk` enters the double sum. Explaining *why* the factor of two appears — `x` occurs on both sides of `A` — is the part that distinguishes a derived answer from a recalled one. ## What good answers sound like A strong candidate writes both results, names the symmetry condition explicitly, sketches the index argument for the factor of two in under a minute, and verifies with the scalar case. A weak one recites `w` and `2Ax` and cannot say what happens when `A` is not symmetric.
- Why does the symmetric case give 2Ax rather than Ax?Because x enters the double sum twice, once through the row index and once through the column index. Differentiating with respect to one entry picks up both, giving (Ax) + (A^T x). When A^T = A those two are identical, so they add to 2Ax rather than cancelling into one copy.
- What is the gradient of x^T x with respect to x?It is 2x. This is the quadratic form with A equal to the identity matrix, and the identity is symmetric, so the formula gives 2 * I * x = 2x. You can confirm directly: x^T x is the sum of squares of the entries, and the partial with respect to entry k is 2 * xk.
- How would you sanity-check a vector-calculus identity like these under interview pressure?Two cheap checks. First, shapes: a gradient with respect to x must be a vector the same size as x, which rules out any matrix-valued answer. Second, reduce to one dimension: w*x differentiates to w, and a*x^2 to 2ax. Any candidate formula that fails either check is wrong.
- Does the antisymmetric part of A affect the quadratic form at all?No. Any antisymmetric matrix contributes zero to x^T A x for every x, so only the symmetric part of A influences the value. That is exactly why the gradient depends on A only through the combination A + A^T, and why assuming symmetry is usually harmless in practice.
saying these in an interview costs you the question
- States the gradient of x^T A x is Ax regardless of symmetry
- Quotes 2Ax without mentioning the symmetry assumption
- Returns a matrix where a vector is required
- Claims the gradient of a linear form depends on x
- Cannot explain where the factor of two comes from