For f(x, y) = x^2 * y, what are the two partial derivatives and the gradient at (2, 3)?
answer
- one variable moves, the rest freeze
- y is just a constant when differentiating x
- power rule fires only where the exponent is
- substitute the point last, not first
- stack the partials into a vector, in order
basics
~20 sDifferentiate one variable at a time, holding the other fixed: df/dx = 2xy and df/dy = x^2. At (2, 3) those evaluate to 12 and 4, so the gradient there is the vector (12, 4).
solid answer
~50 sA partial derivative differentiates with respect to one variable while every other variable is frozen as a constant. For `f(x, y) = x^2 * y`, freezing `y` leaves a constant times `x^2`, so `df/dx = 2xy`. Freezing `x` leaves a constant times `y`, so `df/dy = x^2`. Evaluating at the point `(2, 3)` gives `df/dx = 2 * 2 * 3 = 12` and `df/dy = 2^2 = 4`. The gradient is simply those partials stacked into a vector in a fixed variable order: `grad f(2, 3) = (12, 4)`. The 12 says f climbs about 12 units per unit step in x from that point; the 4 says it climbs about 4 units per unit step in y. The gradient is a vector, not a single number, and it has one component per input variable.
go deeper
Be ready to differentiate a two-variable expression on the spot, holding the other variable constant, and to write the answer as a vector in the given variable order.
Explain why the mechanics work: freezing a variable turns the expression into a single-variable function, and the gradient is only the assembled list of those one-at-a-time rates.
Show that you read the numbers, not just compute them: state what each component means per unit step, note that the components can carry different units, and flag points where the gradient vanishes.
Own the framing question of whether the variables are even on comparable scales. Component sizes depend entirely on the units chosen, so 'which input matters most' is a modelling decision, not something the raw partials settle.
## What a partial derivative is A function of one variable has one derivative: the instantaneous rate at which the output changes as the single input moves. A function of several variables has one rate per input, because the output can change for several independent reasons. The **partial derivative** of `f` with respect to `x`, written `df/dx`, is the rate of change of `f` as `x` moves and **every other variable is held fixed at its current value**. Operationally that means: pretend the other variables are numbers, and differentiate as if the function had only one variable. ## Working the example Take `f(x, y) = x^2 * y`. **With respect to x.** Treat `y` as a constant, say `c`. The function reads `c * x^2`, whose derivative is `c * 2x`. Putting `y` back in its place: `df/dx = 2xy`. **With respect to y.** Treat `x` as a constant, say `k`. The function reads `k^2 * y`, which is a constant multiple of `y`, so its derivative is `k^2`. Putting `x` back: `df/dy = x^2`. Notice the asymmetry: the exponent lives on `x`, so the power rule fires only in the first case. There is no product rule to apply here in either direction, because in each case exactly one factor varies and the other is a frozen constant. ## Evaluating at a point So far both partials are still functions of `x` and `y` — they describe rates everywhere on the surface. A specific point turns them into numbers. At `(x, y) = (2, 3)`: - `df/dx = 2xy = 2 * 2 * 3 = 12` - `df/dy = x^2 = 2^2 = 4` ## Assembling the gradient The **gradient** of `f` is the vector whose components are the partial derivatives, listed in a fixed variable order: `grad f(x, y) = (df/dx, df/dy) = (2xy, x^2)` and at the point `(2, 3)`: `grad f(2, 3) = (12, 4)` Two things to keep straight. First, order matters: `(12, 4)` and `(4, 12)` describe different directions, and swapping them is one of the most common careless errors. Second, the gradient is a vector with one component per input, never a scalar — reporting `16` (the sum) or `12` (just the x-partial) is a category mistake, not an arithmetic one. ## Reading the numbers Geometrically, `f` is a surface sitting above the `(x, y)` plane. Standing at the point `(2, 3)`, the number 12 is the slope of the surface if you walk due east (increasing `x` with `y` pinned at 3), and 4 is the slope if you walk due north (increasing `y` with `x` pinned at 2). Both are *local, first-order* statements: they describe the tangent behaviour right at that point. Step a finite distance and the actual change will differ from `12 * (change in x) + 4 * (change in y)`, by an amount that shrinks as the step shrinks. Units follow the same logic. If `f` is a cost in dollars, `x` is hours and `y` is headcount, then `df/dx` carries units of dollars per hour and `df/dy` dollars per person. The components of a gradient need not share units, which is one reason "the gradient is big" is only meaningful once you have fixed the scaling of the inputs. ## Where partials vanish Evaluate the same gradient at `(0, 3)`: `df/dx = 2 * 0 * 3 = 0` and `df/dy = 0^2 = 0`, so `grad f(0, 3) = (0, 0)`. The gradient is the zero vector along the whole line `x = 0`. At such points the surface is flat to first order and no direction is singled out as uphill — which is exactly why zero gradients get special attention when you look for maxima, minima and flat ridges. ## Checklist for the interview 1. Say the rule out loud: one variable moves, the rest are constants. 2. Differentiate each partial symbolically first. 3. Substitute the point only at the end. 4. Report a vector, in the stated variable order.
- What does the number 12 actually mean at that point?It is the instantaneous rate of change of f per unit increase in x, with y pinned at 3. Geometrically it is the slope of the surface along the x-direction at (2, 3). It is a first-order statement: a small step dx in x changes f by roughly 12 * dx, and the approximation degrades as the step grows.
- What is the gradient of the same function at (0, 3), and why is that interesting?Both partials vanish: df/dx = 2 * 0 * 3 = 0 and df/dy = 0^2 = 0, so the gradient is (0, 0). In fact it is zero along the entire line x = 0. Where the gradient is the zero vector the surface is flat to first order and no direction stands out as uphill or downhill.
- Is a gradient a vector of numbers or a function?Both, depending on where you stop. Symbolically grad f(x, y) = (2xy, x^2) is a vector-valued function defined at every point — a gradient field. Evaluate it at one point and you get a concrete vector of numbers, here (12, 4). Interviewers listen for whether you keep the two apart and only substitute the point at the end.
It is like measuring the slope of a hillside twice: once walking strictly east, once strictly north. Each walk gives one number, and together they describe the tilt at that spot.
saying these in an interview costs you the question
- Applies the product rule to x^2 * y as if both factors varied
- Reports the gradient as a single number such as 16
- Swaps the components and answers (4, 12)
- Substitutes the point before differentiating
- Leaves the answer symbolic when a point was given