How does automatic differentiation differ from symbolic and numerical differentiation?
answer
- three ways to get a derivative
- one rewrites formulas, one perturbs inputs
- expression swell versus step-size error
- exact numbers, fixed cost per operation
basics
~20 sAutomatic differentiation applies exact per-operation derivative rules to numeric values as the program runs, giving machine-precision derivatives at one input point. Symbolic differentiation manipulates formulas and can blow up in size; numerical differentiation perturbs inputs and carries approximation error.
solid answer
~50 sAll three compute derivatives, but they differ in what they produce. Symbolic differentiation transforms an expression into another expression; on a deeply nested composition the product rule keeps duplicating subterms, so the printed derivative can grow explosively — expression swell — and it needs a closed form, which data-dependent control flow does not provide. Numerical differentiation evaluates `(f(x+h) - f(x)) / h`, which is easy but approximate: too large an `h` leaves truncation error, too small an `h` loses digits to cancellation, and it costs one evaluation per input direction. Automatic differentiation sits between them: it decomposes the program into primitive operations, applies a fixed exact rule to each, and propagates *numbers*, reusing shared intermediates. The result is exact to floating point at the current point, at a constant multiple of the cost of running the program.
go deeper
Recall the three-way split and one sentence on each: formulas out, perturbations in, or exact numbers at a point. Naming expression swell and step-size error already puts you ahead.
Explain the mechanics: how the product rule duplicates subterms into expression swell, and how truncation and rounding error trade off against each other as the step size changes.
Show you use the distinction in practice — finite differences as a verification tool for a hand-written derivative rule, with a realistic tolerance, rather than as a way to obtain gradients.
Be ready to argue why differentiability of the whole pipeline is worth engineering for, and where a non-differentiable or black-box component forces a different estimator entirely.
## Three different products Ask for 'the derivative' and you can get three different things. **Symbolic differentiation** takes an expression and returns another expression. Give it `f(x) = x * sin(x)` and it returns `sin(x) + x * cos(x)`, valid for every `x`. That generality is genuinely useful, and where a closed form exists and stays small it is the best of the three. **Numerical differentiation** takes a function you can only call, and estimates the derivative by perturbing the input: the forward difference `(f(x+h) - f(x)) / h`, or the central difference `(f(x+h) - f(x-h)) / (2h)`. It needs no access to the internals at all. **Automatic differentiation** takes the *program* and returns derivative *values* at the input you supplied. It is exact in the same sense the program is exact — up to floating-point rounding — but it produces numbers, not a formula. ## Why symbolic differentiation breaks down on deep compositions A network is a deep composition `f_k( ... f_2( f_1(x) ) ... )`. Apply the product and chain rules symbolically and each level distributes over the level below, duplicating subterms rather than naming them. The printed derivative can grow exponentially in the nesting depth even though the function itself is cheap to evaluate. This is **expression swell**, and it is the classic reason a hand-derived or computer-algebra derivative of a deep model is unusable while the same derivative computed op-by-op is trivial. Automatic differentiation avoids it by construction: the derivative of each primitive is a fixed local rule with no growth of its own, and every shared intermediate is computed once and reused, because it is a *value* sitting in a variable rather than a subterm sitting in an expression tree. Cost stays a small constant multiple of evaluating the function. There is a second limit. Symbolic differentiation needs a closed-form expression. A program with data-dependent branches or loops whose trip count depends on the input has no single formula; automatic differentiation simply differentiates the sequence of primitive operations that actually executed for this input, which is the derivative of the piece of the function you are standing on. Worth knowing that this is only meaningful away from kinks: at a point where the branch flips, the returned derivative is one side's, not a subgradient of the whole. ## Why finite differences are the wrong default Two errors pull in opposite directions. **Truncation error** comes from the Taylor remainder and shrinks with `h`: order `h` for a forward difference, order `h^2` for a central one. **Rounding error** comes from subtracting two nearly equal floating-point numbers and grows like machine epsilon divided by `h`. The total error is minimised at a sweet spot — around the square root of machine epsilon for a forward difference, times the scale of `x` — and even at that optimum you keep only about half the significant digits. Shrinking `h` further makes things *worse*, which is the single most common misconception here. Cost is the other problem. A finite-difference gradient needs one extra evaluation per input direction, so a model with millions of parameters is out of the question. Automatic differentiation's reverse sweep gets the whole gradient for a constant multiple of one evaluation. Finite differences still have one excellent use: as an **independent check**. If you write a new derivative rule by hand, compare its directional derivative against a central difference along a few random directions in double precision, and expect agreement to several digits, not to machine precision. A curiosity worth knowing: for real analytic functions the complex-step estimate `Im(f(x + i*h)) / h` has no subtractive cancellation, so it is accurate for very small `h`. ## Is automatic differentiation just symbolic differentiation in disguise? They are cousins — both apply exact chain-rule identities, and reverse mode can be described as symbolic differentiation with aggressive sharing of common subexpressions, evaluated eagerly. The distinctions that matter in an answer: automatic differentiation carries numeric values rather than expressions, so nothing swells; it has a bounded cost per primitive, so the total cost is predictable; and it follows the executed path, so control flow is not an obstacle. What it gives up is generality — you get the derivative at one point, and you must run it again at the next point. ## A compact summary - Symbolic: expression in, expression out. Exact everywhere, can swell, needs a closed form. - Numerical: black box in, approximation out. Trivial to apply, error-prone, one evaluation per direction. - Automatic: program in, numbers out. Exact at the point, fixed cost multiple, handles arbitrary control flow — which is why every trained network relies on it.
- Why does a finite-difference step size have a sweet spot instead of being as small as possible?Truncation error shrinks as the step shrinks, but rounding error grows: subtracting two nearly equal values loses significant digits, and the difference is then divided by a tiny number. The two curves cross around the square root of machine epsilon for a forward difference, and even there you keep only about half your digits.
- Is reverse-mode automatic differentiation just symbolic differentiation with shared subexpressions?They are closely related — both apply exact chain-rule identities — but automatic differentiation evaluates numbers at one point rather than emitting an expression, spends a bounded amount of work per primitive so its total cost is a constant multiple of the function, and differentiates the path the program actually took, which no closed-form expression can do.
- When is a finite-difference derivative still the right tool?As an independent check on a hand-written derivative rule, and when the function is a black box you cannot instrument. Compare a central difference against the analytic directional derivative along a few random directions in double precision and expect several matching digits, not exact equality.
Symbolic differentiation is deriving a general formula for a journey's speed; numerical differentiation is timing two nearby mileposts with an imprecise stopwatch; automatic differentiation is reading the speedometer at the exact moment you pass one milepost.
saying these in an interview costs you the question
- Says automatic differentiation is finite differences under the hood
- Claims automatic differentiation hands you a derivative formula to print
- Thinks a smaller finite-difference step is always more accurate
- Treats symbolic differentiation as free and ignores expression swell
- Cannot name any error source in a finite-difference estimate