What is ReLU, and why is it the default hidden-layer activation in deep networks?
answer
- a hinge at the origin
- one branch passes, one blocks
- the active-branch slope is constant
- that constant is exactly one
- a clamp, not an exponential
basics
~20 sReLU computes max(0, z): positive pre-activations pass through unchanged and negative ones become zero. Its derivative is exactly 1 wherever the unit is active, so gradients flow back unshrunk, and evaluating it costs a single comparison.
solid answer
~50 sReLU is the elementwise map `f(z) = max(0, z)` applied to a unit's pre-activation `z = w·x + b`. Its derivative is 1 for `z > 0` and 0 for `z < 0`. That constant 1 is the whole point: on the active path the activation contributes no shrinking factor to the backward signal, so the only rescaling comes from the weights themselves and depth stops costing gradient magnitude by default. It is also nearly free to compute — a clamp, with no exponential. That combination of cheap and non-saturating on the active branch is why it became the default hidden nonlinearity. The prices you pay: outputs are never negative, so the next layer sees a positively shifted input mean; the output is unbounded above; and the flat negative branch means a unit whose pre-activation never rises above zero gets no gradient at all.
go deeper
Be ready to write max(0, z), sketch the hinge, and state the derivative on each side: one where the unit is active, zero where it is not. That much is expected in a first screen.
Explain why a slope of exactly one leaves the backward signal unscaled by the activation itself, and why a network of ReLUs is still nonlinear even though each branch is linear.
An interviewer expects the costs you manage in real training runs: non-negative outputs shifting the next layer's input mean, no upper bound on activations, and a flat negative branch that can silence units permanently.
Own the default. Argue when a fleet of models should stay on the cheapest nonlinearity versus paying extra compute per element for a small accuracy gain, and price that against inference cost, not just training curves.
## What the unit computes Each hidden unit first forms a **pre-activation** `z = w·x + b` — a weighted sum of its inputs plus a bias — and then passes it through a nonlinearity. ReLU, the *rectified linear unit*, is that nonlinearity in its simplest useful form: ``` f(z) = max(0, z) ``` Positive pre-activations come through untouched; negative ones are clamped to exactly zero. Graphically it is a hinge at the origin: a flat segment along the negative axis, then a 45-degree line. ## The derivative, which is the reason it exists ``` f'(z) = 1 for z > 0 f'(z) = 0 for z < 0 ``` Backpropagation forms a unit's error signal by multiplying the error arriving from above by that local derivative. With a slope of exactly 1, an **active** unit neither shrinks nor amplifies what passes through it — the activation is transparent, and the only rescaling of the backward signal comes from the weight matrices. An activation whose derivative is always well below one does the opposite: it multiplies in a factor smaller than one at every layer, and those factors compound with depth. Removing that per-layer shrink on the active path is the single mechanical reason ReLU let people train much deeper stacks than the earlier squashing units allowed. ## Piecewise linear, but not linear A common misreading is that ReLU is "basically linear", so a stack of them collapses into one linear map. It does not. For a *fixed input*, the set of active units is fixed, and on that region the whole network is an affine function — but the active set changes as the input moves. The network is therefore **piecewise linear**: the input space is carved into regions, each with its own affine map, and the boundaries between regions are learned. That is a genuinely nonlinear function class, and it is expressive enough for the usual universal-approximation results. ## Sparsity For any given input, a substantial fraction of units sit on the flat branch and emit exact zeros. This is often sold as a feature, and there is a real representational argument: different inputs recruit different subsets of units, so the representation is conditional rather than uniformly dense. What it is *not* is free speed. A dense matrix product multiplies and accumulates the zeros like any other number; you only save arithmetic if something explicitly skips them, and unstructured, input-dependent zeros are hard to exploit. Treat sparsity as a property of the representation, not as an optimisation. ## The costs you sign up for 1. **Non-negative outputs.** Every value a ReLU layer emits is at least zero, so the next layer's inputs have a strictly positive mean. That shifts its pre-activations, and it means all of the incoming weights of a downstream unit receive gradients sharing the sign of that unit's upstream error, so the weight vector tends to move in one direction at a time rather than freely. Normalisation layers, or a nonlinearity that can emit negative values, offset it. 2. **Unbounded above.** Nothing caps the activation, so bad scaling upstream shows up as very large forward values rather than being squashed. 3. **A flat negative branch.** A unit whose pre-activation is negative on *all* data emits zero everywhere and receives exactly zero gradient — the dead-unit failure, which is the price of the clean positive-branch derivative. 4. **A kink at zero.** `f` is not differentiable at `z = 0`. Any value in the interval from 0 to 1 is a valid subgradient there and a fixed convention is simply chosen; a pre-activation that is exactly zero in floating point is vanishingly rare, so the choice has no practical effect. ## How to answer the "why default" part Because it is the cheapest thing that works: one comparison per element, no transcendental function, a derivative of exactly one where the unit is active, and no upper saturation to stall learning on that branch. Later nonlinearities smooth the kink or reshape the negative branch, and they buy small, workload-dependent improvements — but they all cost more per element, and ReLU remains the reasonable starting default that you deviate from with evidence.
- ReLU has no derivative at exactly zero — how is that point handled?Any value between 0 and 1 is a valid subgradient at the kink, and a fixed convention is chosen and used consistently. It does not matter in practice: a pre-activation that lands on exactly 0.0 in floating point is essentially a measure-zero event, and if it did happen the convention would affect one unit for one step.
- ReLU never outputs a negative number. Why does that bother the following layer?The next layer sees inputs with a strictly positive mean, which biases its pre-activations. It also means every incoming weight of a downstream unit gets a gradient carrying the sign of that unit's own error signal, so the whole weight vector is pushed in one direction at a time instead of adjusting component by component. Normalisation, or an activation that can emit negatives, offsets it.
- Does ReLU's sparsity make a dense layer's forward pass faster?Not on its own. A dense matrix product multiplies and adds the zeros like any other value; you only save work if something explicitly skips them, and unstructured, input-dependent zeros are rarely worth exploiting. The real argument for ReLU sparsity is representational — different inputs recruit different subsets of units — not arithmetic.
It works like a one-way valve. Pressure above zero passes through at full strength; anything below is shut off completely, and the valve adds no resistance of its own when it is open.
saying these in an interview costs you the question
- Says a stack of ReLUs collapses into a single linear map
- Claims ReLU eliminates gradient problems entirely
- Thinks ReLU's derivative equals the pre-activation value
- Says the zero outputs speed up dense matrix multiplication
- Calls ReLU differentiable everywhere, including at zero