Your sigmoid hidden layer outputs 0.999 for every training example - what caused it?
answer
- the failure exists before the first update
- invert the curve to recover the input
- check the magnitude of the raw features
- the layer is now a constant
- fix the feature scale, not the optimizer
basics
~20 sThe pre-activations are far too large and positive, because the input features were fed in unscaled. Decibel-scale loudness values in the tens, summed over many channels, push every unit deep into the sigmoid's flat tail.
solid answer
~50 sThat is arithmetic, not an optimization mystery. Each unit computes a weighted sum of its inputs plus a bias, and if the incoming features are decibel-scale loudness values with magnitudes in the tens, summed across sixty channels, the pre-activation lands in the tens or hundreds - and `sigma(7)` is already 0.999. Two things then break at once. Backward, the derivative is `0.999 * 0.001`, about 0.001, roughly 250 times below the sigmoid's 0.25 peak, so nothing beneath the layer learns. Forward, an output that is 0.999 for every example is a constant: the layer has destroyed the input's information, and the model can only fit the output bias, so the loss plateaus near the base rate. The fix is at the input, not the optimizer - standardize each feature channel to roughly zero mean and unit variance so pre-activations land where the curve has slope.
go deeper
Recognise that an output of 0.999 means the input to that unit was large and positive, and that a flat part of the curve means almost no learning. Saying the inputs probably were not scaled is already the right instinct.
Do the arithmetic out loud: invert the sigmoid to recover a pre-activation near 7, compute the derivative as 0.999 times 0.001, and connect the feature magnitudes and the number of summed channels to that pre-activation.
Show both failure directions - the attenuated gradient and the constant, information-free representation - predict the loss curve that follows, and pick the fix at the input rather than at the optimizer. Say how you would confirm it before changing code.
Frame this as a guardrail question: which invariants about input scale should hold before any model trains, who owns them in the pipeline, and how a team catches a step-zero scaling failure automatically instead of rediscovering it as a mysterious plateau.
## Reading the symptom An entire sigmoid hidden layer pinned at 0.999 across the whole dataset is one of the most diagnosable failures in neural network training, because the number tells you the pre-activation. Inverting the sigmoid, an output of 0.999 means an input of about 6.9, and 0.9999 means about 9.2. The layer is not mildly biased; it is sitting several units out into the flat tail. ## Where the number came from The pre-activation of a unit is a weighted sum of the incoming features plus a bias. Its scale is the product of three things: how large the input values are, how large the weights are, and how many terms are being summed. In the classic version of this failure, one of those dominates. Decibel-scale audio loudness features are not small numbers - values in the tens are normal, and they are all-positive and all of similar magnitude. Feed sixty such channels straight into a hidden layer with ordinary small random weights and any unit whose weight row happens to sum positive will produce a pre-activation in the tens. Every such unit reports 0.999 on the first forward pass, before a single update has happened. That last detail is the giveaway: a failure present at step zero is a scaling failure, not a learning-dynamics failure. ## Two failures, not one Candidates usually name the gradient problem and stop. Name both. **Backward.** The sigmoid derivative in terms of its own output is `output * (1 - output)`, so at 0.999 it is about 0.001. That is roughly 250 times smaller than the 0.25 the unit would contribute at its best operating point. The error signal reaching the weights of this layer, and everything beneath it, is attenuated to nothing. **Forward.** This one is worse and is often missed. If the output is 0.999 for *every* example, the layer's output vector is essentially constant. Two very different inputs produce the same representation. Whatever discriminative information the features held has been thrown away, so even a perfectly healthy set of layers above has nothing to separate. Expect the loss to fall for an epoch or two - the output bias is still fitting the class prior - and then flatten at roughly the base rate. ## What actually fixes it **Standardize the features.** Subtract the per-channel mean and divide by the per-channel standard deviation, using statistics computed on the training split only. This moves the pre-activations back toward the origin, where the sigmoid has slope, and it is the direct fix because the input scale is what is out of line here. **Then confirm.** After the change, the layer's outputs should spread across the interval rather than clustering at one end, and the loss should keep falling past the point where it previously flattened. ## What does not fix it **Raising the learning rate.** The gradient reaching the saturated layer has been divided by hundreds; multiplying the step size back up destabilises every part of the network that was training normally, and it does not restore the information the forward pass destroyed. **Adding capacity.** More units or more layers of the same kind, fed by the same unscaled features, saturate identically. The model is not short of parameters; it is short of a usable representation. **Relying on the bias to absorb the offset.** A bias can shift a pre-activation, but it cannot rescale one, and it is a single learned number that would have to fight a fixed input scale of tens - from a starting position where its own gradient is already attenuated a thousand-fold. **Swapping in tanh.** Tanh has the same flat tails. With pre-activations in the tens every unit pins at 1.0 and `1 - tanh(x)^2` is around 1e-6. Changing the squashing function does not change the arithmetic that produced the large pre-activation. ## The general lesson A saturating unit only works in the neighbourhood where its curve has slope, roughly a pre-activation magnitude below three or four. Keeping the inputs to such a layer in that neighbourhood is a precondition for training, not an optimization nicety - and it is cheap to guarantee at the point where the features enter the network. The reason this failure feels mysterious the first time is that it presents as a training problem when it is entirely a data-scaling problem, visible in the very first forward pass if anyone looks at the numbers the layer is actually producing.
- Would swapping the sigmoid for tanh fix this?No. Tanh saturates at 1.0 with the same flat tail, and with pre-activations in the tens its derivative `1 - tanh(x)^2` is around 1e-6 - worse, not better. The problem is that the weighted sum is enormous, and the choice of squashing function does not change that sum. Fix the feature scale and tanh becomes a reasonable choice again.
- The loss drops for two epochs and then flattens well above what you expect. Is that consistent with this diagnosis?Yes, and it is the signature. With the hidden layer constant, the only thing that can still improve the loss is the output bias fitting the class prior, which happens quickly and then has nothing left to do. A plateau at roughly the base-rate loss, reached early and immovable, is what a dead representation looks like from the loss curve.
- How would you confirm the diagnosis before changing anything?Look at the raw feature magnitudes first - decibel values in the tens against weights of order 0.1 already predict a pre-activation in the units-to-tens range. Then check that the pinned outputs are present on the very first forward pass, before any update. If both hold, the cause is input scale and no further investigation is needed.
saying these in an interview costs you the question
- Blames the learning rate rather than the input scale
- Suggests more layers or more units as the remedy
- Assumes the bias term will learn away an input scale of tens
- Mentions the vanished gradient but not the destroyed representation
- Assumes the units can never recover because the gradient is exactly zero
- Proposes clipping the outputs below 0.99 to keep gradients alive