Dive into Deep Learning · §4.4
Numerical stability & initialization
Why deep nets once refused to train: the problem, the diagnosis, and the cure.
Motivation
A deep net composes many layers before any loss is seen. The initial weights decide whether a signal survives the trip, forward and back.
Three ideas made deep training routine:
Get init wrong and the gradient dies or blows up before learning starts. Keep score: ten sigmoid layers tax the gradient to 10^{-6}; a hundred random matrices explode past 10^{24}; and one closing plot shows 10^{80} vs 10^{-15} vs flat (naive, Xavier, He).
01
Unstable Gradients
why the chain rule makes depth dangerous
Unstable Gradients
Backprop multiplies one Jacobian per layer. For a weight in layer \ell,
\partial_{\mathbf{W}^{(\ell)}} \mathbf{o} = \underbrace{\mathbf{M}^{(L)} \cdots \mathbf{M}^{(\ell+1)}}_{L-\ell \text{ Jacobians}}\, \mathbf{v}^{(\ell)}, \qquad \mathbf{M}^{(k)} = \partial_{\mathbf{h}^{(k-1)}} \mathbf{h}^{(k)}.
Forward pass builds z, h, o, L; the backward sweep multiplies a gradient back through every node.
Unstable Gradients
Whether the product grows or shrinks is set by the Jacobians’ scale.
A constant per-layer factor \rho compounds to \rho^{\,L-\ell}. Only \rho \approx 1 stays usable across depth.
Unstable Gradients · vanishing
The sigmoid’s derivative peaks at 0.25 and is flat at zero in both tails.
Stack ten such layers and 0.25^{10} \approx 10^{-6}: the bottom layer sees a millionth of the gradient.
ReLU’s derivative is exactly 1 wherever a unit is active, so it does not attenuate the signal. Hence ReLU is the modern default.
Unstable Gradients · exploding
Multiply one hundred \mathcal{N}(0,1) matrices (for i in range(100): M = M @ randn(4, 4)), exactly what a deep linear stack does to a gradient. Each factor is a little too large, and the product compounds:
a single matrix [[ 0.21544254 -1.5842865 0.73179865 -0.6055198 ]
[-0.41833407 -0.09534191 -0.02739832 0.9578756 ]
[-0.60731924 0.97379816 -0.40570006 0.12149587]
[-0.1136004 0.49238634 -1.2923752 0.5174167 ]]
after multiplying 100 matrices [[ 6.1329105e+22 9.0935597e+21 3.0797470e+22 -2.1217980e+22]
[-5.3375046e+22 -7.9141739e+21 -2.6803193e+22 1.8466116e+22]
[-1.6772870e+22 -2.4869927e+21 -8.4227840e+21 5.8028943e+21]
[ 2.8053260e+21 4.1596003e+20 1.4087421e+21 -9.7055689e+20]]
A poorly scaled initialization does exactly this to the gradient. No optimizer converges from here.
Unstable Gradients · in practice
Unstable Gradients · symmetry
Set every weight in a layer to the same constant c:
An h-unit layer is then stuck behaving like a single unit, forever.
Gradient descent alone never breaks this tie. Random initialization does; so does dropout. Bias may still start at 0.
02
Variance-Preserving Init
keep the signal’s scale constant through depth
Initialization
For a linear layer o_i = \sum_{j=1}^{n_\textrm{in}} w_{ij} x_j with i.i.d. zero-mean weights (\textrm{Var} = \sigma^2) and inputs (\textrm{Var} = \gamma^2):
\mathbb{E}[o_i] = 0, \qquad \textrm{Var}[o_i] = n_\textrm{in}\, \sigma^2\, \gamma^2.
To carry the input’s variance through unchanged (\textrm{Var}[o] = \gamma^2), the only knob is \sigma^2:
\sigma^2 = \frac{1}{n_\textrm{in}}.
Initialization
The same variance count run backward through \mathbf{W}^\top sums over the n_\textrm{out} outputs instead:
\text{forward: } n_\textrm{in}\,\sigma^2 = 1 \qquad \text{backward: } n_\textrm{out}\,\sigma^2 = 1.
Both cannot hold at once unless n_\textrm{in} = n_\textrm{out}, so Xavier splits the difference by averaging the two fan sizes.
Preserve the activation scale on the way in and the gradient scale on the way out: one \sigma^2, two demands.
Initialization
Xavier / Glorot (2010), for \tanh and sigmoid:
\sigma^2 = \frac{2}{n_\textrm{in} + n_\textrm{out}}.
He / Kaiming (2015), for ReLU:
\sigma^2 = \frac{2}{n_\textrm{in}}.
ReLU zeroes half a symmetric signal, halving its second moment (E[\textrm{ReLU}(z)^2] = \tfrac{1}{2}E[z^2]), so He doubles the weight variance to compensate.
Rule of thumb: Xavier for \tanh/sigmoid, He for ReLU. Both ship as named initializers in most libraries.
Initialization · payoff
All three regimes in one experiment: push a unit-scale signal through 50 ReLU layers of width 100 and track the second moment E[(h^{(l)})^2] layer by layer, under three weight scales.
Only the He-initialized stack delivers usable forward signals here. A backward variance calculation gives the same scale under mean-field independence assumptions. Run the sweep yourself in the notebook.
Beyond
Good init buys a deep net that trains without NaNs. To reach hundreds of layers, modern architecture re-normalizes during training:
We return to both in the chapters on modern CNNs.
Wrap-up
Next (the generalization-in-deep-learning section): the model trains, but why does an over-parametrized network generalize at all?