Why does the SELU activation function preserve mean 0?

Viewed 134

From Aurelien Geron's book "Hands-On Machine Learning with Scikit-Learn, Keras & Tensorflow", p. 337:

"The authors showed that if you build a neural network composed exclusively of a stack of dense layers, and if all hidden layers use the SELU activation function, then the network will self-normalize: the output of each layer will tend to preserve a mean of 0 and a standard deviation of 1 during training, which solves the vanishing/exploding gradients problem.

My question is: Why does it preserve a mean of 0? Negative values are moved much more towards 0 than positive values are, so why doesn't the output mean exceed the input mean?

1 Answers

Note that it doesn't preserve mean 0 on its own, but only when starting with variance 1 as well, so negative values are mostly small and not

moved much more towards 0 than positive values are

The paper says is that normalizing the variance is the primary effect and normalizing the mean follows from it:

To give an intuition, the main property of SELUs is that they damp the variance for negative net inputs and increase the variance for positive net inputs. The variance damping is stronger if net inputs are further away from zero while the variance increase is stronger if net inputs are close to zero. Thus, for large variance of the activations in the lower layer the damping effect is dominant and the variance decreases in the higher layer. Vice versa, for small variance the variance increase is dominant and the variance increases in the higher layer...

Therefore SELU networks control the variance of the activations and push it into an interval, whereafter the mean and variance move toward the fixed point. Thus, SELU networks are steadily normalizing the variance and subsequently normalizing the mean, too.

Maybe this link will help too: https://medium.com/@damoncivin/self-normalising-neural-networks-snn-2a972c1d421

Related