Weight initialisation is the procedure of assigning starting values to the trainable parameters of a neural network before training commences. The choice of initialisation scheme affects gradient flow, convergence speed, and the avoidance of vanishing or exploding activations across deep layers. Common schemes such as Xavier (Glorot) and He initialisation scale the variance of initial weights according to layer fan-in and fan-out.
Overview
- Initialising weights to small random values breaks symmetry between neurons so that they learn distinct features. Naive choices, such as all-zero or large-magnitude weights, cause neurons to update identically or to saturate non-linear activations, stalling learning. Variance-scaling schemes keep the signal variance roughly constant as it propagates forward and backward through many layers.
Key aspects
- Symmetry breaking through random sampling so neurons diverge during training
- Variance scaling by layer fan-in/fan-out (Xavier/Glorot, He/Kaiming)
- Matching the scheme to the activation function (He for ReLU, Xavier for tanh/sigmoid)
- Bias initialisation, typically to zero or small constants
- Orthogonal initialisation for recurrent architectures to preserve gradient norms
Applications
- Training deep convolutional networks for image classification
- Stabilising very deep residual and transformer architectures
- Improving convergence speed during hyperparameter search
- Reducing sensitivity to learning-rate selection