The Xavier initializer determines the initial weights of a neural network such that the variance of the activations will remain the same throughout the layers. Xavier Glorot and Yoshua Bengio proposed the initializer in 2010 as a solution for the vanishing and exploding gradients problem of deep feed-forward networks. The formula needs only two numbers: the number of input neurons (fan-in), and the number of output neurons (fan-out) of a particular layer.
Most of the production networks in 2026 include either an initializer along with batch normalization or residual connections. These two factors determine where the benefits of using the Xavier initialization arise. Therefore, it does not eliminate the need for selecting an initializer appropriately. This tutorial describes the formula, derivation of its two forms, implementation in PyTorch and TensorFlow, and conditions for which He initialization is more suitable.
What Will I Learn?
What Is Xavier Initialization?
Xavier initialization is a weight initialization method that sets a neural network’s initial weights using the number of input and output units in each layer. Xavier Glorot and Yoshua Bengio, researchers at the Université de Montréal, introduced the method in “Understanding the Difficulty of Training Deep Feedforward Neural Networks,” presented at the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) in 2010. The method is also called Glorot initialization, after its first author.
The method addresses one specific problem: in a network with several layers, the variance of the signal changes as it passes through each layer. If the variance shrinks, the signal vanishes by the final layers. If the variance grows, the signal explodes. Xavier initialization picks a starting variance for the weights that keeps the signal variance approximately constant from the first layer to the last.
Why Weight Initialization Matters
Weight initialization determines whether a network converges and how many training iterations it needs. Two separate failures result from a poor starting point: a symmetry failure and a variance failure.
What Happens If You Initialize Weights to Zero?
Initializing every weight to zero, or to any single constant, gives every neuron in a layer an identical output and an identical gradient during backpropagation, so the network cannot break symmetry and fails to learn distinct features.
A layer with two inputs and three hidden units illustrates the failure. If every weight equals the same constant, each of the three hidden units computes the same weighted sum of the inputs. Backpropagation then assigns the same gradient to each of the three units, because the calculation depends only on the shared weight value. The three units update identically through every training step, so the layer behaves as if it had one unit instead of three, regardless of how many iterations training runs.
Random initialization solves the symmetry failure by giving each weight a different starting value. It does not solve the variance failure described next.
Agentic AI Course
Average time:6 month(s) + Lifetime Access
Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving
The Variance Failure: Vanishing and Exploding Signals
A network’s activation functions determine how variance failures appear. The sigmoid function is linear near zero and flat at large positive or negative inputs.
- Weights set too small: the variance of the signal shrinks with each layer. By the final layers, the signal collapses toward zero, and the sigmoid function operates in its near-linear region, which removes the non-linearity that gives a multi-layer network its representational power.
- Weights set too large: the variance of the signal grows with each layer. The sigmoid function saturates at 0 or 1, its derivative approaches zero at that range, and the gradient during backpropagation approaches zero as a result. This is the vanishing gradient problem.
The Xavier Initialization Formula
Xavier initialization draws each weight from a distribution with a variance set by the fan-in and fan-out of the layer. Two versions of the formula exist: a uniform version and a normal (Gaussian) version.
Deriving the Target Variance
For a linear neuron y = w₁x₁ + w₂x₂ + … + wₙxₙ + b, with weights and inputs drawn independently from a zero-mean distribution, the variance of the output is:
Setting Var(y) equal to Var(x), so that variance stays constant across the layer, requires N × Var(w) = 1, which gives:
Glorot and Bengio extended this result to account for both directions of signal flow: the forward pass, which depends on fan-in, and the backward pass during backpropagation, which depends on fan-out. Averaging the two produces the published formula:
Normal Xavier Initialization
Normal Xavier initialization draws each weight from a Gaussian distribution with a mean of 0 and a standard deviation of σ = √(2 / (fan-in + fan-out)).
| Term | Meaning |
| fan-in | Number of input units to the layer |
| fan-out | Number of output units from the layer |
| σ | Standard deviation of the weight distribution |
Uniform Xavier Initialization
Uniform Xavier initialization draws each weight from a uniform distribution in the range [−a, a], where a = √(6 / (fan-in + fan-out)).
The value 6 in the uniform formula follows directly from the target variance above. A uniform distribution on [−a, a] has variance a² / 3. Setting a² / 3 equal to the target variance of 2 / (fan-in + fan-out) and solving for a gives:
The uniform and normal forms produce the same target variance; they differ only in the shape of the distribution the weights are drawn from.
Worked Example
A dense layer that maps a flattened 28×28 image (784 input units) to a 256-unit hidden layer has fan-in = 784 and fan-out = 256.
| Formula | Calculation | Result |
| Normal Xavier (σ) | √(2 / (784 + 256)) | 0.0439 |
| Uniform Xavier (a) | √(6 / (784 + 256)) | 0.0760 |
The layer’s weights are drawn either from a Gaussian distribution with a mean of 0 and a standard deviation of 0.0439, or from a uniform distribution in the range [−0.0760, 0.0760]. A layer with a smaller combined fan-in and fan-out produces a larger σ and a wider uniform range, because fewer connections share the responsibility for preserving signal variance.
Fan-In vs. Fan-Out: Why Both Input and Output Size Matter
Fan-in is the number of input connections to a layer. Fan-out is the number of output connections from a layer. PyTorch and Keras use these two terms directly as function parameters and internal variable names.
A layer with a high fan-out sends its output to many downstream units, so a small change in that layer’s weights has a larger combined effect on the next layer. A layer with a high fan-in receives signal from many upstream units, so the gradient flowing back to it during backpropagation is the sum of many contributions. Xavier initialization accounts for both values so that neither the forward signal nor the backward gradient is scaled incorrectly by a layer’s specific shape.
Agentic AI Course
Average time:6 month(s) + Lifetime Access
Skills you’ll build: Autonomous Decision-Making, Reinforcement Learning, Multi-Agent System Design, Natural Language Understanding & Generation, API Integration & Autonomous Execution, Goal-Oriented Planning & Problem Solving
How to Implement Xavier Initialization in Code
Xavier Initialization in PyTorch
PyTorch’s nn.Linear and nn.Conv2d layers do not use Xavier initialization by default; their default initializer is Kaiming uniform. Applying Xavier initialization to a layer’s weights requires an explicit call to torch.nn.init after the layer is created.
- Import torch.nn and torch.nn.init.
- Define the layer, for example nn.Linear(fan_in, fan_out).
- Call nn.init.xavier_uniform_(layer.weight) or nn.init.xavier_normal_(layer.weight) on the layer’s weight tensor.
- Set the bias separately, typically to zero, with nn.init.zeros_(layer.bias).
import torch
import torch.nn as nn
layer = nn.Linear(in_features=256, out_features=128)
nn.init.xavier_uniform_(layer.weight)
nn.init.zeros_(layer.bias)
# Normal (Gaussian) version
nn.init.xavier_normal_(layer.weight)
Xavier Initialization in TensorFlow/Keras
Keras Dense and Conv2D layers use glorot_uniform as the default kernel_initializer. Setting it explicitly makes the choice visible in the model definition rather than relying on the default.
import tensorflow as tf
from tensorflow.keras import layers
model = tf.keras.Sequential([
layers.Dense(128, activation="relu",
kernel_initializer="glorot_uniform"),
layers.Dense(64, activation="relu",
kernel_initializer=tf.keras.initializers.GlorotNormal()),
layers.Dense(10, activation="softmax",
kernel_initializer="glorot_uniform"),
])
glorot_uniform and GlorotUniform() both call the uniform form of the formula. GlorotNormal() calls the normal form.
Xavier vs. He Initialization: Which Should You Use?
Use He initialization when the network’s hidden layers use ReLU or a ReLU variant. Use Xavier initialization when the network uses sigmoid or tanh activation functions.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun, researchers at Microsoft Research, introduced He initialization in “Delving Deep into Rectifiers: Surpassing Human-Level Performance on ImageNet Classification,” presented at ICCV 2015. The formula removes the factor of 2 that Xavier initialization splits between fan-in and fan-out, and applies it to fan-in alone, because a ReLU function sets roughly half its inputs to zero and needs a larger variance to compensate.
| Initializer | Uniform Formula | Normal Formula | Designed For | Introduced |
| Xavier (Glorot) | √(6 / (fan-in + fan-out)) | σ = √(2 / (fan-in + fan-out)) | Sigmoid, tanh | Glorot & Bengio, 2010 |
| He (Kaiming) | √(6 / fan-in) | σ = √(2 / fan-in) | ReLU, Leaky ReLU | He et al., 2015 |
| LeCun | √(3 / fan-in) | σ = √(1 / fan-in) | SELU | LeCun et al., 1998 |
The variance of Xavier initialization for ReLU activation functions is too low at the start since the derivation does not take into consideration that ReLU zeros half of its input. The variance of He initialization for sigmoid/tanh activation functions is too high since the derivation does not support those activation functions.
Does Xavier Initialization Still Matter in 2026?
Introduced in 2015 by Ioffe and Szegedy, batch normalization normalizes the output of a layer and makes the network less sensitive to the specific initialization variance of its weights. Introduced in 2016 in the ResNet architecture by He et al., residual connections provide the gradient a direct way through the layers, which is also making the network less sensitive to initialization. Both techniques do not eliminate the need for proper initialization; both were invented and tested in networks that used weight initializers.
Xavier/Glorot initialization is used in the linear projections in transformer architectures, as described by Vaswani et al. in “Attention Is All You Need” (2017). Networks built on top of batch normalization, residual connections, or transformers still require a weight initializer for their linear and convolutional layers; only the network structure changes the price paid for the wrong one.
Common Mistakes When Using Xavier Initialization
- Pairing Xavier initialization with ReLU activations. The formula’s variance is calibrated for sigmoid and tanh; a ReLU network typically converges faster with He initialization instead.
- Leaving the bias uninitialized. Xavier initialization applies to weight tensors only. Bias terms are conventionally set to zero and require a separate initialization call.
- Applying Xavier initialization to a from-scratch NumPy implementation without dividing by the correct denominator. Using fan-in alone instead of (fan-in + fan-out) produces the He formula’s target variance, not Xavier’s.
- Assuming a framework’s default initializer is Xavier initialization. PyTorch’s default for nn.Linear is Kaiming uniform, not Xavier; Keras’s default for Dense is glorot_uniform. The two frameworks do not share a default.
Frequently Asked Questions
Q1. What is the difference between Xavier and He initialization?
Ans. Xavier initialization uses a variance of 2 / (fan-in + fan-out), and is appropriate for the sigmoid and tanh activation functions. The initialization uses a variance of 2 / fan-in and is appropriate for the ReLU function and its extensions, which set about half their inputs to zero.
Q2. What is normal Xavier initialization?
Ans. Normal Xavier initialization selects each weight according to a Gaussian distribution with mean 0 and standard deviation √(2 / (fan-in + fan-out)). This yields the same target variance as the uniform Xavier initialization, but from a bell curve rather than a fixed range.
Q3. Is Xavier initialization the default in Keras?
Ans. Yes. Keras Dense and Conv2D layers use glorot_uniform as the default kernel_initializer.
Q4. Does PyTorch use Xavier initialization by default?
Ans. No. PyTorch’s nn.Linear and nn.Conv2d layers default to Kaiming uniform initialization. Xavier initialization requires an explicit call to torch.nn.init.xavier_uniform_ or torch.nn.init.xavier_normal_.
Q5. Should I use uniform or normal Xavier initialization?
Ans. They yield the same target variance and behave equivalently in practice. Xavier uniform initialization is the more frequent default setting in deep learning libraries, while Xavier normal initialization can be used as a drop-in replacement in case of a preferred Gaussian distribution for a certain layer.
Q6. Does batch normalization remove the need for Xavier initialization?
Ans. No. The batch normalization method normalizes the outputs of a layer after the application of the weights, hence making it less dependent on the initial variance. However, all layers still require some initial weight value before the first forward propagation through them, which is provided by Xavier/He initialization.
Choosing an Initializer
The Xavier initializer is used in case the layer uses a sigmoid or tanh activation function. The initialization is used in case the layer uses a ReLU or any other variant of the ReLU activation function. The initialization method should be used on a per-layer basis and not on a per-network basis, even when the activation functions are different in various layers.