Initialization
Goal
By the end of this lesson, you can explain why hidden units need symmetry-breaking initialization, show how initialization scale changes activations, and diagnose suspicious starting statistics before training.
Training begins from the parameters you choose
Before the first forward pass, every trainable parameter already has a value.
Those starting values matter because they determine:
- the first activations;
- the first loss;
- the first gradients;
- whether different hidden units begin with distinguishable behavior.
Initialization is therefore part of the training setup, not a cosmetic detail.
Why identical hidden units are a problem
Imagine two hidden neurons with exactly the same:
- incoming weights;
- bias;
- inputs;
- update rule.
They compute the same output.
Because their forward behavior is identical, they also receive identical gradient information in this symmetric setup. After one update, they remain identical. The pattern can continue.
Two neurons are present, but they behave like copies of one neuron.
Small random differences break this symmetry so the units can learn different roles.
Random does not mean any scale is safe
Suppose weights are random but extremely large.
Then weighted sums can become very large on the first forward pass. Depending on the activation, this can lead to:
- saturated sigmoid/tanh values;
- large or unstable outputs;
- unhelpful gradient scales.
Weights that are far too small can also make signals shrink through depth.
So initialization has two jobs:
- break symmetry;
- keep signal magnitudes in a useful range.
Compare zero and random initialization
The Lab sends one batch through two hidden units.
- Click Run.
- Compare the two pre-activation tables.
zero-init pre-activationsis all0.0.random-init pre-activationsstarts with[0.055, 0.238], so the two hidden units already compute different values. - Confirm
zero columns identical: True. With zero weights, both hidden units receive identical inputs, identical gradients, and would stay identical during training. - Find
W_random = rng.normal(0.0, 0.2, size=(2, 2)). The0.2is the scale (standard deviation) of the random weights. Change only0.2to2.0. - Before running, predict what happens to the size and spread of the random pre-activations.
- Click Run.
random weight stdgrows from0.0878to0.8778, and every pre-activation becomes ten times larger (the first row becomes[0.551, 2.379]). The pattern is the same because the seed is the same; only the scale changed. Large starting values like these can push activations such as tanh into their flat, saturated regions. - Press Reset afterward.
Loading lab…
A fixed random seed makes this comparison reproducible. It does not mean that the seeded values are “the correct weights.”
Named initialization schemes come from signal-scale reasoning
Schemes such as Xavier/Glorot and He initialization choose weight variance based on how many connections enter or leave a layer.
The names are less important than the reason:
try to keep activations and gradients from systematically exploding or vanishing as they pass through layers.
Framework helpers automate the calculation after you understand the goal.
What to inspect when training looks wrong immediately
If the very first batch already has strange values, inspect before changing architecture:
- mean and standard deviation of initial weights;
- minimum/maximum weights;
- pre-activation statistics;
- activation statistics;
- whether hidden units are suspiciously identical;
- input scale.
This is faster than waiting many epochs to discover that the network started in a numerically poor state.
Quick Check
Key Takeaways
- Initialization determines the network's starting activations and gradients.
- Identical hidden units can remain redundant because of symmetry.
- Random differences break symmetry, but scale still matters.
- Named initialization schemes aim to keep signal magnitudes usable through depth.
- Inspect starting statistics when training is unhealthy from the first step.
Next Lesson
Next, you will change the training objective itself by adding a preference against unnecessarily large parameter values.
References
- PyTorch, nn.init.
Completion is stored locally on this device.