Skip to main content
L2.12

Initialization

Goal

By the end of this lesson, you can explain why hidden units need symmetry-breaking initialization, show how initialization scale changes activations, and diagnose suspicious starting statistics before training.

Training begins from the parameters you choose​

Before the first forward pass, every trainable parameter already has a value.

Those starting values matter because they determine:

  • the first activations;
  • the first loss;
  • the first gradients;
  • whether different hidden units begin with distinguishable behavior.

Initialization is therefore part of the training setup, not a cosmetic detail.

Why identical hidden units are a problem​

Imagine two hidden neurons with exactly the same:

  • incoming weights;
  • bias;
  • inputs;
  • update rule.

They compute the same output.

Because their forward behavior is identical, they also receive identical gradient information in this symmetric setup. After one update, they remain identical. The pattern can continue.

Two neurons are present, but they behave like copies of one neuron.

Small random differences break this symmetry so the units can learn different roles.

Random does not mean any scale is safe​

Suppose weights are random but extremely large.

Then weighted sums can become very large on the first forward pass. Depending on the activation, this can lead to:

  • saturated sigmoid/tanh values;
  • large or unstable outputs;
  • unhelpful gradient scales.

Weights that are far too small can also make signals shrink through depth.

So initialization has two jobs:

  1. break symmetry;
  2. keep signal magnitudes in a useful range.

Compare zero and random initialization​

The Lab sends one batch through two hidden units.

  1. Click Run.
  2. Compare the two pre-activation tables. zero-init pre-activations is all 0.0. random-init pre-activations starts with [0.055, 0.238], so the two hidden units already compute different values.
  3. Confirm zero columns identical: True. With zero weights, both hidden units receive identical inputs, identical gradients, and would stay identical during training.
  4. Find W_random = rng.normal(0.0, 0.2, size=(2, 2)). The 0.2 is the scale (standard deviation) of the random weights. Change only 0.2 to 2.0.
  5. Before running, predict what happens to the size and spread of the random pre-activations.
  6. Click Run. random weight std grows from 0.0878 to 0.8778, and every pre-activation becomes ten times larger (the first row becomes [0.551, 2.379]). The pattern is the same because the seed is the same; only the scale changed. Large starting values like these can push activations such as tanh into their flat, saturated regions.
  7. Press Reset afterward.

Loading lab…

A fixed random seed makes this comparison reproducible. It does not mean that the seeded values are “the correct weights.”

Named initialization schemes come from signal-scale reasoning​

Schemes such as Xavier/Glorot and He initialization choose weight variance based on how many connections enter or leave a layer.

The names are less important than the reason:

try to keep activations and gradients from systematically exploding or vanishing as they pass through layers.

Framework helpers automate the calculation after you understand the goal.

What to inspect when training looks wrong immediately​

If the very first batch already has strange values, inspect before changing architecture:

  • mean and standard deviation of initial weights;
  • minimum/maximum weights;
  • pre-activation statistics;
  • activation statistics;
  • whether hidden units are suspiciously identical;
  • input scale.

This is faster than waiting many epochs to discover that the network started in a numerically poor state.

Quick Check

1. Why is identical hidden-unit initialization risky?
2. Why use a fixed random seed in a lesson Lab?
3. What should you inspect if activations explode immediately?

0 of 3 questions answered.

Key Takeaways

  • Initialization determines the network's starting activations and gradients.
  • Identical hidden units can remain redundant because of symmetry.
  • Random differences break symmetry, but scale still matters.
  • Named initialization schemes aim to keep signal magnitudes usable through depth.
  • Inspect starting statistics when training is unhealthy from the first step.

Next Lesson

Next, you will change the training objective itself by adding a preference against unnecessarily large parameter values.

References

Lesson actions

Completion is stored locally on this device.

View progress