Layer Normalization
Goal
Normalize one token's feature vector, identify the feature axis in (B,T,C), and explain why normalization changes scale without changing residual-stream shape.
Give the next sublayer a controlled feature scale
A Transformer keeps updating each token's feature vector. Two token vectors can carry a similar relative pattern while arriving with very different common offsets or spreads. In the pre-norm Transformer blocks used in this course, LayerNorm puts that one token's feature values into a more standardized numerical range before the next attention or MLP branch processes them. Other Transformer families can place normalization differently.
For example, take one token vector [101,102,103]. Subtracting its feature mean removes the large shared offset; dividing by feature spread controls scale. The relative pattern remains visible.
Only after that idea is clear do we need the axis notation: in a residual stream (B,T,C), LayerNorm acts per token over the C/features dimension. It does not mix time positions.
Work one tiny example conceptually
For token features [10, 12, 14], the mean is 12. Centering gives:
[-2, 0, 2]
LayerNorm then divides by a measure of feature spread and finally applies learned scale/shift parameters in the real module. The exact normalized numbers are less important here than the invariants:
- the three features belong to one token position;
- their common offset is removed before learned affine adjustment;
- output width stays three;
- no neighboring token is required to compute this token's normalization statistics.
This last point separates LayerNorm from operations that mix examples or time positions. If token 7 changes, the normalization statistics for token 3 should not change simply because both tokens are in the same sequence.
Another misconception is that normalization “erases magnitude information forever.” The learned affine parameters and later residual updates can represent useful scales. LayerNorm standardizes the input to a sublayer; it does not permanently force every hidden state to one fixed distribution.
Practical implementations also add a small epsilon to the variance term so division stays stable when the feature variance is zero or extremely small.
Predict
Make the axis explicit
- Click Run. The Lab normalizes three feature vectors for one token each:
[1.0, 2.0, 3.0],[101.0, 102.0, 103.0](same pattern shifted up by 100), and[10.0, 20.0, 30.0](same pattern scaled by 10). - Compare the outputs. All three become
[-1.2247, 0.0, 1.2247]withmean 0.0. LayerNorm removed the shared offset and the shared scale, but kept the pattern inside the vector. - Add a fourth vector with the same values in reverse order. Change the loop line so the tuple ends with
[3.0, 2.0, 1.0]:
for values in ([1.0, 2.0, 3.0], [101.0, 102.0, 103.0], [10.0, 20.0, 30.0], [3.0, 2.0, 1.0]):
- Before running, predict its normalized output.
- Click Run. It becomes
[1.2247, 0.0, -1.2247]: the order of features inside the vector still matters after normalization. - Press Reset afterward.
Loading lab…
Then deliberately normalize across time positions instead of features in a copy. The code may still return valid numbers; the semantic operation is wrong. Always label dimensions with their meaning before choosing an axis.
Quick Check
Explain it back
Explain why LayerNorm is not a mechanism for “looking across tokens.” Its normalization group is the feature vector of each position.
Key Takeaways
- LayerNorm controls per-token feature scale.
- It preserves
(B,T,C)and does not directly mix time positions. - Axis mistakes can be semantic bugs even when shapes remain valid.
Next Lesson
Next, apply a nonlinear feature transformation independently at every token position.
References
- Ba, Kiros, and Hinton, Layer Normalization.
Completion is stored locally on this device.