Residual Connections
Goal
Explain residual connections as shape-matched learned updates and identify the exact shape contract required at an addition boundary.
A Transformer sublayer usually does not replace the stream from scratch. It returns an update:
new_x = x + update(x)
For x=[1,2] and update [0.1,-0.2], the result is [1.1,1.8]. This small example captures the key idea: the original representation has a direct path, while a learned branch adds a correction.
The addition also creates a powerful interface invariant. If x has (B,T,C), the update branch must return the same shape.
Why the direct path matters
Suppose a sublayer has not yet learned a useful correction. With a residual connection, an update near zero gives:
new_x = x + 0 ≈ x
So the block can initially behave close to an identity mapping instead of being forced to rebuild the whole representation at every depth.
The same idea helps gradients. The computation graph contains a direct additive path from later layers back to the earlier stream. That does not make deep optimization automatically easy, but it gives learning a simpler route than requiring every signal to pass only through a long chain of nonlinear transformations.
Do not confuse this with “residual connections preserve the original vector unchanged.” The stream is repeatedly updated. After several blocks, the values can be very different from the starting embedding even though every update was added through an identity path.
A useful debugging question is therefore not “did x stay the same?” but “did the branch return an update with the right shape and reasonable numerical scale before addition?”
Residual addition creates a common interface between sublayers
Attention and the feed-forward network perform different computations, but both usually return an update with the same width as the residual stream.
That gives a stable pattern:
stream shape: (B,T,C)
branch update: (B,T,C)
add → new stream: (B,T,C)
The preserved interface makes it possible to stack many different sublayers without changing the outer shape at every step.
Watch the scale of the update
Even when shapes match, a branch can still behave badly if its update becomes extremely large compared with the incoming residual stream.
Suppose one feature in x is about 0.5 but the branch suddenly returns 200. The addition is legal, yet the branch overwhelms the previous representation.
Normalization, initialization, optimization, and architecture design all help keep these updates trainable.
This is another recurring lesson in neural-network debugging: shape correctness is necessary, but numerical scale is a separate invariant.
Predict
Inspect the boundary, not the whole network
- Click Run. Compare the three lines:
x: [1.0, 2.0, -1.0],update: [0.1, -0.2, 0.5], andresidual: [1.1, 1.8, -0.5]. With a small update, the output stays close to the originalx. - Make the update large. Change
update = [0.1, -0.2, 0.5]toupdate = [10.0, -20.0, 50.0]. - Click Run. The residual becomes
[11.0, -18.0, 49.0], now dominated by the update. The skip path still carriesx, but a very large branch can drown it out. - Now make the update the wrong width:
update = [0.1, -0.2]. - Click Run. The Lab stops with
ValueError: residual branches must have matching shapes. The failure appears exactly at the addition boundary, which is where you should look first. - Press Reset afterward.
Loading lab…
When a residual error appears in a deep model, print the two shapes immediately before +. Fix that contract before changing optimization or attention logic.
Quick Check
Transfer the idea
Explain why attention and the feed-forward network both need to return to n_embd before their outputs can update the same residual stream.
Key Takeaways
- Residual connections add learned corrections to an existing representation.
- Residual branches enforce a strong shape contract.
- A valid residual shape does not prove the branch semantics are correct, but an invalid shape localizes the failure quickly.
Next Lesson
Next, normalize each token's feature vector before sublayers so the scale presented to repeated transformations is controlled.
References
- Vaswani et al., Attention Is All You Need.
- He et al., Deep Residual Learning for Image Recognition.
Completion is stored locally on this device.