Skip to main content
L5.13

Putting the Pieces Together

Goal

Trace a pre-norm Transformer block boundary by boundary and verify shape, causal, normalization, and residual invariants before treating it as one module.

You now have the component ideas needed for a block. Follow the residual stream rather than memorizing a diagram:

x = x + attention(norm1(x))
x = x + mlp(norm2(x))

The first branch can mix legal token positions through causal attention. The second transforms features independently at each position. Both return to (B,T,C) before residual addition.

Step through the two residual updates

Keep the full block visible while moving one stage at a time. The active stage names the invariant you should check before moving deeper into the implementation.

Stage 1/8 · Residual input
Residual inputx · (B,T,C)Starts both the skip path and the attention branch.
Skip pathxKeep the current stream unchanged.
Norm 1(B,T,C)
Causal attentionupdate · (B,T,C)
Residual add 1x1 = x + attention(norm1(x))unchanged x + attention update
Skip pathx1Keep the updated stream unchanged.
Norm 2(B,T,C)
MLPupdate · (B,T,C)
Residual add 2output = x1 + mlp(norm2(x1))unchanged x1 + MLP update
Block output(B,T,C)Same interface, after two learned updates.

Trace one token through the block without losing the sequence view​

At block input, token position 2 already contains a feature vector shaped by earlier layers. The first branch normalizes that vector, then causal attention lets position 2 mix information from legal positions 0, 1, and 2. The attention update is added back to the original stream.

The second branch starts from this updated stream. LayerNorm prepares each token's features, then the MLP transforms those features independently at every position. Its update is also added back.

That gives a useful division of labor:

attention branch: "which legal positions should contribute information?"
MLP branch: "how should this position's current features be transformed?"
residual path: "keep a direct route while adding each learned correction"

A correct final (B,T,C) shape only proves the interfaces line up. It does not prove causality, normalization axis, or Q/K/V semantics are correct. This is why the lesson asks you to record evidence at intermediate boundaries rather than checking only the block output.

When reading an implementation, locate the two residual additions first. Then open each branch and inspect its internal invariants. This top-down method is easier than reading every matrix multiplication in source-code order because it keeps the purpose of each operation visible.

Predict

Which branch directly mixes information across token positions?

Trace the notebook at named boundaries​

Open the notebook and run both code cells.

  1. The first cell builds one TinyBlock with n_embd=12 and n_head=3, sends a batch x of shape (2, 4, 12) through it, and should print PASS: l05-13 block composition checks.
  2. The second cell prints one line per boundary: block input, normalized input, attention update, first residual, second normalized input, MLP update, and block output. Copy the seven shapes into your notes. Each one should be (2, 4, 12).
  3. Explain why every row has the same shape. (Hint: each update is added back to the residual stream.)
  4. Try one controlled break: in the first cell, change the last layer of self.ff from nn.Linear(4*n_embd, n_embd) to nn.Linear(4*n_embd, n_embd - 2) and rerun. The error appears at the second residual addition—the first boundary where widths must match. Restore the original line afterward.

Loading lab…

Do not diagnose “the Transformer block” as one object. If an invariant fails, identify the first boundary: normalization axis, attention causality/head shape, residual addition, or MLP width.

Quick Check

1. What end-to-end shape should one block preserve?
2. Which invariant must survive the attention branch?
3. What should the MLP branch preserve before addition?

0 of 3 questions answered.

Explain it back​

Explain the block as two learned residual updates rather than a list of class names.

Key Takeaways

  • A block is understandable as tested composition of smaller contracts.
  • Causal attention mixes positions; the MLP transforms features per position.
  • Both branches preserve the residual-stream interface.
  • Debug the first broken boundary, not the full block by guesswork.

Next Lesson

Next, L5.14 — Transformer Block consolidates this composition as a reusable Transformer block.

References

Lesson actions

Completion is stored locally on this device.

View progress