Level 5 — Attention and the Transformer
Level 4 turned text into token representations. Level 5 asks what happens when one token needs information from other tokens in the sequence.
Start with one sentence
Consider:
the cat chased the toy because it rolled
When the model is processing it, some earlier words are more useful than others. toy may be especially relevant. Treating every earlier position exactly the same would throw away that difference.
Attention is a way for a model to build a new representation by mixing information from other allowed positions with different weights.
A simple picture is:
possible context pieces
↓
assign larger or smaller weights
↓
combine their information
↓
new representation for this position
The weights are not permanently fixed. They depend on the representations being compared, so different positions can gather different information.
That is the main idea. The rest of this Level explains how those weights are calculated and how attention becomes a Transformer block.
A few words you will meet
| Term | Plain meaning in this Level |
|---|---|
| attention score | a number used to compare how strongly one position should consider another |
| attention weight | the normalized amount of influence assigned to a position; a row of weights usually sums to 1 |
| query | the representation asking “what information is useful for me?” |
| key | the representation used to decide how well another position matches that request |
| value | the information that is actually mixed after the weights are chosen |
| causal mask | a rule that blocks a position from using future tokens when future information is not allowed |
| attention head | one attention calculation with its own learned projections; multiple heads let the model form several mixtures in parallel |
| residual connection | adding a block's input back to its transformed output so the original stream is preserved while new information is added |
| Transformer block | a repeated unit that combines attention with additional per-position processing and normalization |
You will not need all of these at once. The lessons introduce them in order.
Learning goal
Build causal self-attention and a small Transformer from pieces you can inspect and test.
By the end of this Level, you should be able to:
- explain why attention is useful when different positions need different context;
- explain the distinct query, key, and value roles;
- compute a small dot-product attention row by hand;
- explain why very large comparison scores are scaled before softmax;
- show that blocked future positions receive zero attention probability;
- split and recombine multiple attention heads without changing the overall sequence representation size;
- explain what residual connections, layer normalization, and the feed-forward part each contribute;
- trace one Transformer block from input to output;
- stack blocks while preserving shape and the rule against future access;
- assemble a mini decoder-style Transformer that produces one raw vocabulary score per possible next token.
When this Level uses the word invariant, it means a property that should stay true while other things change. For example, after attention weights are normalized, each valid row should still sum to 1. If an invariant breaks, it gives you a precise place to debug.
The learning path
L5.1 — Why Attention begins with the weighted-mixture idea using ordinary words and small numbers.
L5.2–L5.7 build the scoring and mixing process: words looking at words, query/key/value roles, dot products, scaling, attention weights, and self-attention.
L5.8 — Causal Masking adds the rule that a next-token model must not use future positions. After this lesson, complete the Level 5 mini checkpoint.
L5.9–L5.14 build multi-head attention and then add residual connections, normalization, feed-forward computation, and a complete Transformer block.
L5.15–L5.16 stack the blocks and assemble a mini Transformer from token IDs to vocabulary scores.
How to debug attention
A full Transformer can feel complicated because several small calculations are connected. Debug from the smallest visible boundary outward:
shapes
→ comparison scores
→ scaling
→ mask
→ normalized weights
→ weighted values
→ multiple heads
→ residual stream
For example, if a future token receives a non-zero weight in a causal model, stop there. There is no reason to inspect the final language-model loss until the mask behavior is correct.
Level Project
After L5.16 — Build a Mini Transformer, complete Build a Mini Transformer.
The project combines causal multi-head attention, the residual/normalization/feed-forward block, shape preservation, and one intentional failure/debug path. The goal is to make each important boundary explainable rather than treating the Transformer as one large formula.
Completion is stored locally on this device.