본문으로 건너뛰기
L5.0

Level 5 — Attention and the Transformer

Level 4 turned text into token representations. Level 5 asks what happens when one token needs information from other tokens in the sequence.

Start with one sentence​

Consider:

the cat chased the toy because it rolled

When the model is processing it, some earlier words are more useful than others. toy may be especially relevant. Treating every earlier position exactly the same would throw away that difference.

Attention is a way for a model to build a new representation by mixing information from other allowed positions with different weights.

A simple picture is:

possible context pieces
↓
assign larger or smaller weights
↓
combine their information
↓
new representation for this position

The weights are not permanently fixed. They depend on the representations being compared, so different positions can gather different information.

That is the main idea. The rest of this Level explains how those weights are calculated and how attention becomes a Transformer block.

A few words you will meet​

TermPlain meaning in this Level
attention scorea number used to compare how strongly one position should consider another
attention weightthe normalized amount of influence assigned to a position; a row of weights usually sums to 1
querythe representation asking “what information is useful for me?”
keythe representation used to decide how well another position matches that request
valuethe information that is actually mixed after the weights are chosen
causal maska rule that blocks a position from using future tokens when future information is not allowed
attention headone attention calculation with its own learned projections; multiple heads let the model form several mixtures in parallel
residual connectionadding a block's input back to its transformed output so the original stream is preserved while new information is added
Transformer blocka repeated unit that combines attention with additional per-position processing and normalization

You will not need all of these at once. The lessons introduce them in order.

Learning goal​

Build causal self-attention and a small Transformer from pieces you can inspect and test.

By the end of this Level, you should be able to:

  • explain why attention is useful when different positions need different context;
  • explain the distinct query, key, and value roles;
  • compute a small dot-product attention row by hand;
  • explain why very large comparison scores are scaled before softmax;
  • show that blocked future positions receive zero attention probability;
  • split and recombine multiple attention heads without changing the overall sequence representation size;
  • explain what residual connections, layer normalization, and the feed-forward part each contribute;
  • trace one Transformer block from input to output;
  • stack blocks while preserving shape and the rule against future access;
  • assemble a mini decoder-style Transformer that produces one raw vocabulary score per possible next token.

When this Level uses the word invariant, it means a property that should stay true while other things change. For example, after attention weights are normalized, each valid row should still sum to 1. If an invariant breaks, it gives you a precise place to debug.

The learning path​

L5.1 — Why Attention begins with the weighted-mixture idea using ordinary words and small numbers.

L5.2–L5.7 build the scoring and mixing process: words looking at words, query/key/value roles, dot products, scaling, attention weights, and self-attention.

L5.8 — Causal Masking adds the rule that a next-token model must not use future positions. After this lesson, complete the Level 5 mini checkpoint.

L5.9–L5.14 build multi-head attention and then add residual connections, normalization, feed-forward computation, and a complete Transformer block.

L5.15–L5.16 stack the blocks and assemble a mini Transformer from token IDs to vocabulary scores.

How to debug attention​

A full Transformer can feel complicated because several small calculations are connected. Debug from the smallest visible boundary outward:

shapes
→ comparison scores
→ scaling
→ mask
→ normalized weights
→ weighted values
→ multiple heads
→ residual stream

For example, if a future token receives a non-zero weight in a causal model, stop there. There is no reason to inspect the final language-model loss until the mask behavior is correct.

Level Project​

After L5.16 — Build a Mini Transformer, complete Build a Mini Transformer.

The project combines causal multi-head attention, the residual/normalization/feed-forward block, shape preservation, and one intentional failure/debug path. The goal is to make each important boundary explainable rather than treating the Transformer as one large formula.

Lesson actions

Completion is stored locally on this device.

View progress