Level 5 — Attention and the Transformer
Level 4 turned text into token representations. Level 5 asks what happens when one token needs information from other tokens in the sequence.
Start with one sentence
Consider:
the cat chased the toy because it rolled
When the model is processing it, some earlier words are more useful than others. toy may be especially relevant. Treating every earlier position exactly the same would throw away that difference.
Attention is a way for a model to build a new representation by mixing information from other allowed positions with different weights.
A simple picture is:
possible context pieces
↓
assign larger or smaller weights
↓
combine their information
↓
new representation for this position
The weights are not permanently fixed. They depend on the representations being compared, so different positions can gather different information.
That is the main idea. The rest of this Level explains how those weights are calculated and how attention becomes a Transformer block.
A few words you will meet
| Term | Plain meaning in this Level |
|---|---|
| attention score | a number used to compare how strongly one position should consider another |
| attention weight | the normalized amount of influence assigned to a position; a row of weights usually sums to 1 |
| query | the representation asking “what information is useful for me?” |
| key | the representation used to decide how well another position matches that request |
| value | the information that is actually mixed after the weights are chosen |
| causal mask | a rule that blocks a position from using future tokens when future information is not allowed |
| attention head | one attention calculation with its own learned projections; multiple heads let the model form several mixtures in parallel |
| residual connection | adding a block's input back to its transformed output so the original stream is preserved while new information is added |
| Transformer block | a repeated unit that combines attention with additional per-position processing and normalization |
You will not need all of these at once. The lessons introduce them in order.