Words Looking at Words
Goal
Explain attention as a weighted information lookup between token positions and read one normalized attention row without needing to implement the full attention mechanism yet.
A token often needs context from other positions. In the sentence:
the animal did not cross the road because it was tired
it is easier to interpret if information from animal receives more weight than unrelated words.
Attention gives each query position a mixture of information from legal context positions. The important new idea in this lesson is the mixture itself:
output = weight1 × value1 + weight2 × value2 + ...
The weights are non-negative and, after normalization, add up to 1 across the positions that are allowed to contribute.
Read one attention row
A tiny decoder attention map
Select a row and read it as a distribution over context positions. Future cells are already hidden here; L5.8 — Causal Masking will explain how that no-future rule is implemented.
For the row belonging to predicts, the visible weights are 0.15, 0.45, and 0.40. They add to 1.0.
That means the output at this position is built from a weighted mixture rather than one fixed average of all context.
Predict
Use the Lab only to inspect weighting behavior
The browser Lab contains scores = [0.2, 1.0, 0.4, 2.0] and prints one normalized row for each query position. The starter already handles the decoder's future restriction. You do not need to understand the masking code yet; L5.8 — Causal Masking will build and test that no-future rule directly.
- Click Run unchanged.
- Find the line beginning with
2. That row corresponds to query index2, so only score positions0,1, and2are legal and the fourth weight is0.0. - Confirm the printed
sum=is1.0. - In
scores, change only the first value from0.2to2.0:
scores = [2.0, 1.0, 0.4, 2.0]
- Before running, predict that query row
2should give more weight to position0, because that legal score became larger. - Click Run and compare only row
2with the first run. Confirm the legal weights still sum to1.0and the fourth future position stays at0.0. - Restore the first score to
0.2.
Loading lab…
This is enough for now: scores change → normalized weights change → the weighted information mixture changes.
What you are not expected to know yet
A full Transformer computes those scores from learned vector roles called queries, keys, and values. It also scales scores, applies a causal mask for decoder models, and may split feature channels across several heads.
Those are not prerequisites for this lesson. The next lessons introduce them one at a time:
- L5.3 — Queries, Keys, and Values
- L5.4 — Dot Products as Similarity
- L5.5 — Scaled Dot-Product Attention
- L5.6 — Attention Weights and Weighted Values
- L5.8 — Causal Masking
- L5.9 — Multi-Head Attention
Seeing the names now is a map, not an instruction to understand the implementation early.
Attention weights are evidence, not a complete explanation
A weight row is useful because it shows where one attention operation routed information. It does not prove that a token or head contains a complete human-readable reason for the model's final output. Later layers, value directions, residual connections, and nonlinear transformations also matter.
Quick Check
Key Takeaways
- Attention builds a weighted mixture of information from context positions.
- A normalized legal weight row sums to
1. - Changing one score changes the distribution of weight and therefore the mixture.
- The Lab previews a decoder restriction, but masking and multi-head implementation are taught later rather than assumed here.
- Attention weights are useful evidence, not a complete account of model behavior.
Next Lesson
Next, you will make the scoring rule concrete with L5.3 — Queries, Keys, and Values, separating what a position asks for, what each position advertises for matching, and what information is actually carried forward.
References
- Vaswani et al., Attention Is All You Need.
- PyTorch softmax documentation: https://pytorch.org/docs/stable/generated/torch.nn.functional.softmax.html
Completion is stored locally on this device.