Why Attention
Goal
By the end of this lesson, you can explain why a token may need different amounts of information from different earlier positions, and how attention represents that choice with weights.
Start with one sentence
Consider:
the cat chased the toy because it rolled
Suppose we are building a new representation for the token it.
The earlier words do not all seem equally useful. In this sentence, toy may be more relevant than the or because for understanding what it refers to.
One simple but weak strategy would be to average information from every earlier position equally. If three pieces received equal weights, we might have:
weight on piece A = 1/3
weight on piece B = 1/3
weight on piece C = 1/3
A weight is just a number saying how much influence one piece has in the mixture.
Attention improves on the fixed average by calculating different weights from the current representations. A more relevant position can receive a larger weight.
That is what attention means at this stage of the course:
build a new representation by mixing information from other allowed positions, using larger or smaller weights depending on the current content.
We say allowed positions because some tasks place rules on which positions may be used. In the next-token models you will build later, a position may use itself and earlier tokens but not future tokens. Level 5 will introduce that causal rule explicitly in L5.8 — Causal Masking.
What gets mixed?
Use three simple numbers as stand-ins for information stored at three positions:
values = [1, 5, 9]
weights = [0.2, 0.3, 0.5]
The weighted mixture is:
1×0.2 + 5×0.3 + 9×0.5 = 6.2
The numbers being mixed are called values in attention. Later, a value will be a vector rather than one number, but the idea is the same.
The weights in a valid attention row are non-negative and normally add up to 1, so the output is a mixture of the available values.
From a fixed summary to a different summary for each position
The important contrast is not “attention versus no context.” Earlier sequence models can carry context too. The difference is that attention can build a different context mixture for each query position.
Suppose the earlier value vectors are simplified to three numbers:
cat → 8
toy → 2
rolled → 6
A fixed average always gives (8 + 2 + 6) / 3 = 5.33, no matter which current token is asking for context. But two query positions can need different evidence:
query A weights: [0.8, 0.1, 0.1] → 7.2
query B weights: [0.1, 0.7, 0.2] → 3.4
The values did not change. The routing weights changed because the query changed. That is the useful mental model to carry forward.
A common misconception is that a large attention weight means “this word caused the final answer.” It means something narrower: in this attention operation, more of that position's value representation is routed into the current mixture. Later layers and residual paths still transform the result.
Predict
Compare equal and focused weights in the Lab
The Lab contains two score patterns. Softmax turns each score pattern into weights that add up to 1. You will study softmax more closely later; for now, read it as the step that converts comparison scores into usable mixture weights.
- Click Run without editing anything.
- Read
uniform weights:. The starting scores[0.0, 0.0, 0.0]give all three positions equal weights. - Read
focused weights:. The scores[2.0, 0.0, 0.0]give the first position a larger weight. - The values are
[1.0, 5.0, 9.0]. Compareuniform output:withfocused output:and notice that putting more weight on the first value pulls the output toward1.0. - Find:
focused = softmax([2.0, 0.0, 0.0])
- Change only
2.0to4.0. - Before running, predict that the first focused weight will become even larger and
focused output:will move even closer to1.0. - Click Run and compare the new focused weights and output with the first run.
- Restore
2.0before moving on.
Loading lab…
The useful evidence is not just the final output. Inspect the weight row itself:
- Are all weights non-negative?
- Do they add up to about
1? - When one score becomes much larger, does its weight become larger too?
These checks make the mechanism visible before it is placed inside a larger Transformer.
Quick Check
Explain it back
Explain attention without saying only “the model looks at words.” Your explanation should name:
- the information available at several positions;
- the weights that decide how much each position contributes;
- the weighted mixture produced for the current position.
Key Takeaways
- Attention mixes information from several allowed positions.
- A weight says how much influence one position has in that mixture.
- Equal weights form an average and cannot express relevance differences.
- Attention computes weights from the current representations rather than fixing them forever.
- Inspect the weight row before blaming a much later model output.
Next Lesson
Next, in L5.2 — Words Looking at Words, you will make this weighted-mixture idea concrete with a small sequence and see how one position can assign different weights to other positions.
References
- Vaswani et al., Attention Is All You Need.
Completion is stored locally on this device.