Attention Weights and Weighted Values
Goal
Normalize attention scores and compute a weighted value mixture while checking row-sum and output-shape invariants.
Suppose two values are [2,0] and [0,4] with weights [0.75,0.25]:
0.75*[2,0] + 0.25*[0,4] = [1.5, 1.0]
The score stage answered where to take information from. The value stage determines what content is mixed.
See each weighted value contribution
The numbers match the worked example above. The output is the sum of the two visible weighted contributions.
| Candidate | Weight | Value | Weighted contribution |
|---|---|---|---|
| Value A | 75.0% | [2, 0] | [1.50, 0] |
| Value B | 25.0% | [0, 4] | [0, 1] |
| Weighted sum | [1.50, 1] | ||
The output can land between several values
A weighted mixture does not have to copy one source position. With values
A = [4, 0]
B = [0, 4]
C = [4, 4]
and weights [0.5, 0.25, 0.25], the output is
0.5A + 0.25B + 0.25C
= [2,0] + [0,1] + [1,1]
= [3,2]
There is no source value [3,2]. Attention has created a new mixture from existing value representations.
This also explains why reading only the largest attention cell can be misleading. A row such as [0.45, 0.35, 0.20] has a largest entry, but more than half of the total mass still comes from other positions. To understand the actual attention output, you need both the whole weight row and the value vectors being mixed.
A useful invariant is: changing only the values cannot change a correctly computed softmax weight row, but it can change the output vector immediately.
A normalized weight row is a routing decision for one query
For one query position, softmax might produce:
[0.10, 0.70, 0.20]
Those numbers answer “how much of each value vector enters this query's attention output?”
A different query position gets its own score row and therefore its own weight row.
So an attention matrix is not one global ranking of important words. It contains a separate routing distribution for each query.
Also remember that a large weight multiplies a value vector, not the raw token itself. The value may already contain features created by earlier layers.
This is why an attention heatmap is useful evidence but not a complete explanation of a model prediction. The final behavior depends on value contents, other heads, residual paths, later MLPs, and later blocks.
Predict
Implement the weighted mixture
This Browser Lab now leaves the final mixing step for you to write.
- Run the starter unchanged. The function validates shapes but returns zeros at the TODO, so the expected-mixture checks fail.
- For each feature index
j, compute the sum ofweight * value[j]across all value vectors. - Keep the existing shape checks; replace only the TODO return expression.
- Run again. For weights
[0.75, 0.25]and values[[2,0],[0,4]], the output should be[1.5, 1.0]. - Confirm the one-hot case
[0,1]selects the second value exactly. - Then change the weights to another pair that sums to 1 and predict the new output before running.
Loading lab…
If you manually use weights that do not sum to one, arithmetic still produces a vector, but the normalized-mixture interpretation is no longer the same. This distinction will matter when masking is introduced.
Quick Check
Explain it back
Explain why an attention heatmap alone is not the output of attention. The weights still need to be applied to value vectors.
Key Takeaways
- Softmax turns scores into normalized routing weights.
- Values carry the content being routed.
- Legal rows should be non-negative and sum to one.
- A concentrated weight row makes the output resemble the dominant value.
Next Lesson
Next, L5.7 — Self-Attention applies this complete lookup to one sequence with self-attention.
References
- Vaswani et al., Attention Is All You Need.
Completion is stored locally on this device.