Queries, Keys, and Values
Goal
Distinguish query, key, and value roles and trace how changing a key affects scores while changing a value affects retrieved content.
Attention separates two questions that are easy to mix up:
- How relevant is this position? — compare a query with keys.
- What information should be retrieved? — mix value vectors using the resulting weights.
A useful analogy is a search system: query = request, key = index description, value = content returned after matching. The analogy is imperfect, but it keeps the roles separate.
Trace one lookup with tiny vectors
Use one query and two candidate positions:
query q = [1, 0]
key A = [1, 0] value A = [10, 0]
key B = [0, 1] value B = [0, 10]
Before softmax, the dot-product scores are 1 for A and 0 for B. So A should receive more normalized weight, and the output should lean toward [10, 0].
Now imagine changing value A to [0, 20] while leaving its key unchanged. The score for A should stay the same, because scoring uses the query and key. But the retrieved content changes because the value changed.
That separation gives you a powerful debugging test:
- key-only edit → score can change;
- value-only edit → score should not change;
- either edit can change the final mixed output for different reasons.
Do not memorize Q/K/V as three mysterious letters. Ask two ordinary questions: what decides relevance? Q and K. What information is carried after relevance is decided? V.
Keep key and value roles separate
Select a candidate position. Its key is compared with the query; its value is the content that can be mixed after weighting.
[1, 0]asks what is useful now[1, 0]is compared with the query[10, 0]carries content if this position receives weightQ, K, and V are learned views of the same hidden state
In self-attention, each position begins with one hidden vector x, but learned projection matrices create different views:
q = x Wq
k = x Wk
v = x Wv
The original token position is the same. The projections give the model different feature spaces for asking, matching, and carrying content.
For a model width C=8, one head might use a head dimension D=4:
x: (8)
Wq: (8,4) → q: (4)
Wk: (8,4) → k: (4)
Wv: (8,4) → v: (4)
This is why Q/K/V should not be imagined as three separate token sequences. They are three learned transformations of the current hidden representations.
A useful debugging boundary is to verify the Q, K, and V shapes before computing any attention scores. If those projections are wrong, a later softmax can still return perfectly normalized probabilities for the wrong computation.
Predict
Change one role at a time
The Lab has one query = [1.0, 0.0], two keys, and two values. Keys decide how much attention each position gets; values decide what content is mixed into the output.
- Click Run. Read
weights: [0.731, 0.269]andoutput: [7.311, 5.379]. The query matches the first key better, so the first value gets most of the weight. - Change only a key. Replace
keys = [[1.0, 0.0], [0.0, 1.0]]withkeys = [[0.0, 1.0], [0.0, 1.0]], so the first key no longer matches the query. - Before running, predict the new weights.
- Click Run. The weights become
[0.5, 0.5]and the output becomes[5.0, 10.0]. Changing a key changed the scores, and therefore the mix. - Press Reset. Now change only a value: replace
values = [[10.0, 0.0], [0.0, 20.0]]withvalues = [[30.0, 0.0], [0.0, 20.0]]. - Click Run. The weights stay exactly
[0.731, 0.269], but the output becomes[21.932, 5.379]. Changing a value changed the content that gets mixed, not the attention scores. - Press Reset afterward.
Loading lab…
A silent semantic bug is swapping key and value roles when dimensions happen to match. The code may run, but “what gets scored” and “what gets retrieved” no longer implement the intended algorithm.
Quick Check
Explain it back
Use one sentence to define each role without saying that Q, K, and V are three different tokens. They are learned views/roles applied to the same sequence representations during self-attention.
Key Takeaways
- Queries ask; keys are compared; values carry retrieved information.
- Scoring and value mixing are separate stages.
- Changing one role at a time makes semantic wiring bugs visible.
- Q/K/V are roles in the attention operation, not token categories.
Next Lesson
Next, turn query-key alignment into one numerical score using the dot product.
References
- Vaswani et al., Attention Is All You Need.
- Bahdanau, Cho, and Bengio, Neural Machine Translation by Jointly Learning to Align and Translate.
Completion is stored locally on this device.