본문으로 건너뛰기
L5.3

Queries, Keys, and Values

Goal

Distinguish query, key, and value roles and trace how changing a key affects scores while changing a value affects retrieved content.

Attention separates two questions that are easy to mix up:

  1. How relevant is this position? — compare a query with keys.
  2. What information should be retrieved? — mix value vectors using the resulting weights.

A useful analogy is a search system: query = request, key = index description, value = content returned after matching. The analogy is imperfect, but it keeps the roles separate.

Trace one lookup with tiny vectors​

Use one query and two candidate positions:

query q = [1, 0]

key A = [1, 0] value A = [10, 0]
key B = [0, 1] value B = [0, 10]

Before softmax, the dot-product scores are 1 for A and 0 for B. So A should receive more normalized weight, and the output should lean toward [10, 0].

Now imagine changing value A to [0, 20] while leaving its key unchanged. The score for A should stay the same, because scoring uses the query and key. But the retrieved content changes because the value changed.

That separation gives you a powerful debugging test:

  • key-only edit → score can change;
  • value-only edit → score should not change;
  • either edit can change the final mixed output for different reasons.

Do not memorize Q/K/V as three mysterious letters. Ask two ordinary questions: what decides relevance? Q and K. What information is carried after relevance is decided? V.

Keep key and value roles separate

Select a candidate position. Its key is compared with the query; its value is the content that can be mixed after weighting.

Selected: Position A
Query[1, 0]asks what is useful now
Position A key[1, 0]is compared with the query
Position A value[10, 0]carries content if this position receives weight

Q, K, and V are learned views of the same hidden state​

In self-attention, each position begins with one hidden vector x, but learned projection matrices create different views:

q = x Wq
k = x Wk
v = x Wv

The original token position is the same. The projections give the model different feature spaces for asking, matching, and carrying content.

For a model width C=8, one head might use a head dimension D=4:

x: (8)
Wq: (8,4) → q: (4)
Wk: (8,4) → k: (4)
Wv: (8,4) → v: (4)

This is why Q/K/V should not be imagined as three separate token sequences. They are three learned transformations of the current hidden representations.

A useful debugging boundary is to verify the Q, K, and V shapes before computing any attention scores. If those projections are wrong, a later softmax can still return perfectly normalized probabilities for the wrong computation.

Predict

Two keys score equally against one query, but their values differ. What should happen after normalization?

Change one role at a time​

The Lab has one query = [1.0, 0.0], two keys, and two values. Keys decide how much attention each position gets; values decide what content is mixed into the output.

  1. Click Run. Read weights: [0.731, 0.269] and output: [7.311, 5.379]. The query matches the first key better, so the first value gets most of the weight.
  2. Change only a key. Replace keys = [[1.0, 0.0], [0.0, 1.0]] with keys = [[0.0, 1.0], [0.0, 1.0]], so the first key no longer matches the query.
  3. Before running, predict the new weights.
  4. Click Run. The weights become [0.5, 0.5] and the output becomes [5.0, 10.0]. Changing a key changed the scores, and therefore the mix.
  5. Press Reset. Now change only a value: replace values = [[10.0, 0.0], [0.0, 20.0]] with values = [[30.0, 0.0], [0.0, 20.0]].
  6. Click Run. The weights stay exactly [0.731, 0.269], but the output becomes [21.932, 5.379]. Changing a value changed the content that gets mixed, not the attention scores.
  7. Press Reset afterward.

Loading lab…

A silent semantic bug is swapping key and value roles when dimensions happen to match. The code may run, but “what gets scored” and “what gets retrieved” no longer implement the intended algorithm.

Quick Check

1. Which roles form compatibility scores?
2. Which vectors are mixed by the normalized weights?
3. A score changes after editing a value but the key is unchanged. What should you suspect?

0 of 3 questions answered.

Explain it back​

Use one sentence to define each role without saying that Q, K, and V are three different tokens. They are learned views/roles applied to the same sequence representations during self-attention.

Key Takeaways

  • Queries ask; keys are compared; values carry retrieved information.
  • Scoring and value mixing are separate stages.
  • Changing one role at a time makes semantic wiring bugs visible.
  • Q/K/V are roles in the attention operation, not token categories.

Next Lesson

Next, turn query-key alignment into one numerical score using the dot product.

References

Lesson actions

Completion is stored locally on this device.

View progress