Skip to main content
L5.7

Self-Attention

Goal

Explain what makes attention self-attention: the queries, keys, and values are all produced from the same sequence representations. Preview the decoder's no-future-information requirement without needing to implement causal masking or multi-head reshaping yet.

One sequence plays all three roles​

In the previous lessons, you separated the roles:

  • a query asks what information is useful now;
  • a key is compared with that query;
  • a value carries the information that gets mixed.

In self-attention, all three roles start from the same sequence of token representations. Each token representation is projected into a query, key, and value view.

That is different from saying that query, key, and value are three different kinds of token. They are three learned views of the same sequence state.

For a sequence such as:

a small model learns

each position can form a query and compare it with keys from legal positions in that same sequence. The resulting weights mix value vectors from those positions.

Predict

What makes this operation self-attention?

A decoder adds one more rule​

A decoder-only language model predicts left to right. That means an earlier position must not use information from a token that lies in its future.

The viewer below shows that rule visually. Select different query positions. The × cells are future positions that a causal decoder must not use.

Self-attention with a causal boundary

Select a query token. For now, focus on which positions are legal; the next lesson builds and tests the mask numerically.

Query: model · strongest visible link: model
query ↓ / key →
a
small
model
learns
a
100%
small
40%
60%
model
20%
30%
50%
learns
10%
20%
30%
40%

At query position 2, positions 0, 1, and 2 are legal. Position 3 is future information.

Do not worry yet about the exact masking code or the tensor reshaping used for several heads. Those are the next two dedicated lessons.

Inspect the notebook at the right depth​

The notebook contains a fuller causal self-attention implementation because later lessons reuse the same worked example.

For this lesson, use it only to verify two ideas:

  1. Q, K, and V are produced from the same input sequence tensor.
  2. the reported total probability assigned to future positions is zero.

You do not need to explain the multi-head reshape yet.

Loading lab…

The notebook runs on CPU; you do not need a GPU.

If it shows a head dimension or causal-mask implementation you have not studied yet, treat those as preview details. L5.8 — Causal Masking makes the no-future rule an explicit test; L5.9 — Multi-Head Attention then explains how feature channels are split across heads.

A common misconception​

“Self-attention means every position may read every other position.”

Not necessarily. Self describes where Q, K, and V come from. A separate masking rule decides which positions are legal to use.

A decoder uses self-attention and causal masking.

Quick Check

1. In self-attention, where do query, key, and value views come from?
2. Does the word 'self' by itself guarantee that future positions are blocked?

0 of 2 questions answered.

Key Takeaways

  • Self-attention produces query, key, and value views from the same sequence representations.
  • Q, K, and V are roles/projections, not different token categories.
  • Decoder self-attention also needs a causal boundary so earlier positions cannot use future information.
  • You do not need multi-head mechanics yet; those are introduced after causal masking.

Next Lesson

Next, L5.8 — Causal Masking turns the decoder's no-future-information rule into an executable check and proves numerically that future attention mass is zero.

References

Lesson actions

Completion is stored locally on this device.

View progress