본문으로 건너뛰기
L4.10

Position Matters

Goal

Demonstrate why token identity alone cannot distinguish reordered sequences and explain the separate jobs of token and position representations.

Compare:

dog bites person
person bites dog

The token inventory is nearly identical, but the order changes the event. If a representation simply sums token embeddings, a permutation can collapse to the same result because addition ignores order.

Token information answers what is here? Position information answers where is it? A sequence model needs both.

Identity alone cannot recover order​

Suppose the model receives these token embeddings:

dog → D
bites → B
person → P

If a representation only sums them, both sentences below produce the same total:

D + B + P
P + B + D

Addition is commutative, so the order disappears.

But the sentences “dog bites person” and “person bites dog” describe different events.

A sequence model therefore needs some way to distinguish not only which token is present but where it occurs.

Token identity and position answer different questions​

Think of each input position as combining two signals:

token information: what is in this slot?
position information: which slot is this?

A simplified input vector might be:

input_vector = token_embedding + position_vector

Now the same token can enter the model differently at position 0 and position 5 even though its token ID is unchanged.

Position bugs can survive every dimension check​

If you insert a beginning token but forget to shift later positions, all tensors may still have legal shapes.

The bug is semantic: each token is paired with the wrong position.

This is why a useful test compares known token-position pairs, not only tensor dimensions. Insert one token at the front and verify that every later position moves by exactly one.

Predict

What can happen if two sequences contain the same token embeddings but the representation only sums them?

Make position explicit in code​

Write [4, 7, 9] as token-position pairs (4,0) (7,1) (9,2). Reverse the token order while keeping positions 0–2. Now the identity attached to each slot changes visibly.

The Lab leaves attach_positions unfinished on purpose.

  1. Run the starter once. A pairs and B pairs are empty, so the exercise checks fail.
  2. Complete the TODO so the function returns one (token_id, position) pair for every token, in order. The position is start + offset.
  3. Run again and confirm [4,7,9] becomes [(4,0),(7,1),(9,2)].
  4. Verify that reversing the IDs changes the ordered pairs even though bag_signature stays the same.
  5. Finally test the provided start=1 case and explain why all positions shift by one.

Loading lab…

A classic semantic bug is an off-by-one shift after <bos> is inserted. Shapes still match; every token simply receives the position intended for its neighbor.

Quick Check

1. Why is token identity alone insufficient for order-sensitive text?
2. What does adding `<bos>` require from position indexing?
3. Why can a position bug survive shape checks?

0 of 3 questions answered.

Transfer the idea​

Construct two short sequences a bag-of-tokens representation cannot distinguish. Then explain how explicit positions separate them.

Key Takeaways

  • Order can change meaning even when token identities are unchanged.
  • Token and position representations answer different questions.
  • Position alignment must survive special-token insertion, padding, and slicing.
  • Semantic position errors can exist with perfectly valid tensor shapes.

Next Lesson

Next, turn sequence position into numerical vectors that can be combined with token embeddings.

References

Lesson actions

Completion is stored locally on this device.

View progress