Token IDs and Special Tokens
Goal
Reserve stable IDs for <pad>, <unk>, <bos>, and <eos>, and explain why changing those mappings can corrupt previously encoded data even when every integer remains valid.
Ordinary pieces represent text. Special tokens represent control information.
<pad> filler position
<unk> piece not represented by the vocabulary
<bos> sequence begins
<eos> sequence ends
Suppose an encoded example is [2, 8, 11, 3]. It is only meaningful if every component that reads it agrees that ID 2 means <bos> and ID 3 means <eos>. If another tokenizer version swaps those meanings, the array is still syntactically valid but semantically wrong.
IDs are conventions that must stay synchronized
Suppose a tokenizer uses:
<pad> = 0
<bos> = 1
<eos> = 2
cat = 17
sat = 42
Then the sequence:
<bos> cat sat <eos>
becomes:
[1, 17, 42, 2]
Those numbers only have meaning relative to this exact vocabulary.
If another tokenizer says 17 = dog, loading the same model weights with that tokenizer silently changes the input meaning. Shapes still match, but the model receives the wrong symbols.
Special tokens are part of the model contract
Special tokens can mark boundaries or carry control information:
<bos>— beginning of sequence;<eos>— end of sequence;<pad>— fill unused positions in a batch;- task- or chat-specific control tokens — distinguish roles or sections.
Their IDs must be consistent across training, evaluation, saving, and generation.
Do not treat special tokens as decorative strings. They become ordinary integer IDs at model input, so a mismatch is a real data bug.
Adding one token can shift later assumptions
Suppose a sequence originally starts at position 0 with cat. After inserting <bos>, cat moves to position 1.
The token ID for cat may stay 17, but its position changes. That is why token identity and position identity must be handled separately in later lessons.
Predict
Verify the contract
The Lab reserves the four special tokens first, then adds the sorted ordinary words from TEXTS.
- Click Run. Read
special ids: {'<pad>': 0, '<unk>': 1, '<bos>': 2, '<eos>': 3}andencoded: [2, 4, 1, 3]for"cats quokka": beginning marker,cats, unknownquokka, end marker. - Add a new training sentence. Change
TEXTS = ["cats nap", "dogs play"]toTEXTS = ["cats nap", "dogs play", "birds sing"]. - Before running, predict: which IDs will stay the same, and will the ID for
catschange? - Click Run. The special IDs are still
0–3, but the encoding becomes[2, 5, 1, 3]:catsmoved from ID4to5, becausebirdssorts before it. Reserved IDs protect the control tokens; ordinary IDs can still shift when the vocabulary is rebuilt, which is why a saved vocabulary must travel with saved data. - Press Reset.
Loading lab…
- Now create the dangerous version. Change
SPECIAL_TOKENS = ["<pad>", "<unk>", "<bos>", "<eos>"]toSPECIAL_TOKENS = ["<eos>", "<unk>", "<bos>", "<pad>"]and run again. - The new encoding is
[2, 4, 1, 0]. Now read the old encoding[2, 4, 1, 3]with the new table: its final3now means<pad>, so the sentence appears to have no end marker. Nothing crashed and every shape still matches. The bug is only in the meaning. - Press Reset afterward.
Quick Check
Explain it back
Explain why a special-token mismatch can look like a model-quality failure even when the model parameters themselves are unchanged.
Key Takeaways
- Special tokens carry control semantics, not ordinary text meaning.
- Their IDs are a versioned compatibility contract.
- Semantic mapping bugs can survive all shape/type checks.
- Verify configured IDs against the artifact that produced encoded data.
Next Lesson
Complete the checkpoint now. After it, you will use <pad> together with masks and truncation to create rectangular model inputs without confusing filler with real text.
References
- Kudo and Richardson, SentencePiece.
Completion is stored locally on this device.