Vocabulary and Unknown Tokens
Goal
Build a deterministic vocabulary, encode known and unknown pieces, and explain what information is lost when different unseen pieces share one <unk> ID.
Suppose a tiny word vocabulary knows cat, dog, and runs, but not quokka or axolotl. A fallback keeps the program running, but both unseen pieces map to the same reserved entry.
Unknown tokens: different inputs can collapse to one ID
A fallback keeps encoding defined, but it can erase which unseen piece was originally present.
quokka<unk>0axolotl<unk>0Information lost: the original unseen pieces are different, but their encoded representation is now identical.
The failure is not a crash. It is information collapse: two different source pieces have become indistinguishable at that position.
A vocabulary therefore has two responsibilities. It must cover useful pieces, and it must map them to IDs deterministically so stored sequences keep the same meaning across runs.
A vocabulary is a finite lookup table
Imagine a vocabulary containing only:
0: <unk>
1: cat
2: dog
3: runs
4: sleeps
The sentence cat runs can become [1, 3].
But if fox is not representable by smaller known pieces, an old-style word tokenizer may fall back to <unk>. Then fox runs, robot runs, and banana runs can begin with the same ID. That is real information loss.
Coverage and vocabulary size pull in opposite directions
Putting every possible word in a vocabulary does not scale: language keeps producing names, spelling variants, numbers, URLs, code, and many writing systems. A huge vocabulary also enlarges embedding and output matrices.
Modern tokenizers instead use reusable pieces, bytes, or fallback strategies so rare text can still be represented.
Unknown does not mean semantically unknown
The token <unk> means the tokenizer could not represent a surface form under its current policy. It does not mean the model has decided the concept is unknown.
Conversely, a tokenizer may encode a rare word perfectly while the model knows very little about its meaning. Tokenizer coverage and model knowledge are different failure layers.
Predict
Measure coverage instead of assuming it
Build a tiny vocabulary with <unk> reserved first, then encode one familiar sentence and one from a different topic. The simple diagnostic unknown pieces / total pieces makes coverage failure visible.
- Click Run. Read
vocab: ['<unk>', 'learn', 'models', 'predict', 'tiny', 'tokens']. - Read
evaluation ids: [4, 0, 0]andunknown rate: 0.667. Onlytinyis known;quokkaandpredictsboth collapse to ID0. - Notice
predicts. The vocabulary knowspredict, but a word-level tokenizer treatspredictsas a completely different word. - Add one training sentence. In
TRAIN_TEXT, add"a quokka naps",as a third line. - Before running, predict the new unknown rate on the same evaluation sentence.
- Click Run.
quokkanow has its own ID and the rate drops to0.333, butpredictsis still unknown. Coverage improved only for the exact word you added. - Click Run a second time without editing. The vocabulary and IDs are identical, because
build_vocabsorts the pieces. That determinism is what keeps stored IDs meaningful across runs. - Press Reset afterward.
Loading lab…
A low unknown rate is not proof that a tokenizer is good; it is one piece of evidence. Byte and subword approaches will show ways to preserve more information while keeping the representation finite.
Quick Check
Explain it back
Explain how a tokenizer can have zero runtime errors while still being unusable because its unknown rate is high on the target users' text.
Key Takeaways
- The vocabulary is the piece-to-ID contract.
<unk>is a safe fallback that deliberately loses identity.- Coverage must be measured on representative evaluation text.
- Deterministic vocabulary construction is necessary for reproducibility.
Next Lesson
Next, L4.4 — Build a Simple Tokenizer builds the simple tokenizer end to end. Use it to connect the boundary and vocabulary ideas to runnable code.
References
- Kudo and Richardson, SentencePiece.
Completion is stored locally on this device.