Subword Learning
Goal
Simulate a frequency-based subword merge and explain why adding larger reusable pieces can shorten common sequences while increasing vocabulary size.
Start with small pieces for low, lower, and lowest. If the same neighboring pieces appear repeatedly, we can replace that pair with one new reusable piece.
l o w
l o w e r
l o w e s t
Repeated merges might eventually create low. The vocabulary gained an entry, but every occurrence of that pattern can now use fewer sequence positions. That is the key mechanism—not a claim that the learned piece is a linguistically perfect word part.
Learn reusable pieces from repeated text
Imagine a tiny corpus containing:
play
playing
played
player
replay
If the tokenizer begins with very small units, it can notice that sequences corresponding to play appear repeatedly. A merge-based learner can gradually combine frequent neighboring pieces.
A simplified history might look like:
p + l → pl
pl + a → pla
pla + y → play
Later, text containing playing might be represented with pieces such as play and ing.
The important idea is not the exact merge algorithm. It is that the vocabulary is learned from repeated patterns in text rather than manually listing every possible word.
Frequency alone does not create meaning
If two characters frequently occur together, a tokenizer may merge them even if the resulting piece is not a meaningful word or morpheme to a human.
Subword learning optimizes a representation objective such as compression or frequency-based reuse. It is not performing full linguistic analysis.
This is why token pieces can look odd while still being useful.
Training data shapes the tokenizer
A tokenizer trained mostly on English prose may split source code, Finnish, Korean, mathematical notation, or biomedical text less efficiently.
That can increase token counts and change which pieces are available to the model.
Tokenizer training therefore has a data-distribution question just like model training: what kinds of text were represented when the vocabulary was learned?
A good evaluation should include the real text domains the final model will receive, not only examples similar to tokenizer-training data.
Predict
Follow one merge
The Lab starts from the word list WORDS = ["low", "lower", "lowest", "low"]. It splits each word into characters and adds </w>, a marker meaning “end of word,” so that the w at the end of low stays different from the w inside lower.
- On paper, count adjacent pairs in the four words. Which pairs appear most often?
- Predict how many positions the whole list will have after merging that pair everywhere.
- Click Run.
pair counts:shows('l', 'o'): 4and('o', 'w'): 4tied at the top. The Lab breaks ties alphabetically, sochosen pair: ('l', 'o'). - Compare
positions before: 21withpositions after: 17. One merge saved four positions, one for each word.
Loading lab…
- Now change the training corpus. Replace the list with
WORDS = ["newer", "newest", "wider", "widest"]. Before running, predict whetherl+ocould still be chosen. - Click Run. Every pair now appears exactly twice, and the tie-break chooses
('d', 'e'). Thelomerge learned from the first corpus would save nothing here. A merge learned from one narrow domain can be nearly useless elsewhere, which is why tokenizer training data shapes downstream efficiency. - Press Reset afterward.
Quick Check
Explain it back
Explain how subword learning resembles compression without claiming that a smaller token count automatically makes a better language model.
Key Takeaways
- Subword learning promotes frequent smaller patterns into reusable larger pieces.
- Merges trade vocabulary capacity for shorter encodings on matching text.
- Corpus choice strongly affects which pieces become efficient.
- Evaluate the learned representation on held-out and diverse text slices.
Next Lesson
Next, reserve stable IDs for control tokens so ordinary vocabulary changes cannot silently change padding, unknown, beginning, or end semantics.
References
- Kudo and Richardson, SentencePiece.
- Sennrich, Haddow, and Birch, Neural Machine Translation of Rare Words with Subword Units.
Completion is stored locally on this device.