본문으로 건너뛰기
L4.6

Subword Learning

Goal

Simulate a frequency-based subword merge and explain why adding larger reusable pieces can shorten common sequences while increasing vocabulary size.

Start with small pieces for low, lower, and lowest. If the same neighboring pieces appear repeatedly, we can replace that pair with one new reusable piece.

l o w
l o w e r
l o w e s t

Repeated merges might eventually create low. The vocabulary gained an entry, but every occurrence of that pattern can now use fewer sequence positions. That is the key mechanism—not a claim that the learned piece is a linguistically perfect word part.

Learn reusable pieces from repeated text​

Imagine a tiny corpus containing:

play
playing
played
player
replay

If the tokenizer begins with very small units, it can notice that sequences corresponding to play appear repeatedly. A merge-based learner can gradually combine frequent neighboring pieces.

A simplified history might look like:

p + l → pl
pl + a → pla
pla + y → play

Later, text containing playing might be represented with pieces such as play and ing.

The important idea is not the exact merge algorithm. It is that the vocabulary is learned from repeated patterns in text rather than manually listing every possible word.

Frequency alone does not create meaning​

If two characters frequently occur together, a tokenizer may merge them even if the resulting piece is not a meaningful word or morpheme to a human.

Subword learning optimizes a representation objective such as compression or frequency-based reuse. It is not performing full linguistic analysis.

This is why token pieces can look odd while still being useful.

Training data shapes the tokenizer​

A tokenizer trained mostly on English prose may split source code, Finnish, Korean, mathematical notation, or biomedical text less efficiently.

That can increase token counts and change which pieces are available to the model.

Tokenizer training therefore has a data-distribution question just like model training: what kinds of text were represented when the vocabulary was learned?

A good evaluation should include the real text domains the final model will receive, not only examples similar to tokenizer-training data.

Predict

What happens immediately when a frequent adjacent pair is merged into one reusable piece?

Follow one merge​

The Lab starts from the word list WORDS = ["low", "lower", "lowest", "low"]. It splits each word into characters and adds </w>, a marker meaning “end of word,” so that the w at the end of low stays different from the w inside lower.

  1. On paper, count adjacent pairs in the four words. Which pairs appear most often?
  2. Predict how many positions the whole list will have after merging that pair everywhere.
  3. Click Run. pair counts: shows ('l', 'o'): 4 and ('o', 'w'): 4 tied at the top. The Lab breaks ties alphabetically, so chosen pair: ('l', 'o').
  4. Compare positions before: 21 with positions after: 17. One merge saved four positions, one for each word.

Loading lab…

  1. Now change the training corpus. Replace the list with WORDS = ["newer", "newest", "wider", "widest"]. Before running, predict whether l + o could still be chosen.
  2. Click Run. Every pair now appears exactly twice, and the tie-break chooses ('d', 'e'). The lo merge learned from the first corpus would save nothing here. A merge learned from one narrow domain can be nearly useless elsewhere, which is why tokenizer training data shapes downstream efficiency.
  3. Press Reset afterward.

Quick Check

1. What drives the toy merge decision?
2. Why can vocabulary size rise while sequence length falls?
3. Why evaluate on held-out text?

0 of 3 questions answered.

Explain it back​

Explain how subword learning resembles compression without claiming that a smaller token count automatically makes a better language model.

Key Takeaways

  • Subword learning promotes frequent smaller patterns into reusable larger pieces.
  • Merges trade vocabulary capacity for shorter encodings on matching text.
  • Corpus choice strongly affects which pieces become efficient.
  • Evaluate the learned representation on held-out and diverse text slices.

Next Lesson

Next, reserve stable IDs for control tokens so ordinary vocabulary changes cannot silently change padding, unknown, beginning, or end semantics.

References

Lesson actions

Completion is stored locally on this device.

View progress