Characters, Words, and Subwords
Goal
By the end of this lesson, you can compare three ways to split text into token pieces and explain why smaller pieces usually need more sequence positions while larger pieces need a larger vocabulary to cover many words.
One word can be split several ways
Take the word playing. A tokenizer could represent the same text using character-sized pieces, one whole-word piece, or reusable subword pieces.
A position here means one token slot in the sequence sent toward the model. Compare how many positions each choice needs:
Tokenization: one text, different piece sizes
Compare character, whole-word, and subword choices for the same text. Smaller pieces are easier to reuse, but they usually take more sequence positions.
playing| Strategy | Pieces | Positions |
|---|---|---|
| Characters | p | l | a | y | i | n | g | 7 |
| Whole word | playing | 1 |
| Subwords | play | ing | 2 |
A whole-word tokenizer uses a complete word as one token when that word exists in its vocabulary. This is compact for familiar words, but a vocabulary cannot realistically contain every name, spelling variation, or newly invented word.
A character tokenizer uses much smaller pieces. It can build unfamiliar spellings from a small reusable set of characters, but common words take many more positions.
A subword tokenizer uses reusable pieces that are often larger than one character but smaller than a complete word. For example, playing can reuse play and ing. The split does not always match the word parts a grammar teacher would choose; the goal is useful reuse for the tokenizer.
So there is a tradeoff:
larger pieces
→ fewer positions for familiar text
→ more distinct pieces may be needed in the vocabulary
smaller pieces
→ more reusable pieces
→ more positions are often needed for the same text
Here coverage means whether the tokenizer has a way to represent the text it receives. A strategy with good coverage can still encode unfamiliar text instead of failing only because the complete word was unseen.
Predict
Compare the same text in the Lab
The Lab prints character, word, and simple subword versions of the same text.
- Click Run without editing anything.
- For
players replayed, compare the three lines. Pay attention to bothcount=andpieces=. - Notice that the simple subword rule splits the suffixes, producing pieces such as
play+ersandreplay+ed. - Find:
text = "players replayed"
- Change only the text to:
text = "play replay"
- Before running, predict the direction of the change. The word representation still needs 2 positions. The simple subword representation should also need only 2 positions now because neither word has one of the suffixes that the starter splits off.
- Click Run and compare all three
count=values with the first run. - Restore
text = "players replayed".
Loading lab…
This small program is only one toy subword rule. Real subword tokenizers learn or choose reusable pieces from much larger text collections. The point of the Lab is the tradeoff between piece size, reuse, and sequence length—not to claim that this suffix rule is a production tokenizer.
Quick Check
Transfer the idea
Suppose a chat product works well on common English words. Why would that alone not prove that its tokenizer works efficiently for names, emoji, compounds, misspellings, or other writing systems?
The answer should mention the kind of text tested. A tokenizer's behavior depends on the inputs it must actually represent.
Key Takeaways
- The same text can be split into character, whole-word, or subword tokens.
- Smaller pieces are broadly reusable but usually create longer sequences.
- Larger pieces can shorten familiar text but require broader vocabulary coverage.
- Subwords aim for a useful middle ground between reuse and sequence length.
- Judge tokenization choices on the kinds of text the system actually needs to handle.
Next Lesson
Next, in L4.3 — Vocabulary and Unknown Tokens, you will make the piece-to-ID mapping explicit and see what happens when a vocabulary cannot represent a piece directly.
References
- Kudo and Richardson, SentencePiece.
- Sennrich, Haddow, and Birch, Neural Machine Translation of Rare Words with Subword Units.
Completion is stored locally on this device.