본문으로 건너뛰기
L4.5

Byte-Level Tokenization

Goal

Encode Unicode text as UTF-8 bytes, explain the coverage benefit of a 256-value byte vocabulary, and identify the sequence-length and boundary costs of multi-byte characters.

Bytes give a strong coverage guarantee​

The earlier tokenizer lesson showed a simple fixed vocabulary. Now consider text the vocabulary never anticipated: é, 🙂, or another writing system.

UTF-8 converts text into byte values from 0 to 255. That gives a powerful guarantee: any valid UTF-8 string can be represented without inventing an unknown character token.

The guarantee has a cost. One visible character is not always one byte. ASCII A uses one byte, while many accented characters and emoji use multiple bytes, so visible text can expand into several sequence positions.

Follow one character through the layers​

Suppose an unfamiliar character becomes four UTF-8 bytes.

The tokenizer may expose those bytes directly or use learned byte-derived subwords. The model then receives token IDs corresponding to those pieces. Later embedding layers turn the IDs into vectors.

The model is not “reading raw Unicode” inside attention. The tokenizer has already converted the text to a finite sequence of discrete IDs.

Coverage is not the same as efficiency​

A tokenizer that can represent every input is not automatically a good tokenizer.

If ordinary text regularly expands into very long byte sequences, the model spends more context-window space and more computation on the same human-visible sentence.

That is why practical systems often combine broad byte coverage with learned merges or subwords. Common patterns become compact tokens, while byte-level behavior remains a fallback for unusual text.

When evaluating tokenizers, check both can this text be represented? and how many tokens does it require?

Predict

Why can a byte tokenizer represent an unfamiliar emoji without `<unk>`?

Inspect actual bytes​

Before the lab, compare len(text) with len(text.encode('utf-8')) for A, é, and an emoji. The difference makes the coverage/length tradeoff concrete.

  1. Click Run. Compare characters= with bytes=: A is 1 byte, é is 2, and the emoji 🙂 is 4. The last line, round trip: Token🙂, shows that bytes decode back to the original text exactly.
  2. Add a character from another writing system. Change SAMPLES = ["A", "é", "🙂", "AI🙂"] to SAMPLES = ["A", "é", "🙂", "AI🙂", "가"].
  3. Before running, predict how many bytes the Korean syllable 가 needs.
  4. Click Run. 가 is 1 character but 3 bytes: [234, 176, 128]. Byte-level tokenization can represent any text, but some scripts cost more positions.

Loading lab…

  1. Now create a deliberate failure by cutting through the middle of a multi-byte character. Change round_trip = decode_bytes(byte_tokens("Token🙂")) to round_trip = decode_bytes(byte_tokens("Token🙂")[:-2]), which drops the last two of the emoji's four bytes.
  2. Click Run. The Lab stops with UnicodeDecodeError ... unexpected end of data. The remaining bytes are not valid UTF-8. This matters later when truncation or windowing cuts a sequence near such a boundary.
  3. Press Reset afterward.

Quick Check

1. How many distinct byte values exist?
2. Why can an emoji occupy several byte-token positions?
3. What can break UTF-8 decoding?

0 of 3 questions answered.

Transfer the idea​

For user-generated text with names, symbols, and many languages, explain why broad coverage may be worth extra sequence positions—and why you would still measure token counts rather than assuming the cost is small.

Key Takeaways

  • UTF-8 gives a fixed byte alphabet that can represent any valid text.
  • Byte coverage avoids unknown characters but can increase sequence length.
  • Visible-character boundaries and byte boundaries are different.
  • Round-trip and boundary tests are essential evidence for byte representations.

Next Lesson

Next, use corpus statistics to combine frequent smaller pieces into reusable subwords that can reduce sequence length while keeping fallback coverage.

References

Lesson actions

Completion is stored locally on this device.

View progress