본문으로 건너뛰기
L4.4

Build a Simple Tokenizer

Goal

Turn text into stable integer token IDs, then turn those IDs back into readable tokens, while making the limitations of a tiny word-level vocabulary visible.

A language model cannot multiply words directly. It needs numbers. Tokenization is the reproducible bridge from text to the integer IDs that later select embedding rows.

Entry skills for this coding lesson​

You should be comfortable reading basic Python variables, lists, dictionaries, functions, and for loops. You do not need to know regular expressions, tensors, matrix multiplication, softmax, PyTorch modules, gradients, or cross-entropy yet.

If those basic Python pieces are unfamiliar, use the Project Workbench bridge first. Do not treat a Python-environment problem as if it were a tokenizer concept problem.

One sentence, two representations​

Start with:

The model predicts the next token.

A simple word-level tokenizer can split it into:

["the", "model", "predicts", "the", "next", "token", "."]

A vocabulary then gives each distinct piece a stable integer ID. The exact numbers are arbitrary; the mapping is not. If "model" is ID 7 today and ID 12 tomorrow while the model weights stay unchanged, the same integer no longer means the same input.

Saved tokenizer data and model weights therefore have to agree on what every ID means.

Predict

What should this tiny tokenizer do with a word that is not in its vocabulary?

Make the splitting rule explicit​

This deterministic splitter is deliberately small:

import re

pattern = re.compile(r"[A-Za-z]+|[0-9]+|[^\w\s]")

def split(text):
return pattern.findall(text.lower())

print(split("Tiny models learn, too!"))
# ['tiny', 'models', 'learn', ',', 'too', '!']

A regular expression (regex) is a compact text pattern. Here it means: take runs of letters, runs of digits, or one punctuation symbol. You do not need to design regex syntax for this Level. You only need to know exactly what rule produced the pieces.

Build a stable vocabulary​

A sorted vocabulary makes the same corpus produce the same IDs every run:

tokens = split(corpus)
vocab = ["<unk>", *sorted(set(tokens))]
stoi = {token: i for i, token in enumerate(vocab)}
itos = vocab

def encode(text):
return [stoi.get(token, 0) for token in split(text)]

def decode_tokens(ids):
return [itos[i] for i in ids]

<unk> uses ID 0. Any piece missing from the vocabulary falls back to that explicit unknown token rather than receiving a random new ID.

Run the tokenizer notebook​

The Lab opens in a notebook and runs on CPU; you do not need a GPU.

  1. Open the Lab and run the cells unchanged.
  2. In the sample cell, find:
sample = "The model predicts the next token."
  1. Read the printed token list, ID list, and decoded text. Confirm that the number of tokens equals the number of IDs.
  2. In the later failure cell, run tokenizer.encode("the zzzxxyy token") and inspect both unknown_ids and decode_tokens(unknown_ids).
  3. Find the 0 in unknown_ids and the matching <unk> piece. That is visible evidence of information the tiny vocabulary cannot represent exactly.
  4. Do not “fix” the example by silently adding zzzxxyy to the vocabulary. The limitation is the point of this experiment.

Loading lab…

After the guided pass, change only the sample string to another sentence made from familiar corpus words and punctuation. Predict its pieces before rerunning that cell.

What can go wrong even when the code runs?​

If two machines build different IDs from the same text, check these boundaries first:

  1. normalization — did both lowercase the same way?
  2. splitting — did both produce the same pieces?
  3. vocabulary ordering — did both assign IDs in the same deterministic order?
  4. saved tokenizer identity — are both actually using the same saved tokenizer and vocabulary?

A set can collect unique pieces, but sorting those pieces makes the teaching mapping explicit and reproducible.

Why real LLM tokenizers go beyond whole words​

A word-level tokenizer is easy to inspect, but unseen words collapse to <unk>. Real LLMs commonly use subword or byte-oriented strategies so they can represent unfamiliar names, spellings, languages, and symbols without requiring one vocabulary entry for every possible word.

You will explore those tradeoffs next rather than hiding them inside this first tokenizer.

Quick Check

1. Why sort the unique tokens before assigning IDs?
2. What does <unk> represent?
3. Why is an ID list not enough by itself?

0 of 3 questions answered.

Key Takeaways

  • Tokenization converts text into discrete pieces and stable integer IDs.
  • ID values are arbitrary addresses; the tokenizer mapping gives them meaning.
  • Deterministic splitting and vocabulary ordering make runs reproducible.
  • <unk> makes an unseen-word limitation visible instead of hiding it.
  • Whole-word tokenization is intentionally simple; the next lessons explore byte and subword alternatives.

Next Lesson

Next, you will study L4.5 — Byte-level Tokenization and see how UTF-8 bytes can represent any input text without a whole-word <unk> fallback, at the cost of potentially longer sequences.

References

Lesson actions

Completion is stored locally on this device.

View progress