Skip to main content
L4.1

Text Is Not Numbers Yet

Goal

By the end of this lesson, you can trace one short sentence from text to pieces to integer IDs and explain why the ID numbers themselves do not contain the meaning of the words.

Start with one short sentence​

Consider:

Cats nap.

A person sees letters and words. A neural network needs a reproducible numerical input.

One possible conversion is:

Cats nap.
→ ["cats", "nap", "."]
→ [2, 3, 1]

The first list contains tokens: pieces of text that the tokenizer decided to handle as units.

A tokenizer is the rule or program that turns text into those pieces and assigns each piece an integer token ID.

The saved mapping between pieces and IDs is called a vocabulary. For example:

"." → 1
"cats" → 2
"nap" → 3

The exact numbers are not universal. Another tokenizer could use "cats" → 17 and still be correct if the same mapping is used consistently.

That is why 17 does not mean “more cat-like” than 5. A token ID is an address in a vocabulary, not a measurement of meaning.

Predict

Two tokenizers map `cat` to different integer IDs. Can both mappings be valid?

Why the mapping must travel with the numbers​

Imagine receiving only this:

[17, 42, 3]

Without the vocabulary, you cannot know whether 17 meant cat, blue, punctuation, or something else. The IDs become interpretable only when you know the mapping that produced them.

Later, the model will use each token ID to select a row from an embedding table. An embedding row is a learned vector of numbers that can carry useful numerical features. The token ID itself remains only the lookup address.

So keep these two stages separate:

token ID = which row to use
embedding = learned numerical values stored in that row

Trace the conversion in the Lab​

You do not need to understand every Python line. Focus on the sentence at the top and the five labeled outputs.

  1. Click Run without editing anything.
  2. Read text:, characters:, pieces:, vocab:, ids:, and decoded pieces: from top to bottom.
  3. Confirm that the pieces for Cats nap. are ['cats', 'nap', '.'].
  4. Find the first editable line:
text = "Cats nap."
  1. Change only the text to:
text = "Dogs run!"
  1. Before running, predict which outputs must change. The characters, pieces, vocabulary, IDs, and decoded pieces should all be rebuilt from the new sentence.
  2. Click Run and trace the new sentence through the same outputs in order.
  3. Restore text = "Cats nap." before moving on.

Loading lab…

The important habit is to debug from the earliest step rather than jumping to the final numbers:

original text
→ pieces
→ vocabulary lookup
→ IDs

If the pieces are already wrong, changing the model will not repair the tokenizer. If the pieces are right but the IDs are unexpected, inspect the vocabulary mapping next.

Quick Check

1. What is the main job of a token ID?
2. Why must the vocabulary mapping be saved with encoded data?
3. The IDs look wrong. What should you inspect first?

0 of 3 questions answered.

Explain it back​

Explain why [17, 42, 3] does not tell you what the original text was unless you also know the tokenizer's vocabulary mapping.

Key Takeaways

  • A tokenizer converts text into pieces called tokens.
  • A vocabulary maps token pieces to integer token IDs.
  • Token IDs are lookup addresses, not measurements of meaning.
  • Embeddings later provide learned numerical features for those IDs.
  • Debug text conversion in order: source text → pieces → mapping → IDs.

Next Lesson

Next, in L4.2 — Characters, Words, and Subwords, you will compare three ways to choose the pieces themselves and see how that choice changes sequence length and vocabulary needs.

References

Lesson actions

Completion is stored locally on this device.

View progress