본문으로 건너뛰기
L6.2

Next-Token Prediction

Goal

Create one-token-shifted input/target pairs and explain how cross-entropy turns (B,T,V) logits into a training signal for the correct next-token class.

Let the text provide its own labels​

A language-model corpus does not come with a teacher writing a separate answer beside every token. So where do the training targets come from? The text already contains them: at each position, the next token is the answer the model should learn to predict.

Ordinary tokenized text becomes supervised training data with one simple shift:

tokens: [A, B, C, D, E]
input: [A, B, C, D]
target: [B, C, D, E]

At the position containing C, the correct class is D. The model is not asked to reconstruct C or predict the entire future at once.

Read the shift as aligned columns​

Write the source twice, offset by one position:

source: A B C D E
input: A B C D
target: B C D E

Now read vertically. The logit vector produced from input position A is judged against class B; the vector produced from B is judged against C, and so on.

An off-by-one error changes the task, not just the bookkeeping. If input and target are identical, the model is rewarded for identifying the token already present at that position. If targets are shifted by two, it learns a two-step-ahead objective instead of the intended immediate next token.

Cross-entropy then answers a local question at every valid position: how much probability did the model assign to the correct next-token class relative to all vocabulary alternatives? If the vocabulary has V=1000, each input position produces 1,000 logits and its shifted target ID selects which class should receive high probability. One short sequence therefore contributes several classification decisions to the loss.

The target shift is a semantic contract​

A code path can have perfectly legal tensor shapes while pairing the wrong target with each position.

Always inspect a tiny hand-written example before trusting a large batch pipeline. Print the input IDs and shifted target IDs side by side and verify that every column means “use this prefix position to predict the token immediately after it.”

Predict

For token IDs `[7,4,9,2]`, what targets align with input `[7,4,9]`?

Check alignment before loss curves​

  1. Click Run once. The shifted TODO returns the same list twice, so the output shows input : [7, 4, 9, 2] and target: [7, 4, 9, 2], and one Lab check fails. That failure is expected.
  2. Before editing, write on paper what the input and target lists should be for [7, 4, 9, 2]: each target is the token that comes right after its input position.
  3. Complete shifted so it returns those two lists for any sequence of IDs. Use slicing rather than typing the numbers.
  4. Click Run again. You should see input : [7, 4, 9] and target: [4, 9, 2], and every check should pass.
  5. Change shifted([7, 4, 9, 2]) to a five-token example of your own and verify that each target is the immediate next token.

Loading lab…

Cross-entropy compares each logits vector with one correct vocabulary class. Standard implementations accept logits directly and perform the numerically stable normalization internally; applying a separate softmax first is unnecessary.

A dangerous bug is using identical input and target rows. The code can train, but it rewards copying the current token instead of learning the next-token objective.

Quick Check

1. Why shift targets by one token?
2. For source `A B C D`, what target aligns with input token `C`?
3. What does cross-entropy encourage?

0 of 3 questions answered.

Explain it back​

Show an off-by-one bug with three or four tokens and explain exactly which wrong behavior it would reward.

Key Takeaways

  • Input and target rows come from one source sequence shifted by one.
  • Each time position is a vocabulary classification problem.
  • Cross-entropy turns next-token logits into an optimization signal.
  • Verify tiny target alignment before trusting any training metric.

Next Lesson

Next, turn long token streams into many bounded examples while keeping held-out text outside the training region.

References

Lesson actions

Completion is stored locally on this device.

View progress