본문으로 건너뛰기
L6.0

Level 6 — Build and Train a Tiny Language Model

Level 5 built a small Transformer that can turn each token position into scores over a vocabulary. Level 6 gives those scores a concrete job: predict the next token.

Start with one next-token example​

Suppose the text begins:

the cat sat

A language model can be asked:

Given the tokens seen so far, which token should come next?

It might assign raw scores to possibilities such as:

on → high score
under → lower score
banana → very low score

Those raw output scores are called logits. They are not probabilities yet. Later, a conversion such as softmax turns them into a distribution that can be used to choose a next token.

A decoder-only language model is a model built around this repeated next-token task. “Decoder-only” names the Transformer architecture used here; “language model” describes the prediction job.

The model is causal, which means a prediction may use the prefix that has already appeared, but it must not use future tokens that would reveal the answer.

During training, we already have text, so we can create many input/target pairs from it:

input: the cat
answer: sat

The model makes a prediction, we measure the loss, and training changes model parameters.

During generation, there is no known next answer. The model predicts a distribution, one next token is selected, that token is appended to the prefix, and the process repeats.

prefix
→ next-token scores
→ choose one token
→ longer prefix
→ repeat

Sampling text does not train the model again. It uses the parameters that were already learned.

A few words you will meet​

TermPlain meaning in this Level
logita raw model score for one possible output token before conversion to probabilities
next-token targetthe token the model should have predicted at one training position
training loopthe repeated process of predict → measure loss → compute updates → change parameters
validation lossloss measured on held-out text that was not used to update the model
checkpointa saved snapshot of model and training state that can be inspected or resumed later
samplingchoosing a next token from the model's output distribution
temperaturea control that makes a token distribution sharper or flatter before sampling
top-k / top-pcontrols that restrict which token choices remain available before sampling
reproducible runa run recorded with enough code, data identity, settings, seeds, and saved state for someone else to repeat it

A seed is a starting value used by a pseudo-random generator. Keeping the same seed helps make randomized experiments repeatable.

Learning goal​

Train, evaluate, save, sample from, and package a tiny decoder-only language model while keeping the important boundaries visible.

By the end of this Level, you should be able to:

  • explain decoder-only next-token prediction in ordinary language;
  • build aligned input and target windows without letting validation text influence training;
  • state the important input/output shapes before training;
  • save enough state to explain and resume a checkpoint;
  • distinguish improving training loss from improving held-out behavior;
  • explain sampling as repeated next-token selection rather than retraining;
  • compare temperature, top-k, and top-p while keeping other conditions fixed;
  • diagnose repetition, collapse, incoherence, and the model drifting away from a prompt using controlled evidence;
  • compare checkpoints using held-out loss and the same fixed behavior examples;
  • package code, configuration, tokenizer/data identity, metrics, commands, and saved outputs so the run can be reproduced.

The learning path​

L6.1–L6.3 define decoder-only next-token prediction and show how ordinary text becomes aligned training examples.

L6.4–L6.6 define the model configuration, assemble the model, and run the training loop.

L6.7–L6.8 save reproducible checkpoints and compare training behavior with held-out validation behavior.

L6.9 — Sampling Text shows the basic repeated next-token generation loop. After this lesson, complete the Level 6 mini checkpoint.

L6.10–L6.12 separate decoding controls from model training, diagnose common generation failures, and evaluate checkpoints on fixed evidence.

L6.13–L6.14 make the run reproducible and finish with an end-to-end training/evaluation workshop.

A rule that protects this Level​

A sentence that looks plausible is not proof that the training experiment is correct.

Keep the important evidence visible:

source text
→ training/validation split
→ input/target alignment
→ model settings
→ checkpoint identity
→ validation measurements
→ prompt + seed + decoding settings
→ generated output

If the output looks surprising, this chain helps you ask whether the problem came from the data, the model, the saved state, or the way text was sampled.

Level Project​

After L6.14 — Tiny LLM End-to-End Workshop, complete Train and Ship a Tiny LLM.

The project asks for one reproducible training and evaluation run plus a documented failure/debug path. A good submission should make it possible for another learner to tell exactly what was trained, what was measured, how text was generated, and how the result can be repeated.

Lesson actions

Completion is stored locally on this device.

View progress