Training Loop
Goal
Train the tiny decoder-only model on next-token prediction and use loss behavior plus intermediate evidence to debug the loop.
Training turns an architecture into a learned model
The model from the previous lesson can compute logits, but its initial weights do not yet encode useful next-token patterns.
Training repeats a small cycle:
sample batch
↓
forward pass
↓
compute loss
↓
backward pass
↓
optimizer step
↺
The data also has a simple alignment rule. If the token window is:
the model predicts the next
then one training pair is:
input: the model predicts the
target: model predicts the next
Each input position is trained to predict the token immediately after it.
Predict
Read the optimizer loop in order
A minimal PyTorch loop looks like this:
for step in range(steps):
x, y = sample_batch(...)
_, loss = model(x, y)
optimizer.zero_grad(set_to_none=True)
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
optimizer.step()
The order matters:
zero_gradclears gradients left from the previous step;loss.backward()computes how the current loss depends on the parameters;- gradient clipping limits unusually large gradient norms in this teaching loop;
optimizer.step()uses the current gradients to update the parameters.
If you call step() before backward(), there are no current gradients to apply.
Run the training loop and collect evidence
- Open the notebook Lab and run the training cells unchanged.
- Find the first reported loss and the best/later loss reported by the notebook.
- Confirm that every reported loss is finite.
- Confirm that the best loss near the end is lower than the first loss.
- Keep the seed and data fixture unchanged for this first pass so the run is reproducible.
Loading lab…
The default teaching run uses CPU, a small repeated corpus, explicit seeds, and a short training schedule. The success condition is behavioral, not an exact floating-point target: the loop should remain numerically stable and show evidence that optimization is reducing the training objective.
A falling training loss does not prove the model generalizes or generates good text. L6.8 — Validation Loss and Overfitting will separate training behavior from held-out validation evidence.
One failure at a time
After the baseline run works, choose one failure experiment rather than changing several controls together.
A useful first guided failure is the learning rate:
- Find the notebook's learning-rate setting.
- Record its original value.
- Make one clearly larger learning-rate change in a copy of the run, leaving seed, data, batch construction, and model configuration fixed.
- Observe whether loss becomes unstable, oscillates, or fails to improve.
- Restore the original value afterward.
Other failures worth testing later include unshifted targets, invalid token IDs, missing gradient clearing, or an impossible block size. Keep them separate so each result has one main cause.
Debug from structure before tuning
If loss does not improve, check these boundaries in order:
- data alignment — are targets shifted by one token?
- ID range — are every input and target ID inside the vocabulary?
- shape — do logits still have
(B, T, V)? - finite values — is the loss a real finite number rather than
NaNorinf? - gradient flow — do model parameters receive gradients?
- parameter updates — do values actually change after
optimizer.step()? - learning rate — only after the earlier checks pass, ask whether update size is sensible.
This order prevents expensive hyperparameter guessing from hiding a simple structural bug.
Under the Hood
Near random initialization, if the model gives roughly equal probability to every vocabulary item, cross-entropy is often in the neighborhood of log(vocab_size). Here log is the natural logarithm. Treat this only as a rough reference scale, not a value the first loss must match exactly.
Explain it back
Explain why these are different claims:
- “training loss decreased”;
- “the model learned something useful on unseen text.”
The first is optimization evidence on the training objective. The second needs held-out evaluation.
Quick Check
Key Takeaways
- Next-token targets are the input sequence shifted by one position.
- Training repeats forward, loss, backward, and parameter update steps.
- Gradient clearing and update order are required parts of the training loop.
- Reproducible evidence needs fixed seeds and data fixtures.
- Falling training loss is optimization evidence, not proof of held-out language quality.
- Debug structural boundaries before tuning the learning rate.
Next Lesson
Next, L6.7 — Checkpoints and Reproducibility makes a successful training run recoverable by saving checkpoints together with the configuration and provenance needed to reproduce it.
References
Completion is stored locally on this device.