Evaluate a Tiny Language Model
Goal
Compare checkpoints with held-out loss, fixed generation prompts, simple failure indicators, provenance, and an explicit limitations statement.
One generated sample has high variance: it can look unusually good or bad because of the prompt and random draw. One average loss answers a narrower next-token question and can hide behavioral failures. A useful tiny-model evaluation combines both kinds of evidence under fixed conditions.
Example scorecard:
checkpoint val_loss repetition fixed-prompt notes
step-40 2.91 low varied but rough
step-60 2.70 low best balance
step-80 2.86 high starts looping
Loss and behavior answer different questions
Imagine checkpoint A has validation loss 2.70 and checkpoint B has 2.75. That small loss difference tells you something about average next-token prediction on the held-out set. It does not tell you, by itself, whether either checkpoint loops on a specific prompt or follows a simple requested format.
Conversely, one attractive generated paragraph is weak evidence because sampling has variance. A checkpoint can get lucky once.
That is why a useful evaluation bundle contains:
- held-out likelihood evidence for average prediction quality;
- fixed behavioral cases for observable generation properties;
- failure indicators such as repetition or prompt drift;
- provenance so the exact checkpoint and decoding configuration are known;
- limitations stating what the small test set cannot establish.
The goal is not to combine every metric into one magic score. It is to keep distinct questions visible so a decision can be explained later.
Separate capability questions from reliability questions
A tiny language model may be asked two different kinds of questions.
Capability: can it continue the small training domain at all?
Reliability: how consistently does it avoid loops, preserve formatting, or stay on-topic under fixed test prompts?
Validation loss helps with average token prediction, but reliability often needs behavioral checks.
For example, create a fixed suite:
prompt A → check maximum repeated phrase count
prompt B → check whether output stays within expected vocabulary/domain
prompt C → check whether required delimiter appears
The checks do not need to pretend the tiny model is a production assistant. They should measure behaviors the current level actually taught.
Record uncertainty in the conclusion
A small validation set and a handful of prompts cannot establish broad language-model quality.
A strong report says what the evidence supports and what it does not.
For example:
Checkpoint B had lower validation loss on this held-out split and fewer repetition failures on five fixed prompts. This does not establish better performance on unrelated text domains.
That is more useful than an oversized conclusion such as “B is the better language model.”
Predict
Keep metric questions separate
Held-out loss measures next-token predictive fit on unseen text. Perplexity is exp(loss) when compatible natural-log loss is used, so lower compatible loss means lower perplexity. Generation cases expose observable behaviors that an average likelihood does not describe.
The Lab prints one line per checkpoint with its held-out loss, perplexity, a repetition score (the same repeated-bigram rate as L6.11), and the sample text.
- Click Run. The best held-out loss is step
60(2.7, perplexity14.88), and the last line confirmsselected by held-out loss: step 60. - Notice step
80: its loss (2.86) looks only a little worse, but its sample loops, withrepetition 0.714. Loss and sample behavior answer different questions. - Change only the step 60 sample to
"the cat learned the cat learned". - Before running, predict which numbers will change: loss, perplexity, repetition, or the selected step?
- Click Run. Only step 60's repetition changes, to
0.4. Its loss is still the lowest, so it is still selected. A metric can only respond to the evidence it measures. - Press Reset afterward.
Loading lab…
Do not compare perplexity casually across different tokenizers or datasets; the token units and evaluation distribution have changed.
Quick Check
Make the decision rule explicit
Before inspecting a final comparison table, state how validation loss, repetition/failure evidence, and fixed prompt cases will affect checkpoint selection. Then document what this tiny evaluation cannot establish.
Key Takeaways
- Evaluation is a fixed set of questions, inputs, metrics, and decision rules.
- Combine held-out likelihood evidence with fixed behavioral cases.
- Perplexity comparisons require compatible tokenization/data setups.
- Record provenance and limitations with results.
Next Lesson
Next, package the chosen run so another learner can restore the artifact and reproduce its evaluation rather than only view its outputs.
References
- Radford et al., Language Models are Unsupervised Multitask Learners.
Completion is stored locally on this device.