본문으로 건너뛰기
L6.7

Checkpoints and Reproducibility

Goal

Define a checkpoint as model state plus experiment context and identify the state needed to resume, compare, and explain a training run.

The training-loop lesson produces changing model parameters. Saving only those weights is enough for some inference use cases, but it is not enough to reconstruct the training experiment.

A useful run checkpoint records at least:

model weights
optimizer state
training step
model configuration
tokenizer/vocabulary revision
data split identity
seed/random-state policy
software versions
recent train/validation metrics

Why optimizer state? Adaptive optimizers remember running statistics. Restoring weights but resetting those statistics can change the very next update.

Three different goals need different saved state​

It helps to separate three use cases.

Inference: load model weights and the exact model/tokenizer configuration needed to turn input IDs into the same computation.

Evaluation: also restore the held-out data identity, preprocessing, prompts, decoding settings, and metric procedure so reported results can be reproduced.

Training resume: additionally restore optimizer state, current step, and relevant random-state policy so the next update continues from the same training process rather than starting a new optimizer history.

A file called best.pt does not tell you which of these goals it supports. The surrounding metadata does.

Fingerprints are useful because filenames are mutable labels. A content hash or other stable artifact identity lets a report point to the exact checkpoint evaluated, even if someone later renames or copies the file.

Reproducing inference and reproducing training are different promises​

For inference, you may need enough state to rebuild the model, tokenizer, and decoding setup.

For exact training continuation, optimizer state matters because adaptive optimizers remember statistics from previous updates. Random state can matter too when dropout, data shuffling, or sampling is involved.

Therefore be precise about the claim:

  • “This checkpoint reproduces the model's evaluation output under these fixed settings.”
  • “This checkpoint can resume the training run from step N.”

The second claim requires more saved context.

Connect checkpoints to evidence​

If an evaluation report says validation loss is 2.31, record which checkpoint produced it.

Otherwise files such as step-100.pt, step-200.pt, and best.pt can become ambiguous. Artifact identity, configuration identity, and metric records should point to one another.

Predict

Which missing checkpoint item can most directly change the next adaptive-optimizer update after resume?

Save evidence for selection as well as restoration​

The Lab turns a small checkpoint record into a fingerprint: a short code computed from every field. If any field changes, the fingerprint changes.

  1. Click Run and record fingerprint: df0179d96d32eb13.
  2. Click Run again without editing. The fingerprint is identical, because the record is identical.
  3. Change only one field: "validation_loss": 2.70, becomes "validation_loss": 2.71,.
  4. Before running, predict whether such a tiny change will change the fingerprint.
  5. Click Run. The fingerprint becomes 3fd3e5d1eb62a4fd, completely different. A fingerprint cannot tell you what changed, only that something changed—so keep the full record next to it.
  6. Press Reset afterward.

Loading lab…

Avoid one mutable best.pt with no history. Record the validation criterion and score that selected a checkpoint so another person can explain why that artifact became “best.”

A fixed seed controls some randomness but cannot substitute for fixed data splits, tokenizer/configuration identities, or documented library/hardware nondeterminism.

Quick Check

1. Which package is sufficient to resume and audit a training run?
2. Does one fixed seed guarantee bit-identical results on every platform?
3. Why record tokenizer identity with the checkpoint?

0 of 3 questions answered.

Explain it back​

Explain why “I saved the model” is an incomplete reproducibility claim. Name the state needed both to resume updates and to reproduce evaluation.

Key Takeaways

  • A checkpoint is part of a run record, not just a weights filename.
  • Optimizer state matters for faithful training resumption.
  • Config, data/tokenizer identity, seed policy, step, metrics, and versions make evidence interpretable.
  • Record why a checkpoint was selected.

Next Lesson

Next, compare training and held-out loss to decide whether continued fitting is improving generalization.

References

Lesson actions

Completion is stored locally on this device.

View progress