본문으로 건너뛰기
L6.13

Package the Training Run

Goal

Package code, configuration, artifact identity, tokenizer/data provenance, metrics, versions, and commands so another learner can inspect and reproduce a tiny-language-model run.

A checkpoint is the product of an experiment, but the binary alone does not explain the experiment. A useful handoff contains both the artifact and the recipe/evidence around it.

A small run package might contain:

run/
README.md
config.json
metrics.json
samples.json
checkpoint fingerprint/metadata
tokenizer + data-split identity
environment/version record
train/evaluate/validate commands

Large binary weights do not need to live in Git to be reproducible; their stable identity, storage/retrieval reference, and provenance do.

Reproducibility is a path another person can follow​

Think of the run package as a sequence of claims another learner should be able to verify:

this code revision
+ this model/tokenizer config
+ this data split
+ this checkpoint identity
+ this evaluation command
→ these reported metrics and samples

If any arrow is missing, the result becomes harder to audit.

For example, a metrics.json file saying validation_loss: 2.4 is incomplete if it does not identify which checkpoint and held-out split produced the number. A checkpoint hash without the matching tokenizer identity is also incomplete because the same integer token IDs might decode differently under another vocabulary.

Portable packages also avoid machine-specific assumptions. Prefer relative repository paths, version records, and commands that can be rerun from a documented working directory. Secrets should be provided through safe external configuration, never copied into the reproducibility record.

Think of reproducibility as a graph of identities​

A run package contains related artifacts:

tokenizer
data/split
model config
checkpoint
evaluation config
metrics
samples
code revision

Each result should be traceable to the inputs that produced it.

For example, metrics.json should identify the checkpoint and evaluation setup. Checkpoint metadata should identify the model configuration and tokenizer. The run record should identify the code revision and data split.

Without those links, individually valid files can be combined incorrectly.

“Works on my machine” usually hides an unstated dependency​

Common hidden dependencies include a package version, local absolute path, manually downloaded data, environment variable, tokenizer artifact, or unrecorded command-line option.

Packaging the run is therefore part of engineering, not administrative cleanup after modeling is finished.

Predict

Which package is most reproducible?

Validate the package itself​

This Lab runs locally with the Python standard library. From the repository root, run:

python labs/notebooks/level-06/l06-13-package-run.py projects/reference/l06/sample-run

The checker also accepts the JSON file directly:

python labs/notebooks/level-06/l06-13-package-run.py projects/reference/l06/sample-run/run.json

Both commands should end with PASS: packaged run p06-reference-reduced has required reproducibility evidence.

Now simulate one missing field without editing the committed fixture. Before running, predict which required-field error should appear if only the seed is removed and which parts of the package should remain otherwise unchanged:

python labs/notebooks/level-06/l06-13-package-run.py projects/reference/l06/sample-run --drop seed

Confirm the command fails with missing required field(s): seed. Then make a separate one-variable prediction for --drop tokenizer_fingerprint: the checker should reject tokenizer lineage rather than the seed. Run that command and explain why correct weights without tokenizer identity are still insufficient to reproduce the run.

Loading lab…

When reproduction fails, classify the missing contract: code revision, environment, tokenizer/data identity, config, artifact, command, or evaluation input. Add the smallest portable evidence needed; do not encode secrets or machine-specific absolute paths.

Quick Check

1. Why keep configuration with checkpoint metadata?
2. Should secrets be committed for reproducibility?
3. What is Git especially suitable for here?

0 of 3 questions answered.

Explain it back​

A reviewer has only the run directory and artifact reference. Explain the minimum path they should follow to validate the checkpoint and reproduce the reported metrics without your original notebook session.

Key Takeaways

  • A shipped run includes artifact identity plus the experiment context around it.
  • Commands and manifests should be portable and repeatable.
  • Separate large-artifact storage from versioned metadata when appropriate.
  • Never trade away security by committing secrets for convenience.

Next Lesson

Next, test the complete chain of contracts from tokenized data through checkpoint evaluation and run-package validation.

References

Lesson actions

Completion is stored locally on this device.

View progress