본문으로 건너뛰기
L8.13

Model Cards and Training Records

Goal

Document an adapted model so someone else can tell what it is, how it was produced, how it was tested, where it should be used, and what important limitations remain.

Start with a model file you find six months later​

Suppose you open a project folder and find:

adapter.safetensors

The file may work, but the filename does not answer the questions you need before using it. Which base model does it belong to? Which dataset version trained it? Was it tuned for customer support or code? What evaluation improved? What became worse? Is it safe to compare this artifact with a newer run?

A model card is the learner-facing record of those behavior questions. It describes intended use, evaluation results, important limitations, and conditions that matter when interpreting the results. The goal is not paperwork for its own sake. The goal is to stop a model artifact from becoming an unexplained file that people reuse based on memory.

For an adapted model, a useful card might say that it starts from a particular base revision, was tuned for a narrow support-labeling task, improved that task on a named evaluation set, and should not be treated as a general factual assistant. Specific limits are more useful than a generic warning that “AI can make mistakes.”

Training record: how this artifact was produced​

A training record should identify:

base checkpoint
data version/hash
split policy
format/template version
training method
LoRA/PEFT configuration
optimizer settings
steps/epochs
seed policy
software/runtime versions
output adapter/checkpoint identity
evaluation version

If a real provider/cloud/GPU run is used, record the environment details needed to interpret the result.

Lineage is a graph​

Suppose adapter v2 was trained from adapter v1 rather than the original base.

base
→ adapter-v1
→ adapter-v2

is different from:

base
→ adapter-v2

Even if both artifacts are named “v2,” their starting states differ. Record parent identity explicitly.

Metrics need context​

A line such as:

accuracy = 92%

is weak without:

  • which case set;
  • how many cases;
  • which model version;
  • which prompt/decoding;
  • what metric definition;
  • what slices;
  • what baseline.

Prefer:

support-format-v3: 92/100 exact schema pass
base: 78/100
retention-general-v2: 95/100 base, 94/100 adapted

Now the result is interpretable.

Limitations should be specific​

Weak:

The model may make mistakes.

Stronger:

Evaluation contains 100 English support cases and 20 abstention cases. It does not test Korean inputs, long documents above 4k tokens, or live tool use. The adapter showed a 3-point regression on terse extraction before data rebalancing.

Specific limitations guide the next experiment.

Documentation can reveal missing evidence​

Writing the card often exposes holes:

  • no data version;
  • no held-out split;
  • no base benchmark;
  • no retention suite;
  • unknown adapter target modules.

That is useful. Documentation is part of quality control because it forces artifact relationships to become explicit.

Predict

Which record best identifies an adapted model?

Validate the local training record​

This Lab runs locally with the Python standard library. Start with the complete record:

python labs/notebooks/level-08/l08-13-training-record.py

It should end with PASS: training record and model card agree on lineage, evaluations, and limitations.

Next, predict what should break if the base identity disappears, then run:

python labs/notebooks/level-08/l08-13-training-record.py --drop base_model_id

The validator should reject the package because the training record is incomplete and no longer agrees with the model card. Try the same experiment with --drop retention_eval_id: a training artifact without its retention evaluation cannot support a claim that general behavior was preserved.

Finally, compare a specific limitation with a vague warning:

python labs/notebooks/level-08/l08-13-training-record.py --vague-limitation

The vague limitation is rejected because it does not identify where the evidence stops.

Loading lab…

Quick Check

1. What is the purpose of a model card?
2. Why record parent/base lineage?
3. What makes a limitation useful?

0 of 3 questions answered.

Explain it back​

Write a five-line adaptation identity record containing base, data, method/config, output artifact, and evaluation version. Then write one specific limitation.

Key Takeaways

  • Adapted models need artifact lineage, not just a filename.
  • Model cards describe behavior, intended use, evidence, and limitations.
  • Training records describe how the artifact was produced.
  • Metrics require case-set and baseline context.
  • Documentation is also a quality-control tool.

Next Lesson

Next, integrate the decision, data, adapter, evaluation, regression checks, and documentation into one reviewable adaptation package.

References

Lesson actions

Completion is stored locally on this device.

View progress