Pretrained Models
Goal
By the end of this lesson, you can describe a pretrained model as a reproducible contract—architecture, weights, preprocessing, label mapping, provenance, and evaluation mode—not merely a file of parameters.
Weights only make sense inside the computation that uses them
Imagine a saved encoder whose first layer expects 64 input values and produces 32 hidden values.
If you build a new encoder with 24 hidden units and try to load the old weight tensor, the shapes no longer match.
That visible error is useful. It prevents one set of learned numbers from being silently interpreted under a different architecture.
But matching shapes are only the beginning.
The same weights can fail under wrong preprocessing
Suppose source training normalized digit pixels into a particular range.
At reuse time, you add 2.0 to every normalized feature before passing it to the encoder.
The architecture and state dictionary still match. The model loads successfully. Yet the input distribution reaching the weights is now different from the one the encoder was trained to interpret.
This can shift activations and damage validation behavior.
So a pretrained-model contract includes at least:
- architecture/configuration;
- exact weight revision/checksum;
- preprocessing and input shape;
- label or vocabulary mapping when applicable;
- training/source provenance when available;
- evaluation mode and expected output semantics.
Inspect the notebook model contract
The notebook trains a small source encoder on scikit-learn digits and reloads it.
- Open the notebook and run the code cell. It trains a small encoder on the digits data, saves its weights with
saved = copy.deepcopy(model.encoder.state_dict()), and reloads them into a freshEncoder(32)calledgood. - Read
mean_feature_shift_from_wrong_preprocessing=.... The notebook feeds the same validation images twice: once correctly (val) and once with every pixel shifted up by2.0(val+2.0). The weights never changed, yet the encoder's features moved by more than0.1on average. Wrong preprocessing alone changed what the model “sees.” - Read
shape_mismatch_detected=True. The notebook deliberately tried to load the saved 32-unit weights intoEncoder(24). PyTorch refused with a shape error, and the notebook recorded that refusal instead of hiding it. - Now make the preprocessing mistake smaller. Find
good(val+2.0)and change2.0to0.5. Before running, predict whether the feature shift should grow or shrink. Run the cell and compare. - Change the offset back to
2.0. The lesson is that the checkpoint, its input preprocessing, and its architecture form one contract: change any one of them and the saved weights no longer mean what they did.
Loading lab…
Loading successfully is necessary, not sufficient
A model can load cleanly and still be used incorrectly because:
- class IDs changed order;
- tokenizer/vocabulary mapping changed;
- pixel normalization changed;
- evaluation mode differs;
- an incompatible data population is being treated as equivalent.
Successful deserialization only proves that certain structural checks passed.
Provenance is part of engineering evidence
For an externally released model, reproducibility and responsible use may also require:
- repository revision;
- license;
- training-data description where available;
- known limitations;
- intended/unsupported uses.
A pretrained model is an artifact with a history and input/output assumptions.
A pretrained weight file is not a complete model description
Suppose you download an encoder checkpoint.
To reproduce its behavior, you may also need:
- the exact architecture;
- input size and channel order;
- normalization constants;
- label mapping for the original task;
- library/version assumptions.
A tensor file can load successfully while the surrounding preprocessing is wrong.
Provenance answers “what exactly am I reusing?”
Useful provenance records include where the checkpoint came from, which version it is, what data/task it was trained for, and any license or usage constraints.
This matters technically as well as administratively. A similarly named checkpoint from another release may have different preprocessing or parameter shapes.
Treat pretrained artifacts like dependencies with identities, not anonymous bags of useful weights.
Quick Check
Key Takeaways
- A pretrained model is more than a weight file.
- Architecture and parameter shapes must match.
- Preprocessing and label/vocabulary mapping are part of model meaning.
- Provenance, revision, and limitations support reproducible use.
- Successful loading does not prove correct evaluation.
Next Lesson
Next, you will allow the pretrained representation itself to move and measure whether fine-tuning helps or damages the target task.
References
Completion is stored locally on this device.