Adaptation Review: Choose, Fine-Tune, Explain
Goal
By the end of this lesson, you can choose among scratch training, frozen transfer, and fine-tuning using a controlled scorecard that includes quality, trainable parameters, representation movement, failure slices, and reproducibility.
Adaptation is an experimental choice, not a prestige ranking
Three strategies can solve the same target task:
- scratch: learn encoder and head from target data;
- frozen transfer: reuse encoder and learn only a head;
- fine-tune: start from transferred weights and update encoder too.
The most advanced-sounding option is not automatically best.
Your decision should depend on target evidence and engineering tradeoffs.
Make the strategy the main changed factor
A fair comparison keeps fixed:
- target train/validation examples;
- preprocessing;
- seed policy;
- metric;
- evaluation set.
Then compare:
| Evidence | Scratch | Frozen transfer | Fine-tune |
|---|---|---|---|
| validation quality | ? | ? | ? |
| trainable parameters | many | few | many/some |
| encoder movement | learned from zero | zero | measured |
| failure slices | ? | ? | ? |
| compute/reproducibility cost | ? | ? | ? |
A scorecard makes it harder to choose based on one attractive accuracy number.
Run all three strategies under the same target evidence
- Open the notebook. Before running, confirm that all three strategies train on the same small target set:
aandbcome fromsmall, which holdstrain_size=40examples, and all three are scored on the sameval. - Run the cell. It prints one table:
strategy | accuracy | trainable_parameters | encoder_movement. - Compare the
trainable_parameterscolumn. The frozen model trains only its small head, while scratch and fine-tuning train the whole network. - Read
encoder_movementforfine_tuned. It shows how far fine-tuning moved the pretrained encoder; the other two rows show0.000because that measurement does not apply to them. - Now change only the amount of target data. Find
train_size=40and change it totrain_size=12. - Before rerunning, predict which strategy should suffer most with fewer target labels. (Hint: which one has no pretrained knowledge to fall back on?)
- Run the cell again and compare the whole table with your first run. The notebook's check requires every strategy to stay above
0.75; with only 12 labels a strategy may fall below that and raise anAssertionError. That is evidence, not a bug in your edit—record which strategy failed. - Restore
train_size=40afterward.
Loading lab…
Similar quality can justify the simpler adaptation
Suppose:
- frozen transfer: 91.5% validation accuracy;
- fine-tuning: 91.8%;
If fine-tuning trains far more parameters, takes longer, moves the encoder substantially, and provides no meaningful robustness improvement, frozen transfer may be the better engineering choice.
The 0.3-point difference is evidence, but it is not the only relevant evidence.
Different evaluation conditions destroy the comparison
If scratch uses one split and fine-tuning uses an easier split, the resulting table may look precise but cannot isolate strategy effect.
Audit the experiment before tuning the models.
State a decision with limits
A useful conclusion might say:
“Under this 40-example target split, frozen transfer is preferred because it matches fine-tuning within the observed validation variation while training fewer parameters and keeping the encoder fixed. We would retest the choice if target data grows or the failure-slice gap changes.”
That is stronger than “frozen is best.” It says what evidence supports the decision and what could change it.
Make the choice from a scorecard, not from the method name
Suppose you compare three strategies:
scratch
frozen transfer
fine-tuning
A useful scorecard can include validation quality, trainable parameter count, training time, representation movement, and important failure slices.
You may find that fine-tuning improves average accuracy slightly but costs far more compute and worsens one important slice. Or frozen transfer may nearly match it with much lower operational cost.
The lesson is not that one strategy should always win. The lesson is to make the decision traceable to evidence.
State the conclusion narrowly
Prefer:
On this target dataset and validation split, fine-tuning improved the chosen metric by X while changing Y parameters.
Avoid:
Fine-tuning is better.
A narrow conclusion records the conditions under which the evidence was collected and makes it easier to revise the decision when the data, model, or operational constraints change.
Quick Check
Key Takeaways
- Scratch, frozen transfer, and fine-tuning are strategies to compare, not a universal ranking.
- Hold target evidence and evaluation conditions fixed.
- Compare quality with trainable parameters, encoder movement, robustness, and reproducibility cost.
- A simpler adaptation can be preferable when quality is effectively similar.
- State decisions narrowly enough that new evidence can revise them.
Next Lesson
Complete Transfer Learning Across a Real Dataset. Then Level 4 turns raw text into model-ready tokens and embeddings.
References
- PyTorch, Transfer Learning for Computer Vision Tutorial.
- PyTorch, nn.RNN.
Completion is stored locally on this device.
Level project unlocked: Transfer Learning Across a Real Dataset