본문으로 건너뛰기
L3.12

Fine-Tuning a Small Encoder

Goal

By the end of this lesson, you can distinguish frozen transfer from fine-tuning, use conservative encoder updates, and interpret parameter movement together with held-out target performance.

Frozen transfer asks the target head to adapt​

In frozen transfer:

  • the pretrained encoder stays fixed;
  • a new target head learns from target labels.

This is cheap, stable, and a useful baseline.

But the representation may not perfectly separate the new target classes.

Fine-tuning lets the representation move​

Fine-tuning unfreezes some or all pretrained parameters so target loss can update them too.

That gives the model more freedom:

pretrained representation + target evidence -> adjusted representation

More freedom can improve target fit. It can also damage useful source features when updates are too aggressive or target data is too small.

For that reason, fine-tuning often begins with a smaller learning rate than training randomly initialized parameters from scratch.

Measure movement instead of asking only whether code completed​

Suppose two fine-tuning runs end with similar validation accuracy.

Run A moves the encoder only slightly.

Run B moves the encoder dramatically.

That movement is useful evidence. It can reveal that two apparently similar results were reached through very different adaptation behavior and may respond differently to more training or distribution shift.

Parameter movement is not a quality metric by itself. Interpret it together with validation/failure evidence.

Compare frozen and fine-tuned models in the notebook​

  1. Open the notebook. Before running, find the two fine-tuning lines. Both start from a copy of the same frozen model: fine=train(fine,a,b,80,0.01) uses learning rate 0.01, and aggressive=train(aggressive,a,b,80,1.0) uses learning rate 1.0.
  2. Predict which run should move the encoder weights farther from the frozen starting point.
  3. Run the cell and record three lines: frozen_accuracy=..., fine_tuned_accuracy=... movement=..., and aggressive_accuracy=... movement=....
  4. Compare the two movement values. movement is the size of the change in all encoder weights. With seed 13, conservative fine-tuning moves the encoder only about 0.2, while the aggressive run moves it about 9—dozens of times farther.
  5. Compare accuracy and movement together. A large movement means the pretrained features were heavily rewritten. That may or may not hurt accuracy on this small validation set, but it throws away more of what pretraining provided.
  6. Try one controlled retry: change only the conservative learning rate from 0.01 to 0.05 and run again. Because every run copies frozen first, you do not need to reload anything by hand. Predict where its movement will land compared with the other two, then restore 0.01.

Loading lab…

Starting every comparison from the same checkpoint matters. Otherwise one run begins from weights already modified by an earlier experiment.

A common misconception​

“Fine-tuning means pretrained is always improved.”

Fine-tuning means pretrained parameters are allowed to change under a target objective. Whether that change is useful is an empirical question.

If validation gets worse while encoder movement becomes huge, aggressive adaptation is a plausible hypothesis. Test it by lowering the learning rate while keeping the target split and starting checkpoint fixed.

Fine-tuning changes the reused representation itself​

In frozen transfer, the pretrained encoder stays fixed and only the new task head learns.

In fine-tuning, some or all encoder parameters are allowed to move.

That gives the target task more flexibility, but it also creates risk: a small target dataset can push useful pretrained features too far toward its limited examples.

Use smaller updates as a controlled starting point​

A common strategy is to use a lower learning rate for pretrained parameters than for a newly initialized head.

The reason is not that pretrained weights are sacred. It is that they already encode useful structure, so you often want adaptation rather than a complete rewrite.

When evaluating fine-tuning, compare at least:

  • target validation quality;
  • how many parameters were trainable;
  • how far encoder parameters moved;
  • whether important failure slices improved or regressed.

“Fine-tuned” describes a training procedure, not proof that the resulting model is better.

Quick Check

1. What changes during frozen transfer?
2. Why use a smaller learning rate for fine-tuning?
3. What evidence can reveal overly aggressive fine-tuning?

0 of 3 questions answered.

Key Takeaways

  • Frozen transfer learns a target head while keeping the encoder fixed.
  • Fine-tuning allows target loss to modify the pretrained representation.
  • Conservative learning rates often reduce the risk of destroying useful starting features.
  • Compare runs from the same checkpoint.
  • Track representation movement together with held-out target evidence.

Next Lesson

Next, you will stop averaging all failures together and test image and sequence assumptions with explicit evaluation slices.

References

Lesson actions

Completion is stored locally on this device.

View progress