Transfer Learning
Goal
By the end of this lesson, you can explain what transfer learning reuses, compare it fairly with training from scratch, and identify conditions where transferred features may help or hurt.
Reuse a learned representation instead of starting from zero
Suppose a model has already learned an encoder for recognizing handwritten digits.
The encoder turns an image into useful internal features. Now you have a related target task—perhaps distinguishing digits with loops from digits without loops—but only 40 labeled target examples.
You have two starting strategies:
- scratch: learn the representation and target classifier from those 40 examples;
- frozen transfer: reuse the pretrained encoder and learn only a new target head.
Transfer learning asks whether earlier learned features reduce how much the target task must relearn.
Why transfer can help
If source and target tasks depend on overlapping visual structure, the pretrained encoder may already detect strokes, curves, and shapes useful for the new target.
With few target labels, reusing those features can be more data-efficient than learning everything from random initialization.
But this is a hypothesis, not a guarantee.
If source features emphasize the wrong information, transfer can provide little benefit or even hurt. That is negative transfer.
Keep a scratch baseline
Without a scratch baseline, “pretrained worked” tells you almost nothing.
A fair comparison should keep fixed:
- the same target examples;
- the same validation set;
- preprocessing;
- seed policy;
- metric.
Then the main changed factor is the starting representation/training strategy.
Run the notebook transfer comparison
The notebook uses the real scikit-learn digits dataset and a fixed seed of 13.
- Run the notebook unchanged.
- Identify the source task: 10-class digit recognition. Identify the target task: loop-like digits
{0, 6, 8, 9}versus the other digits. - Compare
frozen_transfer_accuracyandscratch_accuracyon the same validation set. - Find this line in the setup cell:
small_idx, _ = train_test_split(all_idx, train_size=40, random_state=SEED, stratify=y_target_all)
- Change only
train_size=40totrain_size=20. LeaveSEED, the validation split, model definitions, epochs, and learning rates unchanged. - Before rerunning all dependent cells, predict whether the smaller target-training set could make the reusable pretrained representation more valuable relative to scratch training. Treat that as a hypothesis, not a promised outcome.
- Rerun from the data/setup cell through the comparison cell and record both accuracies again.
- Restore
train_size=40when the comparison is complete.
Loading lab…
After the guided pass, explain why changing both the target sample count and the validation set at the same time would make the comparison harder to interpret.
Do not let transfer hide bad supervision
If target labels are shuffled, a pretrained encoder cannot turn incorrect supervision into a trustworthy target model.
Transfer changes the starting representation. It does not repair label semantics, data leakage, or an invalid evaluation split.
A useful decision question
Ask:
Does the source representation contain information the target needs, and does controlled target evidence show a benefit over scratch?
That question is stronger than “Is the pretrained model famous?” or “Did training finish?”
Quick Check
Key Takeaways
- Transfer learning reuses parameters and representations learned from earlier data.
- Frozen transfer keeps the encoder fixed and learns a new target head.
- Transfer can help when source and target needs overlap, especially with limited target data.
- A scratch baseline is required to test transfer rather than assume it helps.
- Pretraining does not fix bad labels, leakage, or an unfair evaluation.
Next Lesson
Next, you will inspect what must accompany pretrained weights so the reused representation is actually the one you think it is.
References
Completion is stored locally on this device.