Catastrophic Forgetting and Drift
Goal
Recognize adaptation regressions on previously useful behavior, distinguish targeted improvement from broad drift, and design retention checks that reveal forgetting.
Fine-tuning pushes the model toward the new training objective. That can improve the target behavior while moving other behavior unintentionally. Two useful terms are: catastrophic forgetting — previously learned capability degrades substantially after new training; behavioral drift — outputs shift in a way that may be smaller, broader, or simply different from the intended adaptation.
The boundary between these labels is not always sharp. The important engineering job is to measure what changed.
A small target gain can hide a large old-task loss
Imagine:
target-task accuracy:
base 70%
adapted 88%
general extraction:
base 91%
adapted 74%
The adaptation clearly improved its target. It also damaged a capability that may matter in the product. Calling the run “better” without the retention result would be incomplete.
Forgetting is about behavior, not parameter distance alone
Two checkpoints can have a small parameter difference and a large output change on one sensitive slice. A large parameter difference can also leave a particular behavior mostly unchanged. Parameter movement is diagnostic evidence. Retention evaluation is behavioral evidence.
Do not substitute one for the other.
Data imbalance can pull behavior
Suppose 95% of fine-tuning examples use terse one-line answers. Even if the stated goal is only domain terminology, the model may also learn a strong brevity bias. That is drift created by the dataset distribution. Inspect not only labels but also:
- response length;
- formatting;
- tone/style;
- domain mix;
- missing-information frequency;
- refusal/abstention rate.
Mitigation is an experiment, not a slogan
Possible responses include:
- reduce update magnitude/steps;
- improve data balance;
- mix retention examples into training;
- lower adapter capacity;
- change target modules;
- early stop using held-out evidence;
- reject the adaptation.
Which response works depends on the measured failure. Change one factor at a time where practical and rerun the same target/retention suites.
Track drift against the correct baseline
If adapter v2 was trained from the original base, compare against that base and v1 separately when both matter. If v2 was trained on top of v1, the lineage is different. Artifact lineage affects interpretation. A training record should say exactly which model state each adaptation started from.
Drift becomes visible only when you define what should stay stable
Adaptation is supposed to change behavior, so “the outputs changed” is not evidence of forgetting. Catastrophic forgetting or harmful drift means behavior that was previously useful has degraded beyond an acceptable boundary. That requires a retention set and a threshold chosen before looking only at the adapted model's best results.
Slice retention checks by capability when possible. A model might preserve extraction accuracy while becoming much worse at abstention, or keep task accuracy while drifting into an unwanted response style. If a mitigation is tested, keep the evaluation fixed so you can attribute the recovery to the mitigation rather than to a changed benchmark.
Predict
Break and diagnose the Browser Lab model
The Browser Lab contains a toy multi-task score table.
- Click Run once.
regressions: []is empty, but one check fails, because the TODO is not written yet. - Read the scores.
target_formatrose from0.70to0.88, butgeneral_extractionfell from0.91to0.74: the adapter got better at its target and forgot something else. - Complete
regression_flags: return every metric whose adapted score is lower than its base score by more thanmax_drop. - Click Run again. You should see
regressions: ['general_extraction'].abstentiondropped only0.01, which is inside the allowance. - Change
max_allowed_drop = 0.02tomax_allowed_drop = 0.0and run. Nowabstentionis flagged too. The threshold is a policy choice, so write it down before you look at results. - Restore
0.02. Now imagine a gentler adaptation (smaller updates or a more balanced training mix) and changegeneral_extractionto{"base": 0.91, "adapted": 0.90}. Run again: no regressions are flagged.
Loading lab…
Quick Check
Explain it back
Give one example of a desired adaptation and one retained capability that could regress. Define a measurable threshold that would make you stop or revise the run.
Key Takeaways
- Adaptation can improve the target while damaging older behavior.
- Retention evaluation is stronger than parameter distance alone.
- Data distributions can create unintended style or behavior drift.
- Mitigations should be tested under fixed evaluation.
- Model/adaptor lineage must be recorded.
Next Lesson
Next, package the evidence into a model card and training record that another person can audit.
References
Completion is stored locally on this device.