Evaluate Before and After Fine-Tuning
Goal
Compare base and adapted models on identical held-out cases, separate target-task gains from regressions, and report paired differences instead of cherry-picked generations.
Fine-tuning is a model change. A model change should be evaluated like an experiment:
same cases
same prompt/context policy
same decoding/evaluator
base model vs adapted model
If several other variables change at the same time, you cannot isolate the effect of adaptation.
Use paired cases
Suppose five target-task cases produce:
case base adapted
A fail pass
B pass pass
C fail pass
D pass fail
E pass pass
Aggregate accuracy:
base = 3/5
adapted = 4/5
The average improved. But the table also shows a regression on case D. That regression identity matters.
Separate target and retention suites
Use at least two suites:
Target suite
- behavior the adaptation is supposed to improve.
Retention suite
- behavior the base model should keep.
Example:
target:
domain-specific support formatting
retention:
general summarization
basic factual extraction
abstention behavior
structured-output compliance
An adaptation can improve target performance while harming retention. Without the second suite, the regression is invisible.
Keep stochastic evaluation controlled
If generation is sampled, use repeated trials or a fixed seed policy. Record:
- model/checkpoint/adapter identity;
- prompt version;
- case-set version;
- decoding settings;
- evaluator version;
- retry policy.
A before/after comparison should differ primarily in the model adaptation being tested.
Use slices, not only one average
Fine-tuning may improve common cases but harm long inputs, minority classes, rare formats, or abstention cases. Report slices tied to plausible failure mechanisms. For example:
short complete cases
missing-information cases
long cases
rare label
instruction-like source text
This turns “adapted model seems better” into a more precise statement.
Decide thresholds before looking at the result
A release rule might be:
target success +8 percentage points or more
AND
retention regression <= 1 percentage point
AND
zero critical safety regression
The exact thresholds depend on the application. Writing them before looking at results reduces the temptation to move the goalposts after seeing a favored checkpoint.
Compute paired deltas, not only two independent averages
For each case, define:
delta = adapted_score - base_score
For a binary pass/fail case, the most informative transitions are:
fail → pass improvement
pass → fail regression
pass → pass retained success
fail → fail unresolved failure
This paired view answers a different question from two separate averages. Two models could both score 80% while succeeding on different cases. An average-only report would call them tied even though 20% of cases may have changed in each direction.
Small evaluation sets have uncertainty
If an adapted model improves from 8/10 to 9/10, that is one additional passing case. It may be a real improvement, but ten examples provide limited evidence about the broader task population. Do not hide small sample size behind percentages. Report counts and, for important decisions, expand the evaluation set or use an appropriate statistical uncertainty method.
For stochastic generation, repeated trials add another source of uncertainty. The evaluation protocol should state whether each case is run once, under a fixed seed, or across multiple samples.
Keep evaluator changes separate from model changes
If you rewrite the rubric after seeing adapted outputs, the comparison can move even when model behavior does not. Version the evaluator or human-review instructions. A clean adaptation experiment changes the model while keeping the measurement boundary stable.
Predict
Complete the evaluation Browser Lab
The Browser Lab contains paired base/adapted case results.
- Click Run once. Every list in the report is empty, so the checks fail. The last line already shows the averages:
pass rate base -> adapted: 0.6 -> 0.8. - Complete
compare_cases: a case is improved if it went from fail to pass, regressed if it went from pass to fail, and otherwise stays inunchanged_passorunchanged_fail. - Click Run again. You should see
'improved': ['A', 'C']and'regressed': ['D']. The average improved, but caseDgot worse. - Add a retention regression: change case
Eto{"id": "E", "base": True, "adapted": False}. - Click Run. The pass rates are now
0.6 -> 0.6—no average change at all—whileregressedlists['D', 'E']. The case-level report shows a trade that the average hides. - Now apply a simple release rule: “no more than one regression.” Does this adaptation pass? Explain which regressed case you would investigate first and why. The Lab should report
Result: experiment ranbecause the original paired-case baseline changed while the transition-classification invariants still pass.
Loading lab…
Quick Check
Explain it back
Design a target suite and a retention suite for one adaptation. Define one release rule before seeing the results, then explain why an average alone is insufficient.
Key Takeaways
- Use paired before/after evaluation.
- Keep target and retention suites separate.
- Record all model/prompt/decoding/evaluator identities.
- Inspect failure slices and regressed case IDs.
- Predeclare acceptance thresholds when practical.
Next Lesson
Next, study catastrophic forgetting and behavioral drift as specific forms of adaptation regression.
References
- Liang et al., Holistic Evaluation of Language Models.
- Mitchell et al., Model Cards for Model Reporting.
Completion is stored locally on this device.