본문으로 건너뛰기
L8.1

When to Fine-Tune

Goal

Decide whether a model failure should be addressed with prompting, retrieval, application logic, or fine-tuning, and state what evidence would justify changing model weights.

Level 7 showed that many “model problems” are really workflow problems. If the model lacks a current policy document, adding that document through retrieval is usually more direct than fine-tuning. If an output must be valid JSON, schema validation is stronger than hoping fine-tuning makes malformed output impossible. Fine-tuning becomes useful when the desired change is behavioral and repeated enough that carrying all of it in the prompt is inefficient or unreliable.

Start from the failure, not from the technique​

Suppose a support assistant has four problems:

  1. It does not know yesterday's price change.
  2. It sometimes returns malformed JSON.
  3. It consistently uses the wrong domain-specific response style despite a clear prompt.
  4. It needs to apply a deterministic refund limit.

These suggest different interventions.

  • Current price → retrieval or tool.
  • Malformed JSON → structured-output/schema enforcement.
  • Stable domain behavior/style → fine-tuning may be worth testing.
  • Refund limit → application logic.

The point is not that fine-tuning can never affect the other problems. The point is to use the simplest controllable lever that matches the failure.

Fine-tuning changes a learned prior​

A pretrained model enters a request with learned tendencies from its original training. Fine-tuning supplies new examples and updates some parameters so desired responses become more likely under similar conditions. That can help with:

  • recurring format/style behavior;
  • task conventions that appear repeatedly;
  • domain-specific response patterns;
  • compact adaptation when prompt examples would consume too much context.

But changing weights also creates new risks:

  • behavior outside the target task may move;
  • bad labels can become learned behavior;
  • evaluation leakage can create false confidence;
  • later debugging is harder if data and checkpoint identity are not recorded.

Use a decision table​

For one observed failure, ask:

QuestionIf yes, consider
Is the missing information current/private and retrievable?Retrieval/tooling
Is the requirement deterministic and enforceable in code?Application logic
Is a clear prompt sufficient under fixed evaluation?Prompting
Is the same behavioral gap repeated across many examples?Fine-tuning
Can you create trustworthy target examples and held-out tests?Fine-tuning becomes more defensible

Fine-tuning without reliable training examples or evaluation cases is not a controlled experiment.

Define success before adaptation​

Suppose the target gap is:

The base model answers internal troubleshooting questions correctly but writes long free-form explanations; the product requires a three-field diagnostic record.

Before fine-tuning, create held-out cases and checks:

  • required fields present;
  • unsupported fields remain null;
  • diagnostic label accuracy;
  • no regression on general support questions;
  • token/latency cost if relevant.

Then the adaptation can be judged against a base-model baseline.

Predict

A model does not know a policy published this morning. Which intervention should usually be tested first?

Run the Browser Lab​

The Browser Lab scores several failure scenarios against four possible interventions.

  1. Click Run. Read the three decisions: today_policy -> retrieval, refund_limit -> application_logic, and domain_style -> fine_tuning.
  2. For each case, find the field that decided it. For example, refund_limit has "hard_rule": True, so it never reaches the fine-tuning question.
  3. Turn domain_style into a current-fact problem: change its "fresh_external_knowledge": False to True.
  4. Before running, predict the new decision.
  5. Click Run. domain_style now goes to retrieval. A fact that changes should be supplied, not trained in.
  6. Press Reset. Add a case where fine-tuning looks plausible but there is no held-out evaluation yet. Insert this line inside cases:
{"name": "no_eval_yet", "hard_rule": False, "fresh_external_knowledge": False, "prompt_works": False, "repeated_behavior_gap": True, "training_examples": True, "heldout_eval": False},
  1. Click Run. The new case returns insufficient_evidence. Without a way to measure success, a decision to fine-tune should stay provisional. The Lab should report Result: experiment ran because the case-list baseline changed while the intervention rules still satisfy their invariant checks.

Loading lab…

Quick Check

1. Which problem is most directly solved by retrieval rather than fine-tuning?
2. What should exist before a fine-tuning experiment?
3. Why are deterministic application rules preferable for a hard constraint?

0 of 3 questions answered.

Explain it back​

Choose one model failure and justify one of four interventions: prompting, retrieval/tooling, application logic, or fine-tuning. State what evidence would make you reconsider the choice.

Key Takeaways

  • Start from the observed failure, not from a preferred technique.
  • Fine-tuning is most defensible for repeated behavioral gaps with trustworthy examples.
  • Fresh knowledge usually belongs in retrieval/tools.
  • Hard deterministic constraints usually belong in application logic.
  • Define fixed before/after evaluation before changing weights.

Next Lesson

Next, turn the desired behavior into instruction-tuning examples that make the input/output contract visible.

References

Lesson actions

Completion is stored locally on this device.

View progress