Supervised Fine-Tuning
Goal
Explain supervised fine-tuning as next-token training on desired examples, distinguish prompt tokens from target-response tokens, and interpret one small masked-loss example.
Think of a student practicing with an answer key. The page contains both the question and the answer, but the teacher may choose to grade only the answer-writing part. The question still matters because it tells the student what to answer; it simply does not need to count as a mistake to be corrected.
Supervised fine-tuning can use the same idea. An instruction, input, and target response are joined into one token sequence. The model reads the whole sequence for context, while a loss mask can mark only selected token positions as places where prediction error contributes to training. This makes the later 0/1 mask less mysterious: it is a bookkeeping rule for which next-token mistakes the optimizer should care about.
Supervised fine-tuning (SFT) is not a completely new training algorithm. For a decoder-only model, it still uses next-token prediction. The important change is which sequences you train on and which token positions contribute to the objective.
Turn an instruction example into one token sequence
A simplified example:
Instruction: classify the issue
Input: battery drains quickly
Response: battery_drain
After formatting and tokenization, imagine:
[SYS, classify, USER, battery, drains, ASSIST, battery_drain, EOS]
The model predicts the next token at every position it is asked to learn from.
Response-only loss is a policy choice
Some SFT pipelines calculate loss on all tokens. Others mask prompt/input positions and optimize only assistant-response positions. A mask might look like:
tokens: SYS classify USER battery drains ASSIST battery_drain EOS
loss?: 0 0 0 0 0 0 1 1
Here the input still influences the hidden state, but its token-prediction errors do not contribute to the average loss. This does not mean response-only loss is universally superior. It is a choice about what behavior the objective should emphasize.
Compute a tiny masked average
Suppose response token losses are:
battery_drain: 0.6
EOS: 0.2
and all prompt positions are masked. Then:
masked mean loss = (0.6 + 0.2) / 2 = 0.4
Do not divide by all eight sequence positions. Only active loss positions belong in that average. A wrong mask can silently train the wrong objective while every tensor shape remains valid.
Fine-tuning still needs a held-out set
Training loss tells you whether the model is fitting the SFT examples. It does not tell you whether the adapted model:
- generalizes to new task cases;
- retains useful base capabilities;
- becomes more reliable on the product metric;
- overfits formatting quirks.
Keep adaptation evaluation separate from optimizer updates.
Small updates can still move behavior
A large pretrained model may need relatively little target data to shift a narrow behavior. That is useful, but it means bad examples can also matter. Inspect target examples before treating “more steps” as the answer to a weak evaluation result. If training loss falls but held-out behavior gets worse, consider data quality, target mismatch, overfitting, or update magnitude—not only architecture.
The loss tells you exactly which token errors matter
Decoder-only supervised fine-tuning still uses next-token prediction. The main change is which sequences are presented and which token positions contribute to the objective. In a response-only setup, prompt tokens provide context for predicting the answer but their losses are masked out. The model can read the instruction without being rewarded for reproducing it.
That detail matters when interpreting training loss. A lower masked loss means the model became better at the selected target-token objective on the training examples. It does not by itself show better held-out task behavior, factual accuracy, or retention. Those claims require separate evaluation on data that was not optimized directly.
Predict
Complete the Browser Lab objective
The Browser Lab gives token-level losses and a 0/1 response mask.
- Click Run once. The starter averages all eight losses, which is the wrong objective, so the checks fail.
- Complete the TODO in
masked_meanso it averages only the losses whose mask value is1. - Click Run again. You should see
response-only loss: 0.4, the average of0.6and0.2, and every check should pass. - Change one prompt-token loss: in
token_losses, change the second value from0.9to9.0. - Before running, predict whether the response-only loss changes.
- Click Run. It is still
0.4. The model is not trained to predict the prompt, so prompt losses do not count. - Change a response loss instead: change the seventh value from
0.6to1.0(keep the second value at9.0or restore it). Run again. The loss becomes0.6.
Loading lab…
Quick Check
Explain it back
Draw a token sequence with instruction, input, and response. Mark which positions you would include in response-only loss and explain why masked prompt tokens still matter to the forward pass.
Key Takeaways
- SFT is next-token training on curated desired behavior.
- Formatting determines the sequence the model actually sees.
- Loss masking controls which positions contribute to optimization.
- Prompt tokens can provide context even when masked from loss.
- Held-out evaluation is still required.
Next Lesson
Next, reduce the number of trainable parameters by adapting a small set of weights while freezing most of the pretrained model.
References
- Wei et al., Finetuned Language Models Are Zero-Shot Learners.
Completion is stored locally on this device.