Skip to main content
L8.4

Supervised Fine-Tuning

Goal

Explain supervised fine-tuning as next-token training on desired examples, distinguish prompt tokens from target-response tokens, and interpret one small masked-loss example.

Think of a student practicing with an answer key. The page contains both the question and the answer, but the teacher may choose to grade only the answer-writing part. The question still matters because it tells the student what to answer; it simply does not need to count as a mistake to be corrected.

Supervised fine-tuning can use the same idea. An instruction, input, and target response are joined into one token sequence. The model reads the whole sequence for context, while a loss mask can mark only selected token positions as places where prediction error contributes to training. This makes the later 0/1 mask less mysterious: it is a bookkeeping rule for which next-token mistakes the optimizer should care about.

Supervised fine-tuning (SFT) is not a completely new training algorithm. For a decoder-only model, it still uses next-token prediction. The important change is which sequences you train on and which token positions contribute to the objective.

Turn an instruction example into one token sequence​

A simplified example:

Instruction: classify the issue
Input: battery drains quickly
Response: battery_drain

After formatting and tokenization, imagine:

[SYS, classify, USER, battery, drains, ASSIST, battery_drain, EOS]

The model predicts the next token at every position it is asked to learn from.

Response-only loss is a policy choice​

Some SFT pipelines calculate loss on all tokens. Others mask prompt/input positions and optimize only assistant-response positions. A mask might look like:

tokens: SYS classify USER battery drains ASSIST battery_drain EOS
loss?: 0 0 0 0 0 0 1 1

Here the input still influences the hidden state, but its token-prediction errors do not contribute to the average loss. This does not mean response-only loss is universally superior. It is a choice about what behavior the objective should emphasize.

Compute a tiny masked average​

Suppose response token losses are:

battery_drain: 0.6
EOS: 0.2

and all prompt positions are masked. Then:

masked mean loss = (0.6 + 0.2) / 2 = 0.4

Do not divide by all eight sequence positions. Only active loss positions belong in that average. A wrong mask can silently train the wrong objective while every tensor shape remains valid.

Fine-tuning still needs a held-out set​

Training loss tells you whether the model is fitting the SFT examples. It does not tell you whether the adapted model:

  • generalizes to new task cases;
  • retains useful base capabilities;
  • becomes more reliable on the product metric;
  • overfits formatting quirks.

Keep adaptation evaluation separate from optimizer updates.

Small updates can still move behavior​

A large pretrained model may need relatively little target data to shift a narrow behavior. That is useful, but it means bad examples can also matter. Inspect target examples before treating “more steps” as the answer to a weak evaluation result. If training loss falls but held-out behavior gets worse, consider data quality, target mismatch, overfitting, or update magnitude—not only architecture.

The loss tells you exactly which token errors matter​

Decoder-only supervised fine-tuning still uses next-token prediction. The main change is which sequences are presented and which token positions contribute to the objective. In a response-only setup, prompt tokens provide context for predicting the answer but their losses are masked out. The model can read the instruction without being rewarded for reproducing it.

That detail matters when interpreting training loss. A lower masked loss means the model became better at the selected target-token objective on the training examples. It does not by itself show better held-out task behavior, factual accuracy, or retention. Those claims require separate evaluation on data that was not optimized directly.

Predict

In response-only SFT, prompt tokens are masked from loss. Are they still part of the model input?

Complete the Browser Lab objective​

The Browser Lab gives token-level losses and a 0/1 response mask.

  1. Click Run once. The starter averages all eight losses, which is the wrong objective, so the checks fail.
  2. Complete the TODO in masked_mean so it averages only the losses whose mask value is 1.
  3. Click Run again. You should see response-only loss: 0.4, the average of 0.6 and 0.2, and every check should pass.
  4. Change one prompt-token loss: in token_losses, change the second value from 0.9 to 9.0.
  5. Before running, predict whether the response-only loss changes.
  6. Click Run. It is still 0.4. The model is not trained to predict the prompt, so prompt losses do not count.
  7. Change a response loss instead: change the seventh value from 0.6 to 1.0 (keep the second value at 9.0 or restore it). Run again. The loss becomes 0.6.

Loading lab…

Quick Check

1. What remains the core objective in decoder-only SFT?
2. What does a response-only loss mask change?
3. Training loss falls but held-out task performance worsens. What is a defensible conclusion?

0 of 3 questions answered.

Explain it back​

Draw a token sequence with instruction, input, and response. Mark which positions you would include in response-only loss and explain why masked prompt tokens still matter to the forward pass.

Key Takeaways

  • SFT is next-token training on curated desired behavior.
  • Formatting determines the sequence the model actually sees.
  • Loss masking controls which positions contribute to optimization.
  • Prompt tokens can provide context even when masked from loss.
  • Held-out evaluation is still required.

Next Lesson

Next, reduce the number of trainable parameters by adapting a small set of weights while freezing most of the pretrained model.

References

Lesson actions

Completion is stored locally on this device.

View progress