Preference Data
Goal
Build preference examples by comparing two responses to the same request, explain what a chosen/rejected pair does and does not tell you, and spot ambiguous or biased comparisons before training.
Start with two answers to the same request
A customer asks:
I lost my password. How can I get back into my account?
Two candidate answers appear:
A: gives the normal reset steps and tells the user what to do if the reset email never arrives.
B: invents an “administrator bypass code” that the product does not have.
A reviewer chooses A. A training record can store the shared request, A as the chosen response, and B as the rejected response. That is a preference pair.
Supervised fine-tuning usually gives one desired target response. Preference data gives a relative signal instead: for the same context, one candidate is preferred over another.
The pair says something specific: under this context and the annotation rule, A was preferred to B. It does not prove that A is the best possible answer, that every reviewer would agree, or that A is correct in every account-recovery situation. The reviewer may also be reacting to tone, length, formatting, or another hidden preference.
That is why preference data needs the same care as other training data. Keep the compared context controlled, make the judging rule explicit, and preserve uncertainty when the choice is genuinely ambiguous.
Keep the context identical inside a pair
If chosen and rejected responses use different prompts, you no longer know whether the preference came from response quality or different input conditions. A clean pair holds context fixed:
same prompt
same source evidence
same tool state
different candidate responses
This is another controlled-comparison problem.
Ambiguous pairs should be surfaced
Suppose:
A: "The battery lasts about ten hours."
B: "Battery life is approximately 10 hours."
If the task only cares about factual content, the preference may be arbitrary. Forcing a label can teach noise. Useful policies include:
- allow ties/uncertain;
- send ambiguous cases for review;
- record annotator disagreement;
- remove pairs with no meaningful distinction.
Preference data can encode hidden biases
If annotators systematically prefer longer answers, formal tone, one dialect, or one cultural style even when the task does not require it, the model may learn those tendencies. Inspect preference statistics:
- length difference;
- response position order;
- label balance;
- annotator identity if available;
- domain/source slices.
A preference dataset is not automatically neutral because it was created by humans.
Prevent train/eval leakage
The same prompt or near-duplicate pair should not appear across training and held-out evaluation. If multiple candidate pairs come from one source interaction, group them before splitting. The same leakage rules from instruction data still apply.
Randomize candidate order when humans compare responses
Human preference collection can develop a position bias. If the better response is usually shown on the left during data creation, an annotation interface or annotator habit can accidentally correlate “left” with “chosen.” Randomizing candidate display order and recording the original candidate identity helps detect this shortcut. After collection, compare:
- chosen-left versus chosen-right frequency;
- annotator-specific preference rates;
- disagreement rates;
- preference rate by response length and format.
A strong imbalance is not proof of bias, but it is evidence worth investigating.
Write the annotation rule before labeling at scale
Two annotators can disagree because one optimizes factual support while another optimizes helpfulness. A short rubric can specify priority:
1. must be supported by the supplied evidence
2. must follow the requested format
3. among equally correct outputs, prefer the clearer response
Now disagreements can be interpreted against a shared rule instead of being treated as unexplained noise. Keep rubric version alongside the preference dataset because changing the rubric changes what the labels mean.
Predict
Run the Browser Lab
The Browser Lab audits a small preference dataset.
- Click Run. Read
{'uncertain_ids': ['p2'], 'avg_chosen_words': 5.5, 'avg_rejected_words': 4.5}. - Look at pair
p2: “Battery duration is 10 hours.” versus “It lasts ten hours.” Both are correct, so forcing one to be “better” would teach an arbitrary style preference. It is markeduncertaininstead. - Add a pair where the longer answer is chosen only because it is longer. Add this line to
pairs:
{"id": "p3", "prompt": "Give the refund window", "chosen": "The refund window, as described in our current policy document, is thirty days.", "rejected": "Refunds: 30 days.", "uncertain": False},
- Before running, predict how the average lengths will change.
- Click Run.
avg_chosen_wordsjumps to8.0whileavg_rejected_wordsfalls to4.0. If most chosen answers are simply longer, a model can learn “longer wins” instead of “more correct wins.” That hidden shortcut is why you audit length. - Press Reset afterward.
Loading lab…
Quick Check
Explain it back
Create one preference pair for a task you know. State the preference criterion, keep the context fixed, and describe one condition that would make the pair too ambiguous to use.
Key Takeaways
- Preference data expresses relative ranking, not absolute truth.
- Keep context fixed inside a pair.
- Document the criterion behind the preference.
- Surface ambiguity and annotator disagreement.
- Inspect superficial shortcuts such as length or ordering.
- Group related records before train/eval splitting.
Next Lesson
Next, translate pairwise preference evidence into an optimization signal that increases relative preference for chosen behavior.
References
- Rafailov et al., Direct Preference Optimization.
Completion is stored locally on this device.