본문으로 건너뛰기
L15.3

Human Evaluation

Goal

Design human review that turns subjective judgment into useful evidence with clear criteria, examples, and disagreement handling.

Use people where exact matching fails​

Some AI outputs cannot be judged with exact matching. Two summaries can use different words and both be correct. A tutoring answer may be factually right but too confusing for a beginner. A refusal can follow policy but sound needlessly hostile. Human evaluation is useful when the property you care about needs informed judgment.

The difficult part is not finding people to click a score. It is defining the question well enough that different reviewers are judging the same property. “Rate quality from 1 to 5” is vague. One reviewer may reward detail, another may reward brevity, and a third may focus only on correctness.

Separate rubric dimensions​

A stronger rubric separates dimensions. For a factual support answer, you might rate correctness, completeness, groundedness, and clarity separately. Each dimension needs anchors. A correctness score of 4 might mean “all material claims are correct,” while a 2 might mean “at least one material claim is wrong but the answer still contains useful information.” Anchors turn a number into an observable decision.

correctness: 0–4
completeness: 0–4
clarity: 0–4
policy compliance: pass / fail

Examples help reviewers calibrate. Show one response that clearly meets the criterion and one nearby response that does not. The examples should explain the reason, not merely label the answer. Reviewers then have a shared reference for difficult cases.

Reduce avoidable reviewer bias​

Blind comparison can reduce avoidable bias. If reviewers know that “Candidate B” is the expensive new model, that knowledge can influence judgment. Randomized labels and balanced ordering help when comparing systems. They do not remove every source of bias, but they make the procedure easier to inspect.

Multiple reviewers are valuable when disagreement is expected. Disagreement is not automatically noise. It can reveal an ambiguous rubric, a genuinely subjective preference, or a case where product policy is unclear. Record the individual ratings before producing an aggregate so you can distinguish consensus from a split decision.

Treat disagreement as evidence​

Suppose two reviewers score helpfulness as 4 and 4 while a third gives 1. The mean is 3, but the pattern matters. Read the comments. The third reviewer may have noticed a dangerous factual error that the broad “helpfulness” label hid. This is another reason to keep critical checks separate from broad preference ratings.

Human evaluation also needs sampling discipline. Reviewing only the first 20 outputs is easy, but ordering effects may make that sample unrepresentative. Define how cases are selected, how many come from each important slice, and whether reviewers see repeated or control items.

Sample carefully and protect reviewers​

The NIST Generative AI Profile emphasizes pre-deployment testing and risk measurement as part of broader risk management. Human review can contribute evidence, but it should not be treated as magical ground truth. Reviewer expertise, instructions, fatigue, cultural context, and task framing all affect the result.

Protect reviewer privacy and wellbeing as well. Some red-team or safety cases can contain disturbing material. Limit unnecessary exposure, give clear handling procedures, and avoid collecting reviewer information that the evaluation does not need.

A good human-evaluation report therefore contains the rubric, anchors, sampling method, reviewer count, disagreement information, and result distribution. “Humans preferred B 63% of the time” is more useful when the reader also knows what question was asked and how uncertain the judgment was.

Predict

Three reviewers strongly disagree on a small set of cases. What is the best first interpretation?

Run the local Lab​

Run:

python3 labs/notebooks/level-15/l15-03-human-evaluation.py

The Lab summarizes reviewer ratings and disagreement.

  1. Run it unchanged. Record mean_score and high_disagreement_cases; the two cases should initially have no high-disagreement flag.
  2. In case b, find reviewer r1 with "score":3. Before editing, predict whether changing only that score to 1 will cross the Lab's disagreement threshold while leaving the number of cases fixed.
  3. Change only that one score from 3 to 1, then rerun.
  4. Confirm high_disagreement_cases increases even though the output still has two cases. Compare the mean with the disagreement count and explain why the mean alone hides reviewer spread.

Loading lab…

Quick Check

1. Why are rubric anchors useful in human evaluation?
2. What can reviewer disagreement reveal?
3. Which evidence package is sufficient for auditing a human evaluation?

0 of 3 questions answered.

Explain it back​

Write one 1–5 rubric dimension for an AI response. Define what a 1, 3, and 5 mean using observable behavior, then name one reason two careful reviewers might still disagree.

Key Takeaways

  • Human evaluation needs specific dimensions and observable anchors.
  • Blind ordering can reduce some comparison bias.
  • Preserve disagreement instead of hiding it inside an average.
  • Sampling and reviewer procedures are part of the evaluation.
  • Human ratings are evidence, not infallible ground truth.

Next Lesson

Next, L15.4 — Model and System Safety Basics expands evaluation from output quality to risk across the whole AI system.

References

Lesson actions

Completion is stored locally on this device.

View progress