Skip to main content
L15.1

Evaluation as a System

Goal

Turn a vague question such as “Is this AI good?” into a repeatable evaluation with cases, measures, slices, and a decision rule.

Start with the decision​

Suppose a student gets 92 out of 100 questions right on a practice set. That number is useful, but it does not tell you whether the eight mistakes were harmless spelling errors or eight failures on the one safety rule that mattered most. A single average can hide the shape of the mistakes.

An AI evaluation works the same way. Before choosing a metric, decide what decision the evidence must support. Then define cases, expected behavior, separate measures, important groups of cases, and hard rules that cannot be averaged away. The technical word slice simply means a group you inspect separately because its behavior may matter on its own.

Now make the decision explicit. A release team might need to decide whether a candidate version can replace the current version. That decision creates concrete evidence needs: the candidate should preserve task success, avoid critical policy failures, stay within latency limits, and not regress important user slices. The evaluation exists to answer those questions.

Build cases, measures, and slices​

A useful evaluation therefore has several linked parts. Cases are the inputs and context you test. Expected behavior says what a correct or acceptable response should do. Measures turn observed behavior into evidence. Slices group cases that should be inspected separately, such as task type, language, customer tier, or safety category. A decision rule says how the evidence affects release, rollback, or further review.

{"case_id":"refund-policy-07","slice":"high-risk","task_pass":true,"policy_pass":false,"release_blocker":true}

The same case can support more than one measure. Imagine a question asking an assistant to summarize a return policy. You might check whether required facts are present, whether invented facts are absent, whether cited evidence matches the source, and whether the response stays within a latency budget. One answer produces several pieces of evidence because the product has several responsibilities.

Keep averages from hiding critical failures​

Averages are useful but dangerous when they hide structure. If nine ordinary cases pass and one critical authorization case fails, a 90% success rate should not automatically mean “ship.” The critical case may represent a boundary where one failure is unacceptable. The decision rule must preserve that distinction instead of blending every outcome into one score.

HELM is useful here because it argues for broader, more transparent evaluation across scenarios and metrics rather than relying on one headline number. You do not need to copy a benchmark suite into every product. The transferable idea is to state what is being measured, under what scenario, and with which metrics so readers can understand what a result does and does not support.

Risk evidence also changes across the system lifecycle. NIST AI RMF describes risk management as continuous and organizes activities around Govern, Map, Measure, and Manage. In evaluation terms, that means evidence is not collected once and forgotten. New users, tools, data, model versions, or deployment contexts can change the risk map and therefore change what should be tested.

Preserve raw evidence and debug the first failure​

A practical release record should keep raw case outcomes, not only the final score. If a field says release_passed: true, the validator should be able to recompute that conclusion from individual observations and the declared rule. Otherwise the same data is both the student and the grader.

This distinction also helps debugging. When a release fails, ask which boundary failed first. A low aggregate score may come from one broken task family. A safety gate may fail even while overall quality improves. A latency regression may affect only long-context cases. The evaluation should preserve enough detail to identify the responsible slice instead of forcing you to rerun everything blindly.

Evaluation design is therefore a system-design task. You are choosing what evidence to preserve, how to summarize it, and what decisions it may justify. A metric without a decision can become dashboard decoration. A decision without traceable evidence becomes an opinion that is hard to review.

Predict

A candidate improves average task success from 90% to 94%, but one zero-tolerance authorization test fails. What should a release rule designed with independent critical gates do?

Run the local Lab​

Run:

python3 labs/notebooks/level-15/l15-01-evaluation-system.py

The Lab recomputes a release decision from per-case evidence.

  1. Run it unchanged. Record case_success: 1.0, critical_failures: 0, and release_passed: true.
  2. Find the auth-1 case with "passed": True and "critical": True. Before editing, predict what should happen if only that critical case fails. The other nine cases stay successful, so the aggregate rate should remain high.
  3. Change only auth-1 to "passed": False, then rerun.
  4. Confirm the aggregate success rate is still 0.9, but critical_failures becomes 1 and release_passed becomes false. Explain why a zero-tolerance gate should not disappear inside an average.

Loading lab…

Quick Check

1. What is the strongest starting point for an evaluation design?
2. Why preserve evaluation slices instead of reporting only one average?
3. Why should a validator recompute a release flag from raw outcomes?

0 of 3 questions answered.

Explain it back​

Choose one AI feature you know. State the release decision, three kinds of cases, one slice that could hide a regression, and one failure that should remain an independent blocker.

Key Takeaways

  • Begin evaluation design with a decision and the evidence that decision needs.
  • Keep raw case outcomes, slices, and critical failures visible.
  • Use multiple measures when a product has multiple responsibilities.
  • Do not let an average erase a zero-tolerance boundary.
  • Recompute important release conclusions from primary evidence.

Next Lesson

Next, L15.2 — Golden Sets and Regression Suites turns important cases into repeatable protection against regressions.

References

Lesson actions

Completion is stored locally on this device.

View progress