본문으로 건너뛰기

Level 15 — Evaluation, Safety, Security, and Frontier Practice

From a good demo to a defensible release​

Imagine you are deciding whether a school robot is ready to help students. "It worked in my demo" is not enough. You would want to know which tasks were tested, what kinds of mistakes happened, and whether the robot can reach things it should not. You would also want to know what information it stores and what happens if a new version behaves worse.

Level 15 organizes those questions into professional engineering ideas. Evaluation asks what evidence supports a decision. Safety asks what harmful outcomes matter in the real use context. Security asks what an attacker or untrusted input could influence. Governance asks who owns the rules and data lifecycle. Frontier practice asks how to treat fast-changing claims without turning a benchmark or demo into a universal fact. The formal terms come later; keep these ordinary questions as the map.

A production AI system can be fast, available, and reproducible and still fail at the thing that matters: helping people without creating unacceptable risk. A model may answer common questions well but fail on rare inputs. A tool-using assistant may produce useful text while taking an action the user never authorized. A release may look better on one benchmark while becoming worse for a particular workflow, language, or safety boundary.

Level 15 treats evaluation and defense as part of the product, not as a final score added after development. You will build an evaluation system that connects product goals to test cases, metrics, human review, security tests, governance decisions, and release rules. The main question changes from “Is the model good?” to “What evidence would justify shipping this system for this use, and what evidence would make us stop?”

Learning path​

You begin with evaluation as a system. A useful evaluation has a purpose, a defined population of cases, repeatable procedures, and decisions attached to the results. One average number rarely tells the whole story. You will learn to preserve slices, error categories, and zero-tolerance failures so that an improvement in one area cannot hide a serious regression in another.

Next you will build golden sets and regression suites. A golden set is not a pile of favorite examples. It is a reviewed collection of cases with expected behavior, provenance, and enough diversity to reveal important failures. Regression testing keeps known failures from quietly returning when prompts, models, tools, retrieval, policies, or serving settings change.

Some important qualities cannot be reduced to exact-match checks. Human evaluation helps when correctness depends on judgment, usefulness, clarity, or policy interpretation. You will design rating instructions, use multiple reviewers when practical, separate disagreements from bugs, and avoid pretending that one person's preference is objective truth.

Safety work then expands the unit of analysis from model output to the whole AI system. NIST's AI Risk Management Framework organizes risk work around Govern, Map, Measure, and Manage. The Generative AI Profile applies that style of risk management to generative systems and emphasizes that risks can appear at the model, application, use-case, and ecosystem levels. In this course, the framework is used as a way to organize evidence and decisions, not as a compliance checklist.

Security lessons make the threat boundary concrete. You will build a small threat model, test prompt injection and data-exfiltration paths, and apply least-privilege rules to tools and agents. The core principle is that model text is not authority. Retrieved documents, web pages, tool outputs, and messages from other agents can contain useful data and malicious instructions at the same time. Trusted identity, permissions, policy, and secrets must remain under application control.

Privacy and governance add another lifecycle question. What data should the system collect, and why may it keep that data? Who may access it? When should it be deleted? You will practice data minimization, purpose limits, retention rules, incident evidence, and reviewable decisions rather than treating “we do not train on it” as a complete privacy policy.

The final third moves into adversarial testing and frontier practice. Red teaming uses structured abuse cases to find weaknesses before or during deployment. The goal is not to “beat” a model for sport. The goal is to find realistic failure paths, measure their severity, and turn those findings into fixes and regression tests.

Frontier lessons are dated snapshots​

Three Lessons are explicitly marked Frontier. They are snapshots, not timeless promises. L15.10 examines capability claims and limits without assuming that benchmark gains transfer to every product. L15.11 revisits agent interoperability using the curriculum-pinned MCP 2026-07-28 release and the published A2A v1.0.0 specification, while recording the later A2A v1.0.1 patch release as compatibility evidence. L15.12 practices reading and reproducing research while separating a paper's reported result from what you can independently reproduce in your environment.

Level Project​

The Level Project, Ship, Evaluate, and Defend an AI Product, brings these ideas together. You will implement a deterministic release controller that recomputes regression results, human-rating summaries, threat and authorization violations, retention checks, and release decisions from raw evidence. The required path is offline so you can debug evaluation logic without depending on a paid model, network service, GPU, or external account.

What mastery looks like​

By the end of Level 15, you should be able to explain what evidence supports a release, what risks remain, which failures are blockers, how you would detect regressions, and which claims are stable versus version-sensitive. That is the final builder skill in this course. Making the system work is only part of the job. You also need to test it, constrain it, review it, and defend the decision to ship it.

Key Takeaways

  • Evaluation connects product goals, test cases, evidence, and release decisions.
  • Regression suites preserve important cases and failure slices across changes.
  • Human judgment needs clear instructions, disagreement handling, and observable anchors.
  • Safety, security, privacy, and governance are system properties, not model-only properties.
  • Frontier claims must be date/version pinned and independently checked before they become product assumptions.

Next Lesson

Start with L15.1 — Evaluation as a System.

Lesson actions

Completion is stored locally on this device.

View progress