Model and System Safety Basics
Goal
Separate model behavior from system risk and organize safety work around context, evidence, controls, and ongoing review.
Risk belongs to the whole system
The same kitchen knife can be low risk when cutting fruit with supervision and much higher risk in a different setting. Risk depends not only on the object but also on who uses it, what it can affect, and what happens when something goes wrong.
AI safety needs the same system view. A language model is one component. The surrounding tools, permissions, users, data, review steps, and consequences determine what a model mistake can turn into. Start by asking four plain questions: What can happen? Who can be affected? What controls reduce the chance or consequence? What risk remains afterward? The formal safety frameworks in this Lesson give those questions a repeatable structure.
In AI, the same model can therefore carry very different risk in different applications. Drafting internal meeting notes for human review is not the same risk as sending money after reading user messages, even if the model's error rate is unchanged.
Draw the system boundary
That is why safety analysis needs a system boundary. List the model, prompts, retrieval sources, tools, people, policies, data stores, deployment environment, and external services that participate in the task. A failure can enter through any of those parts or through the way they interact.
user → application → model → tool → external system
policy output side effect
A safety review asks which failures can cross each arrow and which control stops them.
Risk is not the same as “bad output.” NIST's Generative AI Profile describes risk in terms of both likelihood and consequence and notes that risks can differ by lifecycle stage, system scope, source, and time scale. For a product team, this means the same failure can deserve different treatment depending on who is affected and what the system is allowed to do.
Use a repeatable risk framework
NIST AI RMF provides four high-level functions: Govern, Map, Measure, and Manage. Govern covers organizational policies, roles, and accountability. Map develops context: intended use, affected people, assumptions, and possible impacts. Measure collects evidence about identified risks. Manage prioritizes and acts on that evidence. The functions support one another; they are not a one-time sequence that ends after release.
Consider an assistant that can draft and send an email. Mapping the system identifies the side effect: sending a message can create commitments or expose information. Measuring might include tests for wrong recipients, prompt injection, sensitive-data leakage, and unauthorized send attempts. Managing could require a confirmation step and recipient allowlist. Governing defines who owns the policy and how incidents are reviewed.
Match controls to the failure boundary
Controls should match the boundary. If the risk is an unauthorized tool call, a stronger system prompt is not enough because the model is still the component proposing the action. The application should enforce permissions outside the model. If the risk is unsupported factual claims, retrieval evidence and citation checks may be appropriate. If the risk is harmful content exposure, access, filtering, review, and escalation controls may all matter.
Safety work also needs residual risk, the risk that remains after controls. A guardrail can reduce a failure mode without making it impossible. Report what the control was tested against, what evidence supports it, and what still remains uncertain. “We added a safety filter” is not a complete risk statement.
Keep residual risk and uncertainty visible
Uncertainty matters especially for generative systems. NIST AI 600-1 notes that some risks are difficult to estimate because evidence and measurement science are still developing. When evidence is weak, say so. A cautious decision can include monitoring, restricted rollout, human approval, or narrower permissions while more evidence is gathered.
Incidents should feed the evaluation system. A harmful output, unauthorized action, or privacy event should become a reviewed case, a root-cause question, and often a regression test. Otherwise the team may fix the immediate symptom without preserving the lesson.
Safety therefore connects directly to previous Lessons. Golden sets preserve known risk cases. Human evaluation handles judgment-heavy criteria. Production observability from Level 14 preserves evidence about which version acted. The safety layer turns those pieces into an explicit risk-management process.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-15/l15-04-safety-system.py
The Lab maps a small AI workflow into assets, controls, and residual-risk review.
- Run it unchanged with
"permission":"read". Record the controls andresidual_risk_review_required. - Before editing, predict which additional controls should appear if the same workflow gains write authority.
- Change only
"permission":"read"to"permission":"write", then rerun. - Confirm
authorizationandside_effect_confirmationare added and residual-risk review becomes required. The asset did not change; only the authority did.
Loading lab…
Quick Check
Explain it back
Choose one AI workflow and draw its system boundary in words. Name one hazard, one control, the evidence you would collect, and one residual risk that would still need review.
Key Takeaways
- Safety is a property of the whole sociotechnical system, not only the model.
- Use context and consequences to identify relevant risks.
- Govern, Map, Measure, and Manage organize continuous risk work.
- Match controls to the boundary where the failure can actually occur.
- Record residual risk and uncertainty instead of claiming perfect safety.
Next Lesson
Next, L15.5 — Threat Modeling for AI Systems turns broad safety concerns into concrete adversaries, assets, paths, and mitigations.
References
Completion is stored locally on this device.