본문으로 건너뛰기
L11.10

Human Approval Points

Goal

Pause an agent immediately before selected high-impact actions, ask a person to approve the exact proposed action, and define what happens when that approval is granted, denied, or expires.

Why must the refund pause here?​

A support agent receives a request to refund order 4172 for $120. It may read the order and check the refund policy. It may also prepare a proposal. But the moment before money moves is different.

A useful flow is:

inspect order
→ calculate allowed refund
→ propose: refund $120 to order 4172
→ ask an authorized person
→ execute only that approved action

The pause immediately before the external effect is a human approval point. It lets automation do the reversible preparation while keeping the consequential step under human control.

Compare that with a vague approval such as “approve refunds for this customer.” The vague version leaves too much room for later interpretation. A stronger approval names the action. It also records the important arguments: order 4172, amount $120, destination, and an expiry or run identity when relevant.

Approval therefore does not make an agent generally trusted. It authorizes one proposal under stated conditions. If the amount changes after approval, the old approval should not silently cover the new action.

Bind approval to arguments​

Suppose a human approves a refund of 20, but the model later changes the amount to 80. The old approval should fail the exact-match check.

Likewise, approval for order 4172 should not authorize a refund on order 9910.

Binding reduces a dangerous ambiguity: “the human approved something around this task.”

Approval is not permanent authority​

After an approved action executes, mark that approval as consumed when appropriate. A later side effect should need its own approval unless policy explicitly says otherwise.

Expiration can also matter. A stale approval may no longer reflect current state.

The application owns these rules. A model statement such as “approval still applies” is not enough.

Denial changes the trajectory​

If the human denies the action, the agent should not keep proposing the same action until someone eventually clicks yes.

Possible next states include:

  • report that the action was denied;
  • choose a permitted non-impactful alternative;
  • ask for different information if the denial reason points to a fix;
  • stop.

Repeatedly pressuring for approval is a failure mode, not persistence.

Choose approval points by risk​

Not every read-only lookup needs a human. Approval friction should be placed where consequences justify it.

Examples include actions that spend money, change permissions, delete data, send external messages, or commit to consequential decisions.

NIST AI RMF 1.0 treats AI risk management as an ongoing activity, not a one-time model property. The practical lesson here is simpler. Decide where human oversight belongs in the workflow, then make that pause technically enforceable.

OWASP guidance for LLM applications also emphasizes that prompt instructions alone are not a sufficient security boundary. The exact risk taxonomy evolves, so this course keeps the stable control principle: high-impact actions need enforceable authorization and oversight outside generated text.

Predict

A human approved refund_order for order 4172 with amount 20, but the new proposal uses amount 80. What should happen?

Run the local Lab​

Run:

python labs/notebooks/level-11/l11-10-human-approval.py

The script compares one proposed action, a refund of 20 for order 4172, against four approval records. An approval counts only if it says yes, names the same action, has exactly the same arguments, and has not expired.

  1. Run it unchanged. Read approved => True, then mismatch => False (approved amount 10), expired => False (expired at 17:00, checked at 18:00), and denied => False.
  2. Now let the agent change its plan after approval. In the proposal = ... line, change "amount": 20 to "amount": 25. Do not change any approval record. Rerun.
  3. Now approved => False too. The human approved refunding 20, not 25, so the old approval cannot be reused for the new action.

An approval is bound to one exact action. If the arguments change, the agent needs a new approval.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-11/l11-10-human-approval-exercise.py

Implement approval binding to the exact action, exact arguments, explicit approval flag, and expiry time. The mismatch and expired cases must remain blocked.

Run the starter after each change:

python3 labs/notebooks/level-11/l11-10-human-approval-exercise.py

A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.

Quick Check

1. When is a strong time to request human approval for a high-impact action?
2. Why bind an approval to exact action arguments?
3. What should an agent generally do after a human denies a proposed side effect?

0 of 3 questions answered.

Explain it back​

Choose one high-impact action for your agent. Write the smallest approval record that identifies the actor, action, exact arguments, and validity period. Then describe the state transition for approval, denial, and expiration.

Key Takeaways

  • Human approval should pause the run before selected high-impact actions.
  • Approval is strongest when bound to the exact proposal and arguments.
  • Approval can expire or be consumed; it is not automatically permanent authority.
  • Denial should change the trajectory rather than trigger repeated pressure.
  • Human oversight works best when it is implemented as an enforceable workflow boundary.

Next Lesson

Next, study common agent failure modes and identify the earliest boundary where each one becomes visible.

References

Lesson actions

Completion is stored locally on this device.

View progress