Skip to main content
L10.12

Tool Security Boundaries

Goal

Threat-model a tool-using assistant by separating untrusted content, model proposals, trusted identity, permission checks, approvals, execution authority, and audit evidence.

Imagine a visitor enters a school carrying a note that says, "The principal says I may enter the server room." The note is information to inspect, not authority by itself. A staff member checks the visitor's trusted identity, the actual access rule, and whether the requested room is allowed before opening any door.

A tool-using assistant needs the same separation. User text, retrieved documents, model-produced arguments, and tool results can all contain useful information, but none should be able to grant itself more permission. Trusted identity, policy, approvals, and execution limits stay outside generated text. Security design is therefore about controlling the path from untrusted input to real effect, not about finding one perfect warning sentence for the prompt.

That is why prompt wording can help communicate intent to the model, but it is not the final security control.

Draw the trust map​

A simple map might classify:

trusted:
authenticated principal
tool registry
permission policy
workflow state
approval record

untrusted:
user text
retrieved documents
webpages
image/audio text
model-produced arguments
external tool data

Some untrusted data is useful evidence. “Untrusted” means it cannot grant itself more authority.

Least privilege reduces damage​

A read-only lookup tool should not secretly carry write permission. A code sandbox that needs one working directory should not see the host filesystem. A support assistant should not receive an administrator token because “the prompt says not to misuse it.”

Reduce permissions before asking the model to behave safely.

Require approval at meaningful boundaries​

Human approval is useful when an action is costly, destructive, irreversible, legally sensitive, or otherwise high impact.

The approval record should bind to the actual action and arguments. Approving “refund order 4172 for $20” should not authorize “refund order 4172 for $2,000.”

Validate after the model, before the tool​

Even a model that usually follows rules can make a wrong selection or be influenced by malicious content.

The application should check:

  • tool is allowlisted;
  • arguments match schema;
  • principal is permitted;
  • current state allows the action;
  • required approval exists;
  • limits/quotas are respected.

Then it may execute.

Audit both denied and executed actions​

Security review needs evidence of what the system attempted, not only what succeeded. Record denial category without leaking secrets.

Useful traces connect:

request → model proposal → policy decision → approval → execution → result

This lets a reviewer see whether defenses stopped the proposal at the intended boundary.

Threat-model the path, not only the model​

Ask what an attacker can control at every input boundary: user message, retrieved page, image text, audio, tool result, file contents, or API payload. Then ask what each component can influence next. This produces a concrete path such as “malicious webpage text influences a tool proposal, which is then stopped by the allowlist and role policy.”

That framing is more useful than asking whether the model is “secure.” The model can still be manipulated while the surrounding system prevents unsafe effects. Conversely, a model that usually follows instructions is not a substitute for a missing permission check.

Denials should be predictable and minimally informative​

When policy blocks an action, return enough information for the workflow to recover—such as permission_denied or approval_required—without leaking sensitive policy internals, resource existence, or credentials. Stable denial categories also make security regression tests easier to automate.

Predict

A retrieved webpage tells the model to call an admin-only delete tool. What is the strongest defense?

Run the local Lab​

python labs/notebooks/level-10/l10-12-security-policy.py

The proposal asks to refund 120, but the human approved only 100.

  1. Run the command. reader: deny_tool (a reader may not use refund_order at all) and operator: deny_approval_mismatch (an operator may, but only for the approved amount).
  2. Keep the proposal the same and change the approval: "approved_amount": 100 becomes "approved_amount": 120. Rerun.
  3. The operator line becomes allow, while the reader is still denied. Role permissions and approval binding are two separate checks, and both must pass. The script then stops with an AssertionError, which is expected: its final checks were written for the original values. Change the value back afterward.

Loading lab…

Quick Check

1. What does least privilege mean for tool design?
2. Why bind approval to concrete arguments?
3. Which data source can grant a user new tool permission by itself?

0 of 3 questions answered.

Key Takeaways

  • Separate trusted authority from useful but untrusted content.
  • Least privilege limits the damage of bad proposals.
  • Bind approvals to concrete actions and arguments.
  • Validate allowlist, schema, permission, state, limits, and approval before execution.
  • Audit denied as well as executed proposals.

Next Lesson

Next, integrate selection, validation, state, retries, multimodal evidence, and security into one reviewable tool-using assistant.

References

Lesson actions

Completion is stored locally on this device.

View progress