Tool and Agent Security
Goal
Apply least privilege, trusted identity, policy checks, approval, and audit evidence to tool-using and agentic workflows.
Keep identity, proposal, and policy separate
Picture a student handing a request slip to an office worker: "Please unlock room 12." The handwriting on the slip can describe the requested action, but it cannot declare, "I am the principal, so I am allowed." The office checks the student's real identity and the school rule before touching the lock.
Tool and agent security uses the same separation. The principal is the trusted identity on whose behalf an action would occur. The model's proposal describes a requested action. The policy is trusted application logic that decides whether that principal may perform that action on that resource. Keeping those three roles separate prevents generated text from granting itself authority.
A tool changes the stakes of a model output. Without a tool, a bad answer may mislead a user. With a tool, the same model can propose sending a message, changing a record, running code, or transferring data. The application must decide whether that proposal is allowed.
{"principal":"user-42","proposal":{"tool":"refund","order_id":"4172","amount":20},"policy_decision":"deny"}
Never let generated text grant authority
Do not let the proposal redefine the principal. A model should not be able to add "role":"admin" or switch tenant_id and thereby gain authority. Trusted identity comes from authenticated application state. Model-generated arguments may select task details only within the scope the principal already has.
Use least privilege and approval at real boundaries
Least privilege means exposing only the tools and permissions needed for the current task. A support assistant that only reads order status should not receive a generic network client or shell. A calendar helper that can create events may still need separate approval before inviting external people.
Approval is most useful at irreversible or high-impact boundaries. Do not ask for confirmation after the action already happened. Present the user with the action, destination, and important parameters before execution. The confirmation itself should bind to that exact proposal so a later model turn cannot silently change the amount or recipient.
Bound loops and delegated authority
Agent loops add persistence and repetition. A single unsafe proposal can be rejected, but an unconstrained agent may retry with modified arguments. Stopping conditions, retry budgets, and immutable principal information prevent “helpful” repair loops from becoming privilege escalation.
Delegation adds another identity question. If one agent asks another to perform work, the receiving side needs to know which authority is being delegated and which is not. “Another agent asked me” is not sufficient authorization. Preserve task identity, principal, allowed scope, and evidence of the delegation boundary.
Layer security principles with local architecture
NIST SP 800-207's zero-trust model is useful as a general security principle: do not grant broad trust merely because a component is inside a network boundary. For AI systems, apply that mindset to model outputs and agent messages too. Each protected resource request should be evaluated using trusted identity and policy.
OWASP and MITRE guidance can help enumerate tool and agent failure modes, but architecture still determines the control. A model prompt cannot reliably replace a permission check. A tool schema cannot replace resource authorization. A sandbox cannot replace a limit on which data may enter or leave it.
Record and test denied paths
Audit evidence should explain the decision. Record principal, requested action, resource, policy version, approval state, execution result, and correlation ID. Avoid logging secrets or complete sensitive payloads when a classification or digest is enough.
Security tests should cover denied paths, not only successful ones. Verify that an unauthorized principal is rejected, that changing model-generated role fields does not change authority, that approval is required where declared, and that retries cannot bypass the same policy.
Predict
Run the local Lab
Run:
python3 labs/notebooks/level-15/l15-07-tool-agent-security.py
The Lab evaluates a generated tool proposal against a trusted principal and application policy.
- Run it unchanged. Record
trusted_role: support, the proposal's generatedrole: admin, andauthorized: truefor the allowedread_orderaction. - Before editing, predict whether authorization should change if only the model-generated role string changes.
- Change only
proposal["role"]from"admin"to"owner", then rerun. - Confirm
generated_role_ignoredchanges buttrusted_roleandauthorizeddo not. Explain why authority comes from the trusted principal, not from model-generated text.
Loading lab…
Write the core logic yourself
Open:
labs/notebooks/level-15/l15-07-tool-agent-security-exercise.py
Implement authorization from the trusted principal role and application policy. The generated proposal includes a spoofed admin role on purpose; your code must ignore it.
Run:
python3 labs/notebooks/level-15/l15-07-tool-agent-security-exercise.py
The starter intentionally stops at TODO until you implement the missing logic. A correct solution reaches the final PASS: marker. Use the solved deterministic Lab as a comparison only after your own attempt.