Skip to main content
L10.3

Validate Structured Tool Arguments

Goal

Validate model-produced tool arguments in layers—parse, shape, domain rules, authorization, and workflow state—before execution.

Imagine a student submits a field-trip form. The office does not jump from "I can read the handwriting" straight to "permission granted." It checks the form format, whether the requested date is allowed, whether the student is eligible, whether the guardian approval exists, and whether registration is still open.

Tool validation is the same layered idea. Parsing asks whether the data can be read. Schema checks ask whether required fields and types are present. Domain checks ask whether the values make sense for the application. Authorization checks trusted identity and permission. Workflow-state checks ask whether this action is legal right now. Passing one layer never silently proves the later layers.

A schema gives you something concrete to check. The next step is to treat every model-produced argument object as untrusted input.

A useful validation sequence is:

parse
→ schema/type checks
→ domain checks
→ authorization
→ workflow-state check
→ execution

Each step can reject the request for a different reason.

Parsing is only the first boundary​

The text:

{"order_id":"4172","quantity":2}

may be valid JSON. That does not prove quantity is allowed for the chosen tool, that order 4172 belongs to the user, or that the workflow is currently allowed to modify it.

Treat “valid JSON” as a syntax result, not a safety result.

Domain checks catch valid-but-impossible values​

A schema may say quantity is an integer, while the business rule says it must be between 1 and the remaining stock.

Likewise, a date can be a valid string but outside the allowed booking window. These constraints should be deterministic application rules when possible.

Authorization uses trusted identity​

Do not let the model choose which tenant or user it is authorized to act as. The current authenticated principal should come from trusted session/application state.

If the tool arguments contain customer_id, application code can compare that value with the principal or ignore it and inject the trusted ID itself.

Workflow state matters too​

A cancellation may be syntactically valid and authorized but disallowed after shipment. The current state therefore participates in validation.

This is important for multi-step workflows: the same action can be valid in one state and invalid in another.

Return structured validation errors​

A useful error tells the next step what failed without leaking secrets:

{
"ok": false,
"category": "validation",
"field": "quantity",
"message": "quantity must be between 1 and 5"
}

The model can then decide whether it can repair the request, ask the user, or stop.

Fail at the earliest deterministic check​

Validation order is useful for diagnosis. If JSON cannot be parsed, authorization is not yet the relevant problem. If the shape is valid but a numeric limit fails, report the domain failure before calling a remote service. If all deterministic fields are valid but the principal lacks permission, stop at authorization.

This ordering makes error traces smaller and safer. It also avoids leaking whether an inaccessible resource exists: an application can choose an authorization policy that rejects before revealing sensitive resource details.

Never “repair” trusted fields from model guesses​

A model may try to fix a rejected request by changing customer_id, tenant, role, or approval metadata. Those values are not ordinary arguments if they represent trusted identity or authority. The repair loop may change user-controlled task fields, but trusted principal and policy inputs must continue to come from application state. Otherwise a helpful-looking retry can turn into privilege escalation.

Predict

Arguments pass JSON Schema but request another customer's order. Which layer should reject the call?

Run the local Lab​

python labs/notebooks/level-10/l10-03-validate-arguments.py

The script checks three model-proposed requests against a trusted principal_customer_id = "c-1" (the logged-in user) and state = "ORDER_LOADED".

  1. Run the command. The results are ok, permission (the request names customer c-2), and domain (quantity 9 is more than the stock of 5).
  2. Simulate a different logged-in user: change principal_customer_id = "c-1" to "c-2" and rerun.
  3. Now the first request becomes permission and the second becomes ok. The third is still domain, because the quantity check runs before the permission check.
  4. Notice what you did not do: you did not edit the model-provided requests to “log in.” Identity comes from trusted session state, never from the arguments. The script then stops with an AssertionError. That is expected: its final checks were written for the original values. Change the value back and rerun to see it pass.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-10/l10-03-validate-arguments-exercise.py

Implement the schema, domain, principal-permission, and workflow-state checks yourself. The assertions require the three requests to end as ok, permission, and domain.

Run the starter after each change:

python3 labs/notebooks/level-10/l10-03-validate-arguments-exercise.py

A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.

Quick Check

1. A field has the correct type but violates a business limit. Which check is missing?
2. Where should the trusted user identity normally come from?
3. Why check workflow state before execution?

0 of 3 questions answered.

Key Takeaways

  • Validate tool arguments as untrusted input.
  • Syntax, schema, domain rules, authorization, and workflow state are distinct checks.
  • Trusted identity should come from application state.
  • Valid structure does not imply a permitted action.
  • Structured errors support safe repair or escalation.

Next Lesson

Next, normalize tool results and return them to model context without confusing data with trusted instructions.

References

Lesson actions

Completion is stored locally on this device.

View progress