Skip to main content
L10.13

Tool-Using Assistant Integration Workshop

Goal

Integrate tool selection, schema validation, permissions, retries, state transitions, multimodal evidence, and evaluation into one reproducible assistant trace.

The final Level 10 workflow is not “LLM + tools.” It is a sequence of controlled boundaries around model proposals and external observations.

A reviewable trace looks like:

request
→ trusted principal + workflow state
→ supplied text/image/audio evidence
→ model proposal
→ tool/schema check
→ permission/approval check
→ execution
→ normalized result
→ state transition
→ next proposal or final response
→ evaluation record

Keep model choice separate from application authority​

The model may choose get_order because the request asks about delivery. The application decides whether that tool is available, whether order 4172 is authorized, and whether the arguments are valid.

For a side-effect tool, the application also checks approval and idempotency requirements.

This division makes the workflow reviewable even if the model implementation changes.

Use one trace ID across the workflow​

A single request may create several tool calls. Give the run a stable trace ID and each call a call ID.

Record:

  • state before call;
  • proposal;
  • validation result;
  • policy/approval decision;
  • attempt number;
  • normalized tool result;
  • state after call;
  • final outcome.

This is enough to reconstruct why an action happened.

Make failures first-class project cases​

The project should include at least:

  • invalid argument;
  • permission denial;
  • retryable timeout;
  • non-retryable not-found;
  • approval-required action;
  • prompt-injection-like text in tool data;
  • multimodal evidence uncertainty;
  • max-step or retry-budget stop.

A workflow that handles only the happy path has not demonstrated reliable tool use.

Release rules should include safety constraints​

Example:

task success >= 0.85
tool selection accuracy >= 0.90
argument validation accuracy == 1.00
unauthorized executions == 0
approval bypasses == 0
runaway workflows == 0

A small gain in task success cannot compensate for an authorization bypass.

The required project path is provider-neutral​

The Level project uses recorded fixtures and Python standard-library logic. It checks the engineering boundaries without requiring a live model, API key, browser, microphone, or image service.

You may add a real model or API as extra evidence. Keep the recorded deterministic path so another person can rerun the same policy and trace checks.

Review capability, then authority​

First review capability correctness. Did the workflow select the intended tool? Did it validate arguments, normalize results, and reach the expected task outcome?

Then review authority correctness. Was every executed action permitted for the principal? Was approval present when required? Did the run stay within its retry and state bounds? Did untrusted instructions remain unable to grant authority?

Keeping these passes separate prevents a successful outcome from excusing an unsafe route. A refund can be factually appropriate and still be a security failure if it bypassed approval.

Make the project extensible without weakening the baseline​

If you add a real model or provider API, keep the same logical tool names and recorded fixture cases. Keep the same security rules and metrics where practical. A live run may add latency and model-quality evidence, but it should not remove the deterministic checks. The two paths serve different purposes: reproducible engineering validation and realistic external-system behavior.

Bridge to agents​

Level 11 will let the system choose actions over longer loops. The boundaries built here remain the foundation: explicit state, narrow tools, trusted permissions, bounded retries, and stop conditions. An agent is not safer because it is called an agent; it is safer when these workflow controls continue to hold across more steps.

Predict

A model selects the right refund tool, but the application executes it without required approval. Which boundary failed?

Run the integration Lab​

python labs/notebooks/level-10/l10-13-integration-check.py projects/reference/l10/sample-run/tool-run.json

Then run the intentional failure fixture:

python labs/notebooks/level-10/l10-13-integration-check.py projects/reference/l10/failure-run/tool-run.json

The first should pass and the second should identify an approval or authorization violation.

Loading lab…

Quick Check

1. What is the strongest reason to use one trace ID across a multi-tool request?
2. Why include intentional failure cases in the project?
3. Can higher task accuracy compensate for one unauthorized execution in the example release rule?

0 of 3 questions answered.

Explain it back​

Trace one request from user input to final response. Identify every place where the model supplies a proposal or interpretation, every place where application code supplies authority, and the evidence you would log if the run failed.

Key Takeaways

  • Reliable tool use is a chain of explicit application boundaries around model proposals.
  • Keep trusted principal, permissions, approvals, and workflow state outside model control.
  • Trace validation, execution, results, retries, and transitions with stable IDs.
  • Evaluate task success, component correctness, and security constraints separately.
  • Provider-neutral deterministic fixtures make the engineering claims reproducible.

Next Lesson

Complete Multimodal Tool-Using Assistant. Level 11 will reuse these tool and state boundaries inside longer-running agent loops.

References

Lesson actions

Completion is stored locally on this device.

Level project unlocked: Multimodal Tool-Using Assistant

View progress