LLMs That Use Tools
Goal
Explain tool use as a model proposing a structured action that application code decides whether and how to execute, and trace the boundary between language generation and side effects.
A normal language-model call returns text or structured data. A tool-using workflow adds another possibility: the model can return a request for an external action.
Consider a weather question. The model may have old training knowledge, but the application can expose a tool that reads current weather data. The useful mental model is not “the LLM browses the weather service. ” It is:
model proposes tool + arguments
→ application validates proposal
→ application executes permitted code
→ application returns result
→ model produces the next response
That separation keeps authority in application code.
A tool call is not the side effect
Suppose a model produces:
{"tool":"get_order","arguments":{"order_id":"4172"}}
This object has not accessed anything yet. It is data. The application can reject it if the tool is unavailable, the argument is malformed, the user lacks permission, or the current workflow state does not allow an order lookup.
This is the same idea you used for structured outputs in Level 7, now attached to a possible action. Parsing is not authorization.
Tools extend capability and failure surface
Tools can provide information the model does not already have, perform deterministic calculations, or trigger real operations. They also introduce new failure classes:
- wrong tool chosen for the task;
- correct tool with wrong arguments;
- unauthorized request;
- network or service failure;
- tool result misunderstood;
- repeated calls that never converge;
- destructive side effect executed too early.
A reliable workflow records enough of the trace to tell these apart.
Use the narrowest useful interface
Compare:
run_any_sql(query)
with:
get_order_status(order_id)
The second interface expresses less power. That is often a feature. Narrow tools are easier to validate, test, authorize, and explain.
Do not give a model a broad capability merely because a prompt asks it to behave carefully. Permission design should reduce what the tool can do even when the model is confused.
One request can need zero, one, or several tools
A model should not be rewarded for using a tool when the answer is already available in trusted context. Tool use has latency, cost, and failure risk. Conversely, a model should not invent a live value when a tool is required to obtain it.
For multi-step tasks, each result can change the next decision:
lookup order
→ result says delivered
→ check refund policy
→ result says approval required
→ request human approval
Later lessons will make that sequence an explicit state machine.
Tool use changes the system boundary, not the model's nature
The model still predicts an output from context. Tool use simply gives one kind of output a special interpretation: application code may treat it as a request to invoke a named capability. This matters because the same model can be used in a read-only assistant or a workflow that can create side effects; the difference is largely in the tools and policies wrapped around it.
When debugging, first ask whether the model proposed the wrong action or whether the application mishandled a reasonable proposal. Mixing these together leads to prompt changes for bugs that actually belong in permissions, schemas, or execution code.
Predict
Run the local Lab
From the repository root:
python labs/notebooks/level-10/l10-01-tool-boundaries.py
- Run the command. You should see three lines:
answer_from_context -> answer_without_tool,read_order -> execute_read, andrefund_order -> request_approval. - Read
application_decisionin the script. The model only suppliestool; the application decides what happens, and it treatsrefund_orderdifferently from the read-onlyget_order. - In
proposals, change the refund entry from"approved": Falseto"approved": Trueand rerun. - The last line becomes
refund_order -> execute_action. The model's proposal did not change at all—only the approval state, which the application owns, changed the decision. The script then stops with anAssertionError. That is expected: its final checks were written for the original values. Change the value back and rerun to see it pass.
Loading lab…
Quick Check
Key Takeaways
- Tool use adds an application-controlled action boundary around model output.
- A valid tool request is not proof that the action is permitted.
- Narrow interfaces reduce unnecessary authority.
- Tool use introduces selection, argument, execution, result, and loop failures.
- Record the trace so each failure boundary can be inspected.
Next Lesson
Next, define tool interfaces precisely with typed schemas so the model and application share the same argument shape.
References
- Schick et al., Toolformer: Language Models Can Teach Themselves to Use Tools.
- Yao et al., ReAct: Synergizing Reasoning and Acting in Language Models.
Completion is stored locally on this device.