Error Handling and Retries
Goal
Classify tool failures, retry only when the failure is plausibly temporary or repairable, and enforce bounded retry and idempotency rules.
Tool workflows fail for ordinary software reasons. A service may time out. A record may not exist. Arguments may be invalid. A permission check may reject the request.
The important skill is not “retry on error.” It is deciding which error changes if you try again.
Separate failure categories
A useful reduced set is:
validation → arguments are wrong
permission → action is not allowed
not_found → requested resource is absent
rate_limit → service asks you to slow down
timeout → result is unknown/late
server → service failed internally
Blindly retrying validation or permission failures usually repeats the same mistake. A timeout or rate limit may be retryable under a bounded policy.
Repair is different from retry
If an argument is missing, the next step may be to construct a corrected request or ask the user. Sending the identical invalid call again is not progress.
Record whether the next attempt changed anything meaningful.
Side effects need idempotency thinking
Suppose a payment tool times out after the request is sent. Did the payment happen? If you retry without an idempotency mechanism, you may create a duplicate charge.
Read-only requests are often easier to retry than actions. For side effects, the workflow needs an operation ID, idempotency key, status check, or another application-specific protection.
Set a retry budget
Every retry loop should have a stopping condition such as:
maximum attempts = 3
maximum elapsed time = 10 seconds
retryable categories = {timeout, rate_limit}
After the budget is exhausted, return a clear failure or escalate. Do not let the model decide to loop forever.
Backoff reduces synchronized pressure
When a service asks clients to slow down, retrying instantly can make the problem worse. Increasing delays—often with jitter—can reduce contention.
The exact timing policy is an application/runtime detail. The stable principle is: retry behavior should be explicit, bounded, and observable.
Walk through one timeout carefully
Suppose get_order times out before returning data. Because it is read-only, a bounded retry is usually straightforward: no external state is changed by repeating the lookup. Now suppose refund_order times out after the request left your process. The missing response does not tell you whether the refund failed or succeeded. Repeating it immediately can duplicate the side effect.
A safer workflow first checks whether the operation has an idempotency key or a way to query completion. If neither exists, the correct next step may be escalation instead of automatic retry. The same error label—timeout—therefore leads to different actions because the operation semantics differ.
Track attempt identity
Record attempt number, call ID, operation/idempotency ID, error category, and next decision. This makes it possible to distinguish three independent failures: the service failed repeatedly, the workflow retried when it should not have, or the workflow exceeded its declared retry budget.
Predict
Run the local Lab
python labs/notebooks/level-10/l10-05-retry-policy.py
Each case is (category, attempt, side_effect, idempotent), and the loop allows max_attempts = 2.
- Run the command. You should see
validation -> stop_or_repair,timeout -> retry,timeout -> check_completion_before_retry(a timeout on a side effect that is not idempotent), andrate_limit -> stop_budget(already at attempt 2). - Use up the budget for the first timeout: change
("timeout", 0, False, False),to("timeout", 2, False, False),and rerun. That line becomestimeout -> stop_budget: the same kind of failure stops once the attempts are used up. - Change it back. Now change the first case,
("validation", 0, False, False),, to("timeout", 0, False, False),and rerun. It becomesretry. - Explain why: a validation error will fail the same way every time, so retrying cannot help, while a timeout may succeed on a second try. Change the case back afterward.
Loading lab…
Quick Check
Key Takeaways
- Classify failures before deciding to retry.
- Repairing an invalid request is different from repeating it.
- Side-effect retries require idempotency or completion checks.
- Retry budgets need explicit stopping conditions.
- Backoff and observability are part of reliable retry behavior.
Next Lesson
Next, represent tool workflows as explicit states and allowed transitions so retries, approvals, and stopping become easier to reason about.
References
- IETF, HTTP Semantics — RFC 9110.
- OpenAPI Initiative, OpenAPI Specification 3.1.0.
Completion is stored locally on this device.