Stopping Conditions
Goal
Give an agent explicit reasons to stop—success, a real blocker, denial, repetition, or a resource limit—and test that it stops before another useless or unsafe action.
Start with an agent that keeps trying
Suppose an agent is asked to find a meeting slot and book it after approval. It checks the calendar, proposes Tuesday at 14:00, asks for approval, and the user denies the proposal.
What should happen next?
A poorly bounded loop may keep trying the same booking, ask for approval again, or call more tools because the model can still produce another action. A better controller has a stop rule for the state it has reached:
proposal denied
→ do not book
→ record terminal reason: denied
→ stop this run
Stopping is part of the agent design, not a feeling the model develops when it is “done.” Different endings mean different things. Success means the requested result has been reached and can be checked. Blocked means a required dependency is unavailable. Denied means an authorized person rejected an action. Repeated means the loop is no longer making progress. Budget exhausted means the run has reached a pre-set limit.
For success in particular, use something observable. “The model says the task is complete” is weaker than “the booking API returned confirmation B-204 for the approved time.”
Blockers are not ordinary failures
Suppose the agent needs a shipping address that the user has not provided and no tool can retrieve it.
Repeatedly searching will not help. The correct terminal state may be BLOCKED_MISSING_INPUT.
That outcome can lead to a user question. It is different from FAILED_TOOL_ERROR, which may need retry or service repair.
Naming the blocker lets the system choose the right follow-up.
Combine local and global stop rules
Local rules react to recent behavior:
- same action repeated too many times;
- critique/revision budget exhausted;
- non-retryable error;
- permission denied.
Global rules apply to the whole run:
- maximum steps;
- maximum cost;
- maximum elapsed time;
- overall task deadline.
A run should stop when any mandatory bound is reached, even if the model proposes another useful action.
Stop before unsafe action
A policy denial is itself a valid stopping or escalation reason.
Do not execute the action first and record a safety failure afterward. The stopping boundary belongs before side effects.
For a high-impact proposal, the sequence can be:
proposal
→ policy check
→ approval check
→ if denied: stop or request human action
→ if allowed: execute
Evaluate stopping quality
Agent evaluation should ask more than “did it eventually stop?”
Useful measures include:
- task success rate;
- premature-stop rate;
- budget-exceeded rate;
- repeated-action stop rate;
- unsafe-execution count;
- average useful steps for successful tasks.
A system that stops quickly by giving up on every hard case is bounded but not useful. A system that succeeds only after many unnecessary calls may be correct but inefficient.
AgentBench evaluates agents across interactive environments and highlights that multi-turn decision quality and long-horizon behavior matter; your local project will use a much smaller deterministic test set rather than trying to reproduce a research benchmark.