본문으로 건너뛰기
L11.3

Agent Loops

Goal

Implement the outer control loop for an agent with an explicit step budget, terminal states, and progress checks so repeated decision making cannot run without a bound.

The observe–decide–act cycle describes one turn. An agent loop repeats that cycle until a stopping condition is reached.

A minimal sketch looks like this:

while not done:
observation = observe(state)
decision = decide(state, observation)
result = act(decision)
state = update(state, result)

The sketch is useful, but it is incomplete. If done never becomes true, the run can continue forever. If the same failed action is chosen repeatedly, cost grows while progress stays at zero.

Agent loop: observe, decide, act, repeat—or stop

The model proposes the next step, but application code decides what is allowed, when to continue, and when the run must stop.

Stage 1/6 · Observe
ObserveDecideCheck limitsActUpdateRepeat?continueStop reasonsuccess · blockeddenied · failed · budget
Terminal reasonMeaning
SUCCESSThe goal is satisfied.
BLOCKEDRequired information or permission is unavailable.
DENIEDPolicy or human approval rejected the action.
FAILEDA non-recoverable error ended the run.
BUDGET_EXCEEDEDThe configured step or cost limit was reached.

Put the bound outside the model​

A prompt can say “finish within eight steps,” but application code should enforce the number.

For example:

for step in range(max_steps):
...
else:
stop_reason = "max_steps"

The model may suggest continuing on step nine. The loop controller should still stop if the configured budget is eight.

This is the same principle as tool permissions: important safety and cost boundaries should not depend only on generated text.

Terminal states should be explicit​

A run should not merely “end.” Record why it ended.

Useful terminal states include:

  • SUCCESS — the goal is satisfied;
  • BLOCKED — required information or permission is unavailable;
  • DENIED — policy or human approval rejected the next action;
  • FAILED — a non-recoverable error occurred;
  • BUDGET_EXCEEDED — the step or cost limit was reached.

Different stop reasons need different follow-up behavior. A blocked run might ask the user for missing information. A denied action should not be retried as if it were a timeout.

Detect repeated actions​

Imagine an agent that calls search_orders("4172") five times and receives the same result each time. The step budget eventually stops it, but repetition detection can stop it earlier.

One simple rule is:

if same normalized action appears 3 times without new evidence:
stop_reason = repeated_action

The threshold is application-specific. The key idea is to make “no progress” observable.

Count attempts even when tools fail​

If a timeout occurs, the attempt still consumed time and possibly money. Step budgets should not reset just because an action did not return a normal result.

Retry rules from Level 10 remain separate. A retry may be allowed, but it still belongs inside the outer agent budget.

The loop controller owns continuation​

The model can propose continue, answer, ask_user, or a tool action. The controller maps those proposals into legal transitions.

This means a model output cannot invent a new terminal state or skip a required approval. It chooses within a controlled vocabulary.

Predict

An agent has max_steps=5 and produces a reasonable proposal for step 6. What should the controller do?

Run the local Lab​

Run:

python labs/notebooks/level-11/l11-03-agent-loop.py

The script contains one run that succeeds and one run that repeats an action.

  1. Run it unchanged and record the terminal reason for both traces.
  2. In run(...), the default is max_steps=4. Before editing, predict which trace will change terminal reason if only the step budget becomes 2.
  3. Change only max_steps=4 to max_steps=2, then rerun.
  4. Compare the terminal reasons. Explain why the same proposed actions can end differently when the controller's outer budget changes.

Loading lab…

Write the core logic yourself​

Open:

labs/notebooks/level-11/l11-03-agent-loop-exercise.py

Implement success, repetition, and step-budget stopping in the outer loop. The starter deliberately contains no completed loop logic.

Run the starter after each change:

python3 labs/notebooks/level-11/l11-03-agent-loop-exercise.py

A correct implementation ends with a PASS: marker. Only after you have a working version, compare your approach with the solved deterministic script used by the Level smoke tests.

Optional: put a real model inside the bounded loop​

Run:

python labs/real-model/l11_agent_loop_real_model.py --max-steps 3

Here the model can choose the order-status tool and then see the tool result, but the controller remains in charge of continuation. It rejects undeclared tools, rejects malformed arguments, blocks an identical repeated action, and stops when the configured step budget is exhausted.

Rerun with --max-steps 1. The model has not changed, but the outer controller now permits less work. That contrast is the point: model reasoning happens inside application-owned bounds.

Quick Check

1. Where should a maximum step limit be enforced?
2. Why record a terminal reason instead of only recording that the loop ended?
3. What is the purpose of repeated-action detection?

0 of 3 questions answered.

Explain it back​

Describe two different ways the same agent can stop: one because the task is complete and one because useful progress has stopped. Explain which application field would distinguish those outcomes.

Key Takeaways

  • Agent loops need application-enforced bounds.
  • Terminal states should record why a run ended.
  • Repetition detection can identify stalled progress before the full budget is used.
  • Failed attempts still consume the outer run budget.
  • The loop controller decides which continuation and stop transitions are legal.

Next Lesson

Next, learn how a generated plan can organize a longer goal without becoming an unquestioned script.

References

Lesson actions

Completion is stored locally on this device.

View progress