본문으로 건너뛰기
L11.0

Level 11 — Agent Fundamentals

Level 10 gave a model access to carefully bounded tools. Level 11 adds a new responsibility: the system may choose what to do next more than once. After each action, it can inspect the new observation, update its state, and decide whether another step is useful.

That repeated decision process is the core idea behind an agent in this level. The model still does not receive unlimited authority. Application code still defines the tools, permissions, approvals, memory rules, and stopping limits.

Learning goal​

Turn a tool-using workflow into a bounded agent that can:

receive a goal
→ observe current state
→ choose the next permitted action
→ act through a bounded tool
→ record the result
→ update working state
→ decide whether to continue
→ stop for success, failure, approval, or budget
→ evaluate the full trajectory

A trajectory is the ordered record of observations, decisions, actions, and results across a run.

Entry skills​

You should already be able to validate structured tool requests and enforce permissions outside the model. You should also be able to bind approval to exact side effects, model legal state transitions, classify retryable failures, and evaluate recorded traces. Those Level 10 boundaries remain in force.

What mastery looks like​

By the end of Level 11, you should be able to:

  • explain what makes a repeated tool workflow agent-like without treating “agent” as a magic category;
  • implement an observe–decide–act loop with explicit state;
  • keep loops bounded with maximum steps, repeated-action detection, success conditions, and failure conditions;
  • split a larger goal into useful subgoals without treating a generated plan as truth;
  • keep short-lived working state separate from selected long-term memory;
  • choose tools from task evidence and tool contracts rather than tool names alone;
  • use critique as a focused check instead of an endless second opinion;
  • place human approval before high-impact actions and bind it to the exact proposed action;
  • recognize common agent failures such as looping, goal drift, stale state, unsafe tool use, and premature stopping;
  • evaluate full trajectories with task, safety, efficiency, and boundedness metrics;
  • package a deterministic bounded-agent run that another person can reproduce.

Running example​

Imagine a support agent with the goal:

“Resolve order 4172 if possible, but never issue a refund without approval.”

The agent may need to look up the order and inspect a policy record. It may then ask for approval. After that, it can execute one permitted refund or stop with a clear reason.

A useful run might be:

goal: resolve order 4172
observe: order status unknown
decide: get_order
act: get_order(4172)
observe: delivered, refund eligible
decide: request approval
observe: approval denied
stop: approval_denied

The run is successful if it follows the policy and reports the correct outcome. “Success” does not mean taking the largest possible number of actions.

Mini checkpoint​

After L11.6 — Long-Term Memory Patterns, complete the checkpoint on the agent boundary, observe–decide–act cycles, bounded loops, planning, working state, and long-term memory.

Level project​

After L11.13, complete Bounded Agent.

The required path uses standard-library Python and recorded trajectories. You will implement planning, action selection, working-memory updates, memory-write rules, approval checks, stopping conditions, state transitions, and run evaluation. No live provider or API key is required.

Three questions for every agent step​

At each step, ask:

  1. What changed? Record the new observation and the current state.
  2. What is allowed next? Apply tool, permission, approval, and transition rules outside the model.
  3. Why continue? A new step should have a reason. If the task is complete, blocked, unsafe, repeated, or over budget, stop.

These questions are more useful than asking whether a system “feels autonomous.”

Stable core and changing implementations​

This level focuses on stable engineering ideas: explicit loops, state, memory boundaries, stopping, approval, and evaluation. Agent frameworks and provider APIs change quickly. They can be useful later, but the required lessons do not depend on one vendor or framework.

Security guidance also changes. The lessons therefore keep authority checks in application code and use versioned or dated references where the curriculum requires them.

The point is not to memorize one fashionable agent architecture. It is to learn the boundaries that remain useful when frameworks change. Ask what the system knows now and what it may do next. Ask why it should continue. Then identify the evidence that proves the run stayed inside its limits.

Working rule​

An agent gets more opportunities to make decisions, not more permission to ignore controls.

Level 12 will build production-grade agent harnesses around these ideas. First, Level 11 makes the loop itself understandable, bounded, and measurable.

Lesson actions

Completion is stored locally on this device.

View progress