본문으로 건너뛰기
L12.1

Agent Harnesses and Runtime Structure

Goal

Explain the jobs of an agent harness and separate model proposals, controller state, tool execution, and durable records into clear boundaries.

What the harness adds​

Level 11 gave the agent a bounded loop. That is enough to reason about one run, but a real application needs more structure around that loop. A process can restart, a tool can time out, a user can return later, and an operator may need to inspect what happened. The agent harness is the application layer that keeps those concerns organized.

A useful harness has several distinct responsibilities. The model proposes a next action. The controller checks whether that action is legal. Tool adapters perform external work. A durable store records task state that must survive a process restart. A trace recorder records observable events. Keeping these roles separate makes failures easier to locate because each boundary has a specific job.

Consider a refund agent. The model may propose refund_order with an amount. The controller checks the user role and approval. The tool adapter sends the request. The durable task record stores whether the refund attempt started, completed, or needs reconciliation—checking durable or external evidence to determine what actually happened before deciding whether to repeat the operation. If all of those responsibilities are hidden inside one large function, a timeout can leave the program unable to tell whether the refund happened.

user task → model proposal → policy check → tool adapter → durable result
↓
trace

Keep recovery state outside the process​

Temporary process memory and durable state are not the same thing. A local variable is useful while the process is alive. Durable state is information the system must recover later, such as task ID, current state, completed effect IDs, approval status, and retry count. If losing the process would make the information unsafe to reconstruct by guessing, that information belongs in durable state.

The harness also defines an execution boundary. Model output is data entering the controller, not executable authority. Tool results are observations entering state, not automatic permission to perform a new side effect. This is the same principle from Level 11, now arranged as a larger runtime structure.

One tempting design is a single agent function that calls the model, executes whatever tool it names, catches every exception, and loops. It looks simple, but it makes durability, testing, and recovery difficult. A smaller set of explicit components is easier to replay, substitute in tests, and inspect when a run stops unexpectedly.

Build a small explicit runtime​

A practical harness can still be small. You might have a TaskRecord, a Policy, a ToolRegistry, an Executor, and a TraceSink. The exact class names are not important. What matters is that state ownership and side effects are visible. The controller should be able to answer what task is running, what operation is next, and what evidence justifies continuing.

Observability belongs beside execution rather than being added only after a failure. When the harness starts a task, chooses an action, calls a tool, retries, blocks for approval, or stops, it should emit structured events. These events will become important in L12.4 when you learn tracing in more detail.

Runtime structure is therefore not a framework brand or a special kind of model. It is an engineering arrangement. You can implement the same boundaries with simple Python, a workflow engine, or a larger distributed system. The learning target is the separation of responsibilities, because that survives changes in libraries and providers.

Predict

A process crashes after a side-effecting tool call. Which information is most important to have outside temporary process memory?

Run the local Lab​

Run:

python labs/notebooks/level-12/l12-01-harness-structure.py

The script moves one task record through four steps: proposal, policy, execute, and complete. Only execute performs a real side effect. Partway through, the script simulates a crash: everything is lost except the saved task record.

  1. Run it unchanged. RESTART_AFTER = "policy", so the saved record has 'state': 'AUTHORIZED' and 'operation_id': None.
  2. Read resume from: execute. The controller looks at the saved state to decide which step comes next. The run then finishes with effect_count: 1.
  3. Open the script and change RESTART_AFTER = "policy" to RESTART_AFTER = "execute". Rerun.
  4. Now the saved record already has 'state': 'EXECUTED' and 'operation_id': 'read-order-4172', and the output says resume from: complete. effect_count: is still 1, so the side effect did not run twice.

Answer this before moving on: which two fields in the saved record let the controller skip the side effect after the restart?

Loading lab…

Quick Check

1. What is the main role of an agent harness?
2. Which information most clearly belongs in durable state?
3. Why separate model proposals from tool execution?

0 of 3 questions answered.

Explain it back​

Describe a four-part harness for a payment-support agent: model proposal, controller, tool adapter, and durable task record. State one fact each part should own and one fact it should not own.

Key Takeaways

  • A harness organizes control around the agent loop.
  • Durable state must survive process loss.
  • Model proposals are not execution authority.
  • Tool adapters isolate side effects from decision logic.
  • Structured events should be emitted as part of normal execution.

Next Lesson

Next, L12.2 — Task Queues and Durable State explains how work can wait, resume, and survive a worker restart.

References

Lesson actions

Completion is stored locally on this device.

View progress