Skip to main content

Level 12 — Agent Engineering and Harnesses

From a bounded agent to a durable system​

Level 11 taught you how to keep an agent bounded: explicit goals, legal actions, memory rules, approval points, stopping conditions, and trajectory evaluation. Those controls answer an important question: what is the agent allowed to do during one run? Here, durable means stored so important task state can survive a worker or process stopping. Level 12 asks the next engineering question: what happens when that run must survive real operational problems? Those problems include process restarts, duplicate deliveries, partial failures, concurrent work, expensive tools, and missing visibility.

The main idea of this level is the agent harness. A harness is the application structure around an agent that manages work before, during, and after model decisions. It stores durable task state, schedules work, enforces permissions, records traces, retries recoverable operations, protects side effects from duplication, and measures reliability. The model still proposes decisions, but the harness owns the operational rules that make repeated execution dependable.

A useful way to picture the difference is to imagine a delivery driver and a dispatch system. The driver decides how to handle the next local situation. The dispatch system does different work. It keeps the job record, assigns work, and records timestamps. It notices missed deliveries, prevents the same package from being marked delivered twice, and gives operators a history of what happened. An agent harness plays a similar role for tool-using software.

What you will build​

You will begin by separating durable controller state from temporary process memory. A Python process can disappear; a durable task record should not. You will then study queues, retries, idempotency, and recovery. Idempotency means designing an operation so repeating the same request does not create an unintended second effect. This matters whenever a timeout makes the caller unsure whether a side effect actually happened.

Next, you will make runs observable. Observability means collecting structured evidence that helps answer questions about a running system without guessing. In this level that evidence includes traces, events, timings, action names, task identifiers, retry counts, permission decisions, and terminal reasons. A trace is not a hidden chain of thought. It is an application-owned record of visible decisions, tool boundaries, state changes, and results.

Security remains part of the controller. You will treat sandboxing and permissions as different but related controls. A permission rule answers whether an action is allowed. A sandbox restricts what code or a process can reach even when something inside it behaves unexpectedly. You will also revisit memory, this time from an operations perspective: retrieval limits, freshness, compression, provenance, and avoiding summaries that silently remove a safety-relevant fact.

The second half of the level expands one run into a small system. Independent tool calls may run in parallel when their dependencies allow it. Larger goals may be delegated to subagents, but delegation does not transfer unlimited authority. The parent must define the subtask, allowed tools, resource budget, evidence expected back, and how the result will be checked before it affects the parent run.

How you will test it​

Testing also becomes more realistic. A single happy-path example is not enough for a durable agent. You will build deterministic fixtures for known cases, then simulation-based evaluations that vary delays, failures, user behavior, and tool results. The point is not to predict every future event. The point is to expose fragile controller assumptions before a real user depends on them.

Finally, you will connect reliability to cost and latency. Faster is not automatically better, and cheaper is not automatically safer. A production harness should make tradeoffs visible. Record how many tool calls were made and how long critical steps took. Record retries and waiting time. Also check whether a cost-saving shortcut changed success or safety metrics.

Level Project​

The Level Project, Production-Grade Agent Harness, asks you to integrate these ideas into a deterministic controller package. You will build a queue-backed task model, replay-safe side effects, structured tracing, permission checks, bounded delegation, evaluation fixtures, and release metrics. The project uses local deterministic validation so you can prove controller behavior without needing a live model, cloud account, or paid API.

What mastery looks like​

By the end of Level 12, you should be able to explain what an agent decided. You should also explain how the surrounding system kept that decision process durable, inspectable, recoverable, and bounded under failure.

Key Takeaways

  • An agent harness is the operational system around model decisions.
  • Durable state, replay-safe side effects, tracing, permissions, and recovery belong in application control.
  • Parallelism and delegation require explicit dependency and authority boundaries.
  • Deterministic fixtures and simulations test failure behavior before deployment.
  • Cost, latency, and reliability should be measured together rather than optimized in isolation.

Next Lesson

Start with L12.1 — Agent Harnesses and Runtime Structure to separate model decisions from the controller that owns durable execution.

Lesson actions

Completion is stored locally on this device.

View progress