본문으로 건너뛰기
L12.12

Operating an Agent in Production

Goal

Combine health checks, release rules, incident evidence, rollback, and safe recovery into an operating routine for a deployed agent harness.

A canary fails before full rollout​

Suppose you deploy agent-controller v18 to only 5% of traffic. Most tasks finish normally, but one run sends an external action without the approval that v17 required.

That one event changes the job immediately. Do not ask only, “Is the new version mostly working?” Ask which version handled the task. Check whether its dependencies were healthy and which trace led to the action. Then decide whether to stop the rollout and how to keep this failure from returning.

A safe response can look like this:

approval bypass appears in canary
→ stop expanding traffic
→ identify affected version and task
→ roll back to the known-good controller
→ preserve trace and incident record
→ add the failure to regression tests

This is what operating an agent in production means. The harness is already built. Now you keep it dependable while releases, traffic, dependencies, and policies change. Version records, health checks, release gates, rollback rules, incident records, and capacity signals make that response deliberate instead of improvised.

Version, check health, and gate releases​

Start with versioned releases. Record the controller version, policy version, tool schema version, evaluator version, and relevant model configuration with each run. When a metric changes, version data helps determine whether the change began after a release rather than guessing from memory.

Health checks answer whether required components are available. A queue may be reachable while a critical tool is not. A shallow health check that only says 'process is alive' can miss a broken dependency. Use checks that are safe and lightweight but still represent the dependencies needed for service.

Release gates should use evaluation evidence. Before increasing traffic, run deterministic fixtures and simulations, then compare key metrics to thresholds. Critical safety constraints stay zero tolerance. Performance targets can use allowed ranges or percentiles.

Rollouts can be gradual. Instead of sending all tasks to a new version, a small fraction can use the new release while the old release remains available. If failures rise, rollback should be an operational action with clear criteria, not an improvised debate during an incident.

{"release":"v18","traffic_percent":5,"task_success":0.94,"approval_bypasses":0,"decision":"continue_canary"}

Use incidents as evidence​

Incidents require evidence. A useful incident record includes affected versions, task IDs, trace IDs, first observed time, failure category, user impact, and temporary mitigation. This record helps separate symptoms from causes and supports later regression fixtures.

Recovery should preserve idempotency and task state. Restarting workers is not enough if queued tasks can repeat side effects without stable operation IDs. Operational procedures must respect the same duplicate-safety rules used in normal execution.

Watch capacity and close the loop​

Capacity matters too. Queue depth, worker saturation, external rate limits, and approval wait can all create backlog. Monitoring only CPU can miss a system that is healthy at the process level but unable to finish tasks within expected time.

Production operation is therefore a feedback loop: observe metrics and traces, classify the problem, mitigate safely, update tests, and release a verified fix. The loop should improve the harness without turning every incident into a one-off exception to policy.

Predict

A new harness version causes approval-bypass count to rise from zero to one during canary traffic. What should the release process do?

Run the Docker-environment Lab preflight​

The full activity is registered for the Docker-oriented environment used by later operational work. Start with this deterministic Python preflight from your repository checkout so you can separate controller-logic failures from Docker, cloud, GPU, or credential setup problems.

Run:

python labs/notebooks/level-12/l12-12-operations.py

The Lab evaluates a small canary release report against health, reliability, latency, and safety gates. Run it once unchanged. Then set approval_bypasses from 0 to 1 and observe that release fails even though the other metrics remain within target.

Loading lab…

Quick Check

1. Why record component and policy versions with each run?
2. What is the purpose of a gradual rollout?
3. Which post-incident artifact most directly helps prevent recurrence?

0 of 3 questions answered.

Explain it back​

Describe a safe release process for a new agent controller: pre-release tests, canary evidence, rollback trigger, and the information recorded if an incident occurs.

Key Takeaways

  • Production runs should carry version metadata.
  • Health checks should represent important dependencies.
  • Release gates combine tests, simulations, performance, and safety constraints.
  • Gradual rollout limits exposure and supports rollback.
  • Incidents should become future regression evidence.

Next Lesson

Next, L12.13 — Agent Harness Integration Workshop brings all Level 12 boundaries together in one end-to-end controller.

References

Lesson actions

Completion is stored locally on this device.

View progress