본문으로 건너뛰기
L12.10

Simulation-Based Evaluation

Goal

Evaluate an agent harness across controlled simulations that vary delays, failures, tool responses, and user behavior while preserving reproducible seeds and metrics.

Vary conditions without risking production​

A flight simulator is useful because it can repeat difficult conditions without putting a real aircraft at risk. You can choose crosswind, equipment faults, or visibility and see whether the procedure still works. The test is useful only if you know which conditions were used and can recreate them.

Simulation-based evaluation does the same for an agent harness. Instead of waiting for production to produce a rare timeout or approval delay, a simulator creates controlled variations. A recorded random seed is like the recipe for one simulated run: it lets you reproduce the same generated conditions when a failure needs investigation. Randomness creates coverage; recording it preserves evidence.

Unlike a fixture, which checks one exact known case, simulation varies controlled conditions across many runs. Keep the scenario definition and seed so a failed run such as run 37 can be recreated with the same inputs.

Choose meaningful simulation variables​

Choose simulation variables that reflect real uncertainty. For a support agent, you might vary whether an order exists, whether a tool times out, whether approval arrives, and how long each service takes. You would not randomly change unrelated fields merely to increase the number of combinations.

seed = 37
timeout_probability = 0.10
approval_delay_probability = 0.20
runs = 100

Measure safety separately from success​

Metrics should connect to harness goals. Task success is useful, but it is not enough. Track unauthorized executions, approval bypasses, duplicate side effects, runaway loops, recovery success, retry count, latency, and cost. A system that succeeds often by violating approval rules is not acceptable.

Zero-tolerance constraints should stay separate from average metrics. If unauthorized execution must never occur, a high overall success rate cannot compensate for one unauthorized side effect. Release rules can therefore combine minimum quality thresholds with maximum counts of critical violations.

Simulations are also useful for comparing controller changes. Suppose one retry policy increases task success from 90% to 94% but doubles duplicate-risk incidents in a failure scenario. The evaluation should make that tradeoff visible rather than reporting only the improved success number.

Know what the simulator cannot prove​

Be careful not to overfit to the simulator. A simulated environment is still a model of reality. It can miss failure modes, user behavior, or timing interactions. Use simulation to test hypotheses and expose fragility, then combine it with deterministic fixtures, staging tests, and real monitoring.

Statistical interpretation matters as the number of runs grows. A difference observed in ten runs may disappear across a thousand. This lesson keeps the arithmetic simple, but you should at least report the number of runs and avoid presenting a tiny sample as a universal guarantee.

A simulator is most valuable when failure diagnosis remains possible. Record the generated conditions and trace IDs for failed runs. That connects aggregate metrics back to individual trajectories you can inspect.

Predict

A simulation reports 96% task success but one unauthorized side effect. The release rule allows zero unauthorized executions. What is the result?

Run the Docker-environment Lab preflight​

The full activity is registered for the Docker-oriented environment used by later operational work. Start with this deterministic Python preflight from your repository checkout so you can separate controller-logic failures from Docker, cloud, GPU, or credential setup problems.

Run:

python labs/notebooks/level-12/l12-10-simulation.py

The Lab runs a seeded set of 100 simulated tasks with controlled timeout and approval probabilities. Change the retry_limit from 1 to 3 and rerun. Compare success, retry count, and duplicate-risk detection rather than looking at success alone.

Loading lab…

Quick Check

1. What makes a simulation reproducible?
2. Why keep zero-tolerance controls separate from average success?
3. What is a limitation of simulation?

0 of 3 questions answered.

Explain it back​

Design a 100-run simulation for a tool-using agent. Name three variables, four metrics, one zero-tolerance rule, and the evidence you would keep for replaying a failed run.

Key Takeaways

  • Simulation explores controlled variations beyond fixed fixtures.
  • Record seeds and generated conditions for replay.
  • Measure safety, recovery, latency, and cost as well as task success.
  • Do not average away critical violations.
  • Simulation complements rather than replaces real monitoring and exact tests.

Next Lesson

Next, L12.11 — Cost, Latency, and Reliability shows how operational tradeoffs can be measured together.

References

Lesson actions

Completion is stored locally on this device.

View progress