Agent Evaluation Basics
Goal
Evaluate an agent with fixed trajectories and separate metrics for task success, action selection, boundedness, approval safety, memory behavior, and efficiency.
A final answer is only one part of an agent run.
Two agents can produce the same answer while taking very different paths. One may use two read-only calls. Another may loop for ten steps. It might retrieve stale memory, request an unnecessary side effect, and only then arrive at the same text.
If evaluation checks only the final answer, those trajectories look equal. They are not.
Evaluate the trajectory
A useful fixed evaluation case includes:
- starting goal and principal;
- initial state;
- available tools;
- expected important observations;
- allowed and disallowed actions;
- approval requirements;
- maximum steps;
- expected terminal category;
- expected final outcome.
Then the evaluator can inspect both what happened and how it happened.
Use separate metric families
For a small bounded-agent suite, useful metrics include:
Task quality
- task success;
- correct terminal state;
- required subgoals completed.
Decision quality
- next-action selection accuracy;
- unnecessary-action count;
- unsupported-action count.
Boundedness
- max-step violations;
- repeated-action loops;
- average steps on successful cases.
Authority and oversight
- unauthorized executions;
- approval bypasses;
- approval-scope mismatches.
Memory
- stale-memory uses;
- cross-scope retrievals;
- prohibited persistent writes.
One average score cannot communicate all of these.
Keep zero-tolerance constraints separate
Suppose five fixed tasks produce:
task success = 4 / 5 = 0.80
unauthorized executions = 1
A single combined score might still look acceptable. A safer release rule can say:
task success >= 0.80
AND unauthorized executions == 0
AND approval bypasses == 0
AND runaway loops == 0
This lets ordinary quality thresholds coexist with hard safety constraints.
Slice the failures
A global success rate can hide one weak class of tasks.
Create slices such as:
- read-only investigation;
- approval-required action;
- missing-input blocker;
- memory retrieval;
- repeated-action trap.
Then compare performance per slice.
If the agent succeeds on every read-only case but fails most approval-required cases, the slice tells you where to investigate.
Evaluate deterministic controls directly
Not every test needs a model.
You can unit-test:
- step-budget enforcement;
- approval binding;
- state transitions;
- memory freshness filters;
- repeated-action detection;
- release-rule recomputation.
These tests are stable and fast.
Model-dependent evaluation can then focus on selection and interpretation behavior.
Record evaluator version and fixtures
An evaluation result is only meaningful if another person knows what was tested.
Record the fixture set and evaluator version. Also record the policy and agent versions. If any of them change, compare results carefully instead of treating the numbers as directly identical.
AgentBench shows the broader research challenge of evaluating agents across interactive environments. This project uses a smaller recorded suite. That keeps the engineering checks reproducible on a normal local machine.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-12-agent-eval.py
The script scores five recorded trajectories. The release rule needs task success ≥ 0.8, selection accuracy ≥ 0.8, and zero unauthorized executions.
- Run it unchanged. Read
task_success: 0.8,selection_accuracy: 0.8,average_success_steps: 2.5,unauthorized_executions: 0, andrelease_passed: True. - In the
approvetrace, change"authorized": Trueto"authorized": False, and keep"executed": True. Rerun. task_successandselection_accuracyare still0.8, butunauthorized_executions:becomes1andrelease_passed:becomesFalse.
One unauthorized action blocks release even when the averages look fine. Safety is a hard gate, not one more score to average.
Loading lab…
Quick Check
Explain it back
Design five fixed cases for your agent. Include at least one success, one blocker, one approval-required case, one repeated-action trap, and one memory case. State one threshold metric and one zero-tolerance constraint.
Key Takeaways
- Agent evaluation should inspect the trajectory as well as the outcome.
- Separate task, decision, boundedness, authority, memory, and efficiency metrics.
- Hard safety constraints should not disappear inside averages.
- Slice metrics reveal classes of tasks that need focused debugging.
- Versioned deterministic fixtures make engineering evidence reproducible.
Next Lesson
Next, integrate the Level into one bounded-agent trace with explicit planning, state, memory, selection, approval, stopping, and evaluation.
References
- Liu et al., AgentBench: Evaluating LLMs as Agents.
Completion is stored locally on this device.