Agent Failure Modes
Goal
Classify common agent failures by the earliest observable boundary—state, selection, planning, memory, authority, stopping, or evaluation—so fixes target the actual cause instead of adding vague prompt instructions.
Agent failures often look similar at the end. The final answer is wrong, a task is unfinished, or a side effect should not have happened.
The useful debugging question is: where did the trajectory first become wrong?
If you find the earliest broken boundary, the fix becomes more specific.
Failure 1: goal drift
The agent begins with one goal but gradually optimizes for a different one.
Example:
goal: answer whether order 4172 can be refunded
later behavior: repeatedly search for discount offers
The tools may be working perfectly. The failure is that the active subgoal no longer supports the original goal.
A fix could compare proposed subgoals and actions with the stored goal before allowing continuation.
Failure 2: stale or corrupted state
Suppose a fresh lookup says status=delivered, but working state still says shipping. Later decisions can be wrong even if the model reasons correctly from the state it sees.
This is a state-update failure. Prompting the model to be “more careful” does not repair an application bug that failed to replace stale data.
Failure 3: bad tool selection
The current subgoal needs policy evidence, but the agent chooses a write tool or an irrelevant lookup.
The tool call can be perfectly valid and authorized while still being the wrong next action. Evaluate selection separately from argument validation.
Failure 4: memory contamination
A retrieved memory from a different user, task, or outdated period enters working state and influences the next action.
This is why long-term memory needs scope, provenance, freshness, and retrieval filters. Stored text can also contain untrusted instructions; memory persistence does not turn those instructions into policy.
Failure 5: authority failure
A model proposes a side effect and application code executes it without the required permission or approval.
This is more serious than a poor answer because the system crossed an external action boundary incorrectly.
The fix belongs in authorization, approval binding, or tool exposure—not merely in a warning sentence inside the prompt.
Failure 6: loop and stopping failure
The agent repeats actions, keeps revising its plan, or continues after the goal is already satisfied.
A max-step budget catches some cases, but earlier signals such as repeated normalized actions, no new evidence, and completed required subgoals can stop the run sooner.
Premature stopping is the opposite problem: the agent ends before required evidence or actions are complete.
Failure 7: evaluation blindness
A system may have a high average task-success rate while hiding rare unsafe trajectories or one weak task slice.
Evaluation blindness happens when the metrics do not expose the failure you care about.
For bounded agents, safety and boundedness constraints should be reported separately from average task scores. One unauthorized side effect should not disappear inside a 95% success number.
Research benchmarks such as AgentBench use interactive environments to study multi-turn agent behavior. Your local project is smaller, but the lesson is transferable: inspect trajectories and failure types, not only final text.
Predict
Run the local Lab
Run:
python labs/notebooks/level-11/l11-11-failure-modes.py
The script checks each trace against six failure categories in order, from goal_drift to stopping, and reports the first one that fails.
- Run it unchanged. Read
stale => stale_state,selection => tool_selection,authority => authority, andclean => none. - Look at the
staletrace in the script. It has two defects:"state_current": Falseand"selection_correct": False. Only the earlier one is reported. - Fix the stale state by changing
"state_current": Falseto"state_current": True. Rerun. - Now
stale => tool_selection. Fixing the earliest failure revealed a second, later defect.
That is why debugging goes in order: a later check can only be trusted once the earlier ones pass.
Loading lab…
Quick Check
Explain it back
Choose one bad agent run and draw a seven-column table: goal, observation, state, decision, authorization, action, stop/evaluation. Mark the first row where reality and the recorded system behavior diverge. Explain why a later prompt change would or would not fix that boundary.
Key Takeaways
- Debug agent trajectories from the earliest broken boundary.
- Goal drift, stale state, bad selection, memory contamination, authority errors, and loop failures need different fixes.
- Correct tools can still be chosen at the wrong time.
- Prompt wording cannot replace missing deterministic authority controls.
- Evaluation should expose safety and boundedness failures separately from averages.
Next Lesson
Next, turn these failure categories into a small evaluation suite that measures the whole trajectory, not only the final answer.
References
- Liu et al., AgentBench: Evaluating LLMs as Agents.
- OWASP GenAI Security Project, Top 10 for LLM Applications.
Completion is stored locally on this device.