Skip to main content
L12.4

Tracing Agent Decisions

Goal

Create structured traces that connect a task, controller decisions, tool calls, retries, and terminal outcomes without recording hidden reasoning.

Trace the journey, not hidden reasoning​

A package-tracking page is useful because it does not only say "delivery failed." It shows a sequence: package accepted, sorting center reached, truck departed, delivery attempted, and so on. Each event has a time and identity, so you can locate the stage where the problem appeared.

A software trace serves a similar purpose for one agent run. The trace links the whole journey, while smaller spans record timed pieces of work inside it. Different identifiers answer different questions: a task ID names the durable job, a trace ID links one observed run, a span ID names one operation inside that run, and an operation ID can connect retries of the same side effect. You do not need to memorize the names first; ask what relationship each ID lets you reconstruct.

A trace is a set of related records for one request or task. A span represents one timed unit of work inside that trace. For an agent harness, useful spans might include task execution, model request, policy check, tool call, memory retrieval, approval wait, and recovery step. Parent-child relationships show which operation caused another operation.

Use structured fields and distinct IDs​

Structured fields make traces searchable. Instead of a log line that says 'tool failed again,' record fields such as task_id, action, tool_name, attempt, error_type, duration_ms, and terminal_reason. Consistent fields let you compare many runs and answer questions such as which tool creates the most retries.

{"task_id":"T-42","trace_id":"tr-9","span":"tool_call","attempt":2,"duration_ms":140}

Tracing should record application-visible decisions, not hidden chain-of-thought. The harness can record that the model proposed search_catalog, that policy allowed it, that the tool returned zero items, and that the controller chose a second search. That is enough to debug behavior without storing private internal reasoning text.

Identifiers matter. A trace ID links the whole run. A span ID names one operation. A task ID links events across restarts or multiple traces. An operation ID links retries of the same side effect. These identifiers answer different questions, so collapsing them into one field makes recovery and analysis harder.

Timing and privacy are part of the trace​

Timing is also evidence. End-to-end latency can be broken into queue wait, model latency, tool latency, approval wait, retry delay, and local processing. If the total run takes thirty seconds, tracing lets you see whether the time was spent waiting for a user, calling a slow service, or repeatedly retrying one tool.

Be deliberate about what you record. Tool arguments may contain secrets or personal data. Traces should apply redaction and retention rules. Observability is not an excuse to copy every prompt, credential, or user field into a permanent log.

OpenTelemetry provides common concepts and data models for traces, metrics, and logs. The exact SDK may change, but the useful mental model is stable: emit structured telemetry at meaningful boundaries and preserve relationships between operations.

Record evidence that answers questions​

A trace is valuable only if it helps answer operational questions. Before adding a field, ask what decision it enables. Before omitting a field, ask whether you could still distinguish a model-selection problem from a policy rejection, tool failure, or retry bug.

Predict

Which trace record best helps distinguish a policy rejection from a tool failure?

Run the local Lab​

Run:

python labs/notebooks/level-12/l12-04-tracing.py

The script prints one trace: a parent task span and three child spans, model, policy, and tool. The children run one after another, so the task span lasts as long as all three together.

  1. Run it unchanged. Check that every child has 'parent': 'task', and that the tool span also records 'attempt': 1.
  2. Read critical path: model 70 + policy 10 + tool 40 = 120 ms.
  3. Open the script and change TOOL_LATENCY_MS = 40 to TOOL_LATENCY_MS = 140. Rerun.
  4. Now the tool span shows 'duration_ms': 140, and the task span becomes 220. The model and policy spans still show 70 and 10.

Because each step has its own span, you can see that the slowdown came from the tool and not the model. One total time number could not tell you that.

Loading lab…

Quick Check

1. Which trace record is sufficient for debugging agent control flow?
2. Why keep task ID and operation ID separate?
3. What is a key privacy rule for tracing?

0 of 3 questions answered.

Explain it back​

Sketch a trace for a task that waits in a queue, calls a model, passes policy, retries one tool once, and finishes. Name the IDs and timings you would want to query later.

Key Takeaways

  • Traces connect observable events across an agent run.
  • Spans represent timed units of work and preserve relationships.
  • Use stable identifiers for tasks, traces, spans, and retryable operations.
  • Do not store hidden reasoning as a debugging requirement.
  • Telemetry needs privacy and retention rules.

Next Lesson

Next, L12.5 — Sandboxing and Permissions separates authorization from the technical isolation that limits what a process can reach.

References

Lesson actions

Completion is stored locally on this device.

View progress