Skip to main content
L14.10

Observability for AI Systems

Goal

Use traces, metrics, and structured logs to explain request behavior, resource pressure, and deployment changes without collecting unnecessary sensitive data.

Start from an operational question​

Suppose one request takes 12 seconds. "The service was slow" is only a symptom. You still need to know whether the request waited in a queue, the model itself ran slowly, the GPU was under memory pressure, or a new deployment changed behavior.

Three kinds of evidence help answer different parts of that question. A trace is like following one request's journey step by step. A metric is like a scoreboard that summarizes many requests, such as p95 latency or error rate. A structured log is a timestamped event record with named fields, such as "deployment v18 rejected request class chat because max_tokens exceeded the limit."

Observability is the practice of collecting and connecting enough of this evidence to explain system behavior without guessing. It is not "save everything." Good observability records the information needed to investigate while avoiding unnecessary sensitive content.

Use traces, metrics, and logs for different jobs​

For a trace, useful spans might cover API validation, queue wait, prefill, decode, tool calls, and response streaming. Their timing shows which step was on the slow path.

For metrics, useful examples include request rate, queue depth, time to first token, latency percentiles, token counts, error rate, GPU memory pressure, cache-admission failures, and replica readiness.

For a structured log, stable fields can record that deployment v17 rejected a request because max_tokens was invalid. Stable fields are easier to search and aggregate than one long free-text message.

trace: request r-17 → queue 80 ms → decode 900 ms
metric: p95_latency_ms = 1480
log: {request_id:r-17, deployment:v18, event:decode_slow}

Correlate signals with stable IDs​

OpenTelemetry defines common concepts for traces, metrics, and logs. The key idea here is correlation: IDs and version fields should let you move from a dashboard spike to the requests and deployment that caused it.

Protect telemetry data and metric cardinality​

Telemetry can contain sensitive data. Prompts, generated text, and tool results may include private information. Record the metadata needed for diagnosis by default. Store content only under an explicit redaction and retention policy.

Metrics also have a cardinality limit. Cardinality means how many distinct label combinations a metric creates. Putting raw request IDs or prompts in metric labels can create huge numbers of series. Keep those high-uniqueness values in traces or logs instead.

Alerts should point to conditions someone can act on. A one-second spike may not matter. Sustained errors, lost readiness, memory exhaustion, or an SLO breach are more useful when each alert has a documented response.

Close the loop after a change​

Observability closes the loop after a change. If a rollout changes latency, throughput, or errors, the telemetry should help separate queue delay, model runtime, resource pressure, and client behavior.

For example, a latency alert might show only that p95 became worse after deployment v18. A correlated trace can then show that queue wait stayed normal while decode time increased. Resource metrics may show higher GPU memory pressure on the same replicas, and structured logs can tie the change to the new model/runtime version. Each signal answers a different part of the incident, but shared request and deployment identity lets them form one explanation.

Predict

A dashboard shows p95 latency rising after a deployment. What evidence best helps locate the cause?

Run the Docker-environment Lab preflight​

This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.

Run:

python3 labs/notebooks/level-14/l14-10-observability.py

The Lab keeps high-cardinality request identity in logs/traces rather than aggregate metric labels.

  1. Run it unchanged. candidate_labels contains only deployment and status, so aggregate metrics are produced.
  2. Before editing, predict why adding one label whose value is different for nearly every request would create an unbounded metric series.
  3. Change only candidate_labels = {"deployment", "status"} to candidate_labels = {"deployment", "status", "request_id"}, then rerun.
  4. Confirm the guard rejects request_id and tells you to keep it in logs and traces instead.

Loading lab…

Quick Check

1. What is a trace best suited for?
2. Why avoid raw request IDs as metric labels?
3. What should observability content collection consider?

0 of 3 questions answered.

Explain it back​

Create one trace, five aggregate metrics, and three structured log fields for a generation request. State which user-content fields you would avoid recording by default.

Key Takeaways

  • Traces explain individual request paths.
  • Metrics summarize service behavior across requests.
  • Logs record discrete operational events with stable fields.
  • Correlate telemetry with deployment/model identity.
  • Protect privacy and control metric cardinality.

Next Lesson

Next, L14.11 — Rollouts, Canaries, and Rollbacks turns telemetry into controlled release decisions.

References

Lesson actions

Completion is stored locally on this device.

View progress