Observability for AI Systems
Goal
Use traces, metrics, and structured logs to explain request behavior, resource pressure, and deployment changes without collecting unnecessary sensitive data.
Start from an operational question
Suppose one request takes 12 seconds. "The service was slow" is only a symptom. You still need to know whether the request waited in a queue, the model itself ran slowly, the GPU was under memory pressure, or a new deployment changed behavior.
Three kinds of evidence help answer different parts of that question. A trace is like following one request's journey step by step. A metric is like a scoreboard that summarizes many requests, such as p95 latency or error rate. A structured log is a timestamped event record with named fields, such as "deployment v18 rejected request class chat because max_tokens exceeded the limit."
Observability is the practice of collecting and connecting enough of this evidence to explain system behavior without guessing. It is not "save everything." Good observability records the information needed to investigate while avoiding unnecessary sensitive content.
Use traces, metrics, and logs for different jobs
For a trace, useful spans might cover API validation, queue wait, prefill, decode, tool calls, and response streaming. Their timing shows which step was on the slow path.
For metrics, useful examples include request rate, queue depth, time to first token, latency percentiles, token counts, error rate, GPU memory pressure, cache-admission failures, and replica readiness.
For a structured log, stable fields can record that deployment v17 rejected a request because max_tokens was invalid. Stable fields are easier to search and aggregate than one long free-text message.
trace: request r-17 → queue 80 ms → decode 900 ms
metric: p95_latency_ms = 1480
log: {request_id:r-17, deployment:v18, event:decode_slow}
Correlate signals with stable IDs
OpenTelemetry defines common concepts for traces, metrics, and logs. The key idea here is correlation: IDs and version fields should let you move from a dashboard spike to the requests and deployment that caused it.
Protect telemetry data and metric cardinality
Telemetry can contain sensitive data. Prompts, generated text, and tool results may include private information. Record the metadata needed for diagnosis by default. Store content only under an explicit redaction and retention policy.
Metrics also have a cardinality limit. Cardinality means how many distinct label combinations a metric creates. Putting raw request IDs or prompts in metric labels can create huge numbers of series. Keep those high-uniqueness values in traces or logs instead.
Alerts should point to conditions someone can act on. A one-second spike may not matter. Sustained errors, lost readiness, memory exhaustion, or an SLO breach are more useful when each alert has a documented response.
Close the loop after a change
Observability closes the loop after a change. If a rollout changes latency, throughput, or errors, the telemetry should help separate queue delay, model runtime, resource pressure, and client behavior.
For example, a latency alert might show only that p95 became worse after deployment v18. A correlated trace can then show that queue wait stayed normal while decode time increased. Resource metrics may show higher GPU memory pressure on the same replicas, and structured logs can tie the change to the new model/runtime version. Each signal answers a different part of the incident, but shared request and deployment identity lets them form one explanation.
Predict
Run the Docker-environment Lab preflight
This activity is registered for the Docker-oriented production environment used in the second half of Level 14. Start with the deterministic Python preflight so service-logic failures remain distinguishable from container, GPU, cluster, or credential setup problems.
Run:
python3 labs/notebooks/level-14/l14-10-observability.py
The Lab keeps high-cardinality request identity in logs/traces rather than aggregate metric labels.
- Run it unchanged.
candidate_labelscontains onlydeploymentandstatus, so aggregate metrics are produced. - Before editing, predict why adding one label whose value is different for nearly every request would create an unbounded metric series.
- Change only
candidate_labels = {"deployment", "status"}tocandidate_labels = {"deployment", "status", "request_id"}, then rerun. - Confirm the guard rejects
request_idand tells you to keep it in logs and traces instead.
Loading lab…
Quick Check
Explain it back
Create one trace, five aggregate metrics, and three structured log fields for a generation request. State which user-content fields you would avoid recording by default.
Key Takeaways
- Traces explain individual request paths.
- Metrics summarize service behavior across requests.
- Logs record discrete operational events with stable fields.
- Correlate telemetry with deployment/model identity.
- Protect privacy and control metric cardinality.
Next Lesson
Next, L14.11 — Rollouts, Canaries, and Rollbacks turns telemetry into controlled release decisions.
References
- OpenTelemetry, Specifications.
Completion is stored locally on this device.