Study notes · 2.6% of the exam

Monitor system performance using logging and observability tools

Monitor a Claude system in production with structured logs, traces, token and latency dashboards, sampled quality grading and drift alerts, while keeping sensitive content inside the trust boundary.

Key points

  1. 1

    Log a structured record per request: request and conversation ids, prompt and model version, index version, input/output/cache token counts, latency, stop_reason, tool calls with inputs and outputs, and an error code. Structured fields feed dashboards and alerts; free text and billing totals do not.

  2. 2

    Agents need traces, not just logs: a span per model call and tool call correlated by a run id, with the retrieved data and its source identifiers at each step, so a failing run can be reconstructed and replayed.

  3. 3

    Dashboards: tokens and cost per task by model and feature, cache hit rate, p50/p95/p99 latency, error and retry rates, stop_reason distribution (truncation), tool-call and escalation rates, all sliced by version so regressions can be attributed to a deployment.

  4. 4

    Quality in production has no labels, so sample: grade a fixed percentage of traffic daily with a rubric-based LLM judge, review a stratified subset with humans weekly to keep the judge calibrated, and track the judged failure and groundedness rates over time.

  5. 5

    Alert on rates over rolling windows compared with a baseline, with minimum sample sizes, and annotate with change events (deployments, model updates, index refreshes). Per-event alerts on a probabilistic system are noise; monthly reviews are too slow.

  6. 6

    Watch for input drift (new topics, new languages, new document types) as well as system drift; refresh the eval set from recent production traffic so offline and online views stay aligned.

  7. 7

    Claude Code exports OpenTelemetry metrics and events: claude_code.session.count, claude_code.token.usage (by type: input, output, cacheRead, cacheCreation; and by model), claude_code.cost.usage, claude_code.lines_of_code.count, claude_code.commit.count, claude_code.pull_request.count, claude_code.code_edit_tool.decision (accept/reject by tool and language) and claude_code.active_time.total, plus events for prompts, tool results, API requests and errors.

  8. 8

    Enable it with CLAUDE_CODE_ENABLE_TELEMETRY=1 and OTEL_METRICS_EXPORTER / OTEL_LOGS_EXPORTER (otlp, prometheus, console) and the OTLP endpoint. Administrators set these centrally in managed settings; Claude Code ignores telemetry variables placed in a repository's .claude/settings.json or .claude/settings.local.json.

  9. 9

    Content is off by default: user prompt text, assistant responses and tool details are exported only if OTEL_LOG_USER_PROMPTS, OTEL_LOG_ASSISTANT_RESPONSES or OTEL_LOG_TOOL_DETAILS are enabled. Cardinality flags (session id, account uuid, repository) control attribute volume.

  10. 10

    In regulated settings, export metrics, traces and metadata to the observability vendor and keep raw prompts and responses in an in-house, encrypted, access-audited store linked by request id. Model-performed redaction is not a compliance control, and vendor role permissions do not undo an export.

  11. 11

    Correlate everything: a request id in the application, prompt.id and request_id in Claude Code telemetry, and version tags on prompts, models and indexes make the difference between diagnosis and guesswork.

Test yourself on Monitor system performance using logging and observability tools

Ten questions, with the answer and explanation after each one.