Monitor system performance using logging and observability tools
Monitor a Claude system in production with structured logs, traces, token and latency dashboards, sampled quality grading and drift alerts, while keeping sensitive content inside the trust boundary.
Key points
- 1
Log a structured record per request: request and conversation ids, prompt and model version, index version, input/output/cache token counts, latency,
stop_reason, tool calls with inputs and outputs, and an error code. Structured fields feed dashboards and alerts; free text and billing totals do not. - 2
Agents need traces, not just logs: a span per model call and tool call correlated by a run id, with the retrieved data and its source identifiers at each step, so a failing run can be reconstructed and replayed.
- 3
Dashboards: tokens and cost per task by model and feature, cache hit rate, p50/p95/p99 latency, error and retry rates,
stop_reasondistribution (truncation), tool-call and escalation rates, all sliced by version so regressions can be attributed to a deployment. - 4
Quality in production has no labels, so sample: grade a fixed percentage of traffic daily with a rubric-based LLM judge, review a stratified subset with humans weekly to keep the judge calibrated, and track the judged failure and groundedness rates over time.
- 5
Alert on rates over rolling windows compared with a baseline, with minimum sample sizes, and annotate with change events (deployments, model updates, index refreshes). Per-event alerts on a probabilistic system are noise; monthly reviews are too slow.
- 6
Watch for input drift (new topics, new languages, new document types) as well as system drift; refresh the eval set from recent production traffic so offline and online views stay aligned.
- 7
Claude Code exports OpenTelemetry metrics and events:
claude_code.session.count,claude_code.token.usage(by type: input, output, cacheRead, cacheCreation; and by model),claude_code.cost.usage,claude_code.lines_of_code.count,claude_code.commit.count,claude_code.pull_request.count,claude_code.code_edit_tool.decision(accept/reject by tool and language) andclaude_code.active_time.total, plus events for prompts, tool results, API requests and errors. - 8
Enable it with
CLAUDE_CODE_ENABLE_TELEMETRY=1andOTEL_METRICS_EXPORTER/OTEL_LOGS_EXPORTER(otlp, prometheus, console) and the OTLP endpoint. Administrators set these centrally in managed settings; Claude Code ignores telemetry variables placed in a repository's.claude/settings.jsonor.claude/settings.local.json. - 9
Content is off by default: user prompt text, assistant responses and tool details are exported only if
OTEL_LOG_USER_PROMPTS,OTEL_LOG_ASSISTANT_RESPONSESorOTEL_LOG_TOOL_DETAILSare enabled. Cardinality flags (session id, account uuid, repository) control attribute volume. - 10
In regulated settings, export metrics, traces and metadata to the observability vendor and keep raw prompts and responses in an in-house, encrypted, access-audited store linked by request id. Model-performed redaction is not a compliance control, and vendor role permissions do not undo an export.
- 11
Correlate everything: a request id in the application,
prompt.idandrequest_idin Claude Code telemetry, and version tags on prompts, models and indexes make the difference between diagnosis and guesswork.
Read the source
Test yourself on Monitor system performance using logging and observability tools
Ten questions, with the answer and explanation after each one.