Analyze observability challenges and select monitoring strategies at scale
Design monitoring for LLM systems at scale: correlate every hop with a trace id, record structured metadata for every request, sample and grade transcripts, alert on quality drift as well as errors, and keep PII out of shared logs.
Key points
- 1
Mint a trace or correlation id at the top-level request and propagate it through every model call, subagent, tool call and MCP server, so any final answer can be walked back to the exact hop and tool result that produced it. More verbose logging on one component does not substitute for correlation across hops.
- 2
Log structured metadata per request: prompt or system-prompt version, model,
effort/thinking configuration, input, output and cache-read tokens, latency, tool calls with outcomes, retrieved document or chunk ids, and stop reason. This is what lets a regression be tied to a release, index refresh or configuration change. - 3
Infrastructure health is not product health: an LLM system can be error-free and fast while its answers degrade. Track quality metrics (grader scores on sampled transcripts, tool-call success rate, escalation or hand-off rate, task completion) and alert on drift from a baseline, not only on 5xx and latency.
- 4
Sample transcripts for human and model-graded review, stratified by intent, region and outcome, and oversample failures, escalations and low-confidence cases. A uniform random sample under-represents rare failure classes; making it larger does not fix that.
- 5
Slice quality metrics along the dimensions where regressions occur (region, tenant, intent, prompt version); a single global average hides a failure confined to one slice.
- 6
At scale, keep lightweight PII-redacted metadata for 100% of traffic and reserve full transcripts for the stratified sample; storing every raw transcript is usually unaffordable and a privacy liability.
- 7
PII-aware logging: redact or tokenize sensitive fields before data enters shared observability tooling, keep any raw copies in a segregated access-controlled store with audit trails and short retention, and remember that encryption at rest and base64 do not stop authorized readers from seeing PII.
- 8
Claude Code exports OpenTelemetry metrics (session count, lines of code, commits, PRs, cost, token usage, code-edit permission decisions, active time) and events (user prompt, tool result, API request, API error, tool decision) with user, session and organization attributes.
- 9
Enable and route Claude Code telemetry fleet-wide through managed settings (
CLAUDE_CODE_ENABLE_TELEMETRY,OTEL_METRICS_EXPORTER,OTEL_LOGS_EXPORTER,OTEL_EXPORTER_OTLP_ENDPOINT); a repository's.claude/settings.jsoncannot turn telemetry on or redirect it. - 10
Prompt text, tool inputs, tool content and assistant responses are redacted from Claude Code telemetry by default;
OTEL_LOG_USER_PROMPTS,OTEL_LOG_TOOL_DETAILS,OTEL_LOG_TOOL_CONTENTandOTEL_LOG_ASSISTANT_RESPONSESopt in. Cardinality flags (OTEL_METRICS_INCLUDE_SESSION_IDand similar) trade granularity for storage cost. - 11
Agent SDK hooks are an audit surface: a
PreToolUseorPostToolUsehook without a matcher fires on every tool call and can write structured audit records; hook inputs carry the tool-use id to correlate the pre and post events andagent_idfor subagents. - 12
Monitor the usage fields the API returns (
input_tokens,output_tokens,cache_read_input_tokens, thinking token details) to verify caching is working and to attribute cost; a zero cache read across repeated requests is a silent-invalidation signal. - 13
Distractors the exam likes: alert only on errors, log everything including PHI to a shared cluster, rely solely on thumbs-down feedback, export logs weekly for analysts, or switch models instead of measuring.
Read the source
Test yourself on Analyze observability challenges and select monitoring strategies at scale
Ten questions, with the answer and explanation after each one.