Study notes · 2.3% of the exam

Support debugging and operational issue resolution

Diagnose and resolve Claude Code problems methodically: inspect what actually loaded, read permission denials and hook output, fix MCP connection failures, recover from context overflow, respond correctly to API errors, and escalate fleet-level issues using observability data rather than guesswork.

Key points

  1. 1

    Start by seeing what loaded: /context shows everything in the context window (memory files, skills, MCP tools, subagents); /memory, /skills, /hooks, /mcp and /permissions drill into each; /status shows active settings sources including whether managed settings apply; /doctor (or claude doctor from the shell) checks the installation and invalid settings files and proposes fixes; /debug [issue] turns on debug logging and asks Claude to diagnose from it.

  2. 2

    Reproduce with logs and transcripts. claude --debug (or --debug-file <path>) records hook evaluation, MCP server lifecycle and stderr, and telemetry errors; debug logs live under ~/.claude/debug/. Session transcripts are JSONL files at ~/.claude/projects/<project>/<session-id>.jsonl; resume one with claude --resume <session-id> or a transcript path, and export a readable copy with /export. Hooks and status-line scripts receive transcript_path, so a SessionEnd hook can archive transcripts. The transcript format is internal and changes between versions, so do not build reporting on parsing it.

  3. 3

    Isolate the cause with a clean configuration: claude --safe-mode disables CLAUDE.md, skills, plugins, hooks, MCP servers and custom agents (managed policy still applies); pointing CLAUDE_CONFIG_DIR at an empty directory from a folder with no .claude, .mcp.json or CLAUDE.md gives a session that loads nothing. If the problem disappears, reintroduce pieces one at a time.

  4. 4

    Settings that seem ignored are usually overridden by a higher level (managed, --settings, local over project over user), set in a file that cannot set them (for example auto or bypassPermissions in project settings, telemetry variables in a repository env block), placed in ~/.claude.json instead of ~/.claude/settings.json, or in a file with invalid JSON that Claude Code skipped.

  5. 5

    Permission denials: rules are evaluated deny, then ask, then allow, so an allow never overrides a deny or an ask from any file, and "Yes, don't ask again" writes an allow to the local file that cannot outrank a project or managed ask rule. In headless runs, denied calls appear in the JSON result's permission_denials (or as permission_denied events in stream-json); fix them with a targeted --allowedTools entry or permissions.allow rule, not by telling the model it is allowed and not by blanket bypass on a runner that holds secrets. dontAsk denies anything that would prompt; nobody can answer a prompt in CI.

  6. 6

    Hooks that do not fire: confirm the hook appears in /hooks (hooks must be under the hooks key of a settings file; there is no standalone hooks file except in plugins). Common matcher mistakes are lowercase tool names (bash instead of Bash), an array instead of a |-separated string (an array makes Claude Code reject the whole settings file), and misspelled tool names. Project hooks run only in a trusted workspace. claude --debug shows which matchers were checked and the hook's exit code and output. To block, a hook exits 2 or returns JSON with permissionDecision: "deny"; exit 0 proceeds and other exit codes are non-blocking errors.

  7. 7

    MCP connection failures: /mcp shows each server's status (connected, needs authentication, failed to connect, pending approval, disabled, rejected) with tool counts and lets you approve, authenticate via OAuth, reconnect or toggle a server. Project .mcp.json servers need a one-time approval; a dismissed prompt leaves them pending. Relative paths in command or args break from other directories, so use absolute paths; .mcp.json must be at the repository root with a top-level mcpServers key, and settings.json does not read mcpServers. Missing environment variables should be set per server in the env entry. claude mcp list and claude mcp get <name> show configuration warnings and error codes; claude --debug (or --debug=mcp) captures the server's stderr. MCP_TIMEOUT governs startup and MCP_TOOL_TIMEOUT per-call timeouts; MAX_MCP_OUTPUT_TOKENS caps tool output.

  8. 8

    Context overflow: Context limit reached · /compact or /clear to continue and Prompt is too long mean the conversation exceeds the model's window. /compact replaces history with a summary and accepts focus instructions (/compact keep only the plan and the diff); /clear starts fresh (the old session can still be resumed); /context shows what is consuming space. If auto-compaction thrashes because a file or tool output immediately refills the window, read the file in smaller ranges, compact with a focus that drops the large output, move the work to a subagent, or clear. A single oversized exchange cannot be compacted. CLAUDE.md can carry standing compaction instructions.

  9. 9

    API errors: API Error: Repeated 529 Overloaded or "<model> is experiencing high load" means the provider is overloaded; Claude Code retries with backoff automatically, and the user options are to wait, reduce request frequency or switch models with /model. Request rejected (429) is your organization's rate limit: wait, space out requests, check TPM and RPM limits in the Console and raise them for known concurrency spikes such as large training sessions (limits are organization-level, so rotating keys does not help). 500 errors are retried with exponential backoff. model not found or "restricted by your organization's settings" means the model is unavailable on the provider or blocked by availableModels. In stream-json output, retries appear as system/api_retry events with attempt, retry_delay_ms and an error category such as rate_limit or overloaded. Usage-limit messages on subscription plans (session or weekly limits) are a different ceiling from API rate limits.

  10. 10

    Performance and stability: /compact regularly, restart between major tasks, add build directories to .gitignore, and use claude --safe-mode to check whether a plugin, MCP server or hook is the cause. Restarting does not lose the conversation; claude --continue resumes it. Search problems usually trace to the bundled ripgrep; install the system package and set USE_BUILTIN_RIPGREP=0.

  11. 11

    Escalate with data, not anecdotes. OpenTelemetry metrics (claude_code.cost.usage, claude_code.token.usage, claude_code.active_time.total) carry user.id, organization.id, session.id, model, query_source, mcp_server.name, skill.name and custom OTEL_RESOURCE_ATTRIBUTES; events include api_request, api_error, tool_decision, tool_result and user_prompt, correlated by prompt.id and request_id. Compare affected and unaffected groups over time to localize a cost or latency regression before changing models, settings or advice. Prompt text is excluded unless OTEL_LOG_USER_PROMPTS is enabled and is rarely needed.

  12. 12

    Traps: reinstalling or switching models before checking /mcp, /hooks or /status; prompting the model to "allow" a tool that the harness denied; using --dangerously-skip-permissions as a fix for denials on shared runners; rotating API keys to escape 429s; blaming context problems on the model rather than compacting or restructuring the work; and collecting conversation content when metrics already carry the needed dimensions.

Test yourself on Support debugging and operational issue resolution

Ten questions, with the answer and explanation after each one.