Support debugging and operational issue resolution
Diagnose and resolve Claude Code problems methodically: inspect what actually loaded, read permission denials and hook output, fix MCP connection failures, recover from context overflow, respond correctly to API errors, and escalate fleet-level issues using observability data rather than guesswork.
Key points
- 1
Start by seeing what loaded:
/contextshows everything in the context window (memory files, skills, MCP tools, subagents);/memory,/skills,/hooks,/mcpand/permissionsdrill into each;/statusshows active settings sources including whether managed settings apply;/doctor(orclaude doctorfrom the shell) checks the installation and invalid settings files and proposes fixes;/debug [issue]turns on debug logging and asks Claude to diagnose from it. - 2
Reproduce with logs and transcripts.
claude --debug(or--debug-file <path>) records hook evaluation, MCP server lifecycle and stderr, and telemetry errors; debug logs live under~/.claude/debug/. Session transcripts are JSONL files at~/.claude/projects/<project>/<session-id>.jsonl; resume one withclaude --resume <session-id>or a transcript path, and export a readable copy with/export. Hooks and status-line scripts receivetranscript_path, so aSessionEndhook can archive transcripts. The transcript format is internal and changes between versions, so do not build reporting on parsing it. - 3
Isolate the cause with a clean configuration:
claude --safe-modedisables CLAUDE.md, skills, plugins, hooks, MCP servers and custom agents (managed policy still applies); pointingCLAUDE_CONFIG_DIRat an empty directory from a folder with no.claude,.mcp.jsonorCLAUDE.mdgives a session that loads nothing. If the problem disappears, reintroduce pieces one at a time. - 4
Settings that seem ignored are usually overridden by a higher level (managed,
--settings, local over project over user), set in a file that cannot set them (for exampleautoorbypassPermissionsin project settings, telemetry variables in a repositoryenvblock), placed in~/.claude.jsoninstead of~/.claude/settings.json, or in a file with invalid JSON that Claude Code skipped. - 5
Permission denials: rules are evaluated deny, then ask, then allow, so an allow never overrides a deny or an ask from any file, and "Yes, don't ask again" writes an allow to the local file that cannot outrank a project or managed ask rule. In headless runs, denied calls appear in the JSON result's
permission_denials(or aspermission_deniedevents instream-json); fix them with a targeted--allowedToolsentry orpermissions.allowrule, not by telling the model it is allowed and not by blanket bypass on a runner that holds secrets.dontAskdenies anything that would prompt; nobody can answer a prompt in CI. - 6
Hooks that do not fire: confirm the hook appears in
/hooks(hooks must be under thehookskey of a settings file; there is no standalone hooks file except in plugins). Common matcher mistakes are lowercase tool names (bashinstead ofBash), an array instead of a|-separated string (an array makes Claude Code reject the whole settings file), and misspelled tool names. Project hooks run only in a trusted workspace.claude --debugshows which matchers were checked and the hook's exit code and output. To block, a hook exits 2 or returns JSON withpermissionDecision: "deny"; exit 0 proceeds and other exit codes are non-blocking errors. - 7
MCP connection failures:
/mcpshows each server's status (connected, needs authentication, failed to connect, pending approval, disabled, rejected) with tool counts and lets you approve, authenticate via OAuth, reconnect or toggle a server. Project.mcp.jsonservers need a one-time approval; a dismissed prompt leaves them pending. Relative paths incommandorargsbreak from other directories, so use absolute paths;.mcp.jsonmust be at the repository root with a top-levelmcpServerskey, andsettings.jsondoes not readmcpServers. Missing environment variables should be set per server in theenventry.claude mcp listandclaude mcp get <name>show configuration warnings and error codes;claude --debug(or--debug=mcp) captures the server's stderr.MCP_TIMEOUTgoverns startup andMCP_TOOL_TIMEOUTper-call timeouts;MAX_MCP_OUTPUT_TOKENScaps tool output. - 8
Context overflow:
Context limit reached · /compact or /clear to continueandPrompt is too longmean the conversation exceeds the model's window./compactreplaces history with a summary and accepts focus instructions (/compact keep only the plan and the diff);/clearstarts fresh (the old session can still be resumed);/contextshows what is consuming space. If auto-compaction thrashes because a file or tool output immediately refills the window, read the file in smaller ranges, compact with a focus that drops the large output, move the work to a subagent, or clear. A single oversized exchange cannot be compacted. CLAUDE.md can carry standing compaction instructions. - 9
API errors:
API Error: Repeated 529 Overloadedor "<model> is experiencing high load" means the provider is overloaded; Claude Code retries with backoff automatically, and the user options are to wait, reduce request frequency or switch models with/model.Request rejected (429)is your organization's rate limit: wait, space out requests, check TPM and RPM limits in the Console and raise them for known concurrency spikes such as large training sessions (limits are organization-level, so rotating keys does not help). 500 errors are retried with exponential backoff.model not foundor "restricted by your organization's settings" means the model is unavailable on the provider or blocked byavailableModels. Instream-jsonoutput, retries appear assystem/api_retryevents withattempt,retry_delay_msand an error category such asrate_limitoroverloaded. Usage-limit messages on subscription plans (session or weekly limits) are a different ceiling from API rate limits. - 10
Performance and stability:
/compactregularly, restart between major tasks, add build directories to.gitignore, and useclaude --safe-modeto check whether a plugin, MCP server or hook is the cause. Restarting does not lose the conversation;claude --continueresumes it. Search problems usually trace to the bundled ripgrep; install the system package and setUSE_BUILTIN_RIPGREP=0. - 11
Escalate with data, not anecdotes. OpenTelemetry metrics (
claude_code.cost.usage,claude_code.token.usage,claude_code.active_time.total) carryuser.id,organization.id,session.id,model,query_source,mcp_server.name,skill.nameand customOTEL_RESOURCE_ATTRIBUTES; events includeapi_request,api_error,tool_decision,tool_resultanduser_prompt, correlated byprompt.idandrequest_id. Compare affected and unaffected groups over time to localize a cost or latency regression before changing models, settings or advice. Prompt text is excluded unlessOTEL_LOG_USER_PROMPTSis enabled and is rarely needed. - 12
Traps: reinstalling or switching models before checking
/mcp,/hooksor/status; prompting the model to "allow" a tool that the harness denied; using--dangerously-skip-permissionsas a fix for denials on shared runners; rotating API keys to escape 429s; blaming context problems on the model rather than compacting or restructuring the work; and collecting conversation content when metrics already carry the needed dimensions.
Read the source
Test yourself on Support debugging and operational issue resolution
Ten questions, with the answer and explanation after each one.