Debugging and Error Handling
Identify what kind of failure a Claude application is having (HTTP error, stop_reason, streaming error, tool or integration defect, or model output), choose the right recovery strategy, and use traces to isolate whether the fault lies in the integration layer or in the model's output.
Key points
- 1
Error types: 400
invalid_request_error, 401authentication_error, 403permission_error, 404not_found_error, 413request_too_large, 429rate_limit_error, 500api_error, 504timeout_error, 529overloaded_error. Every error body hastype,messageand arequest_id. - 2
Retry transient errors (429, 500, 529 and connection errors) with exponential backoff and jitter, honoring
retry-after. The official SDKs already retry these twice by default (configurable withmax_retries). - 3
Do not retry request or credential errors unchanged: 400 (fix the request), 401/403 (fix the key or permissions), 413 (send less data, for example split input or use the Files API).
- 4
429 comes from organization-level limits (RPM, input and output tokens per minute), which replenish continuously. Smooth bursts by capping concurrency and ramping up gradually. A spend-cap 429 has no
retry-afterand keeps failing until access resumes, so alert on it rather than retrying. - 5
529
overloaded_erroris temporary, API-wide load. Back off, retry, and degrade gracefully. Rotating keys or editing the request will not help. - 6
With streaming, a failure can arrive as an SSE
errorevent (for exampleoverloaded_error) after HTTP 200. Treat it as a failed request and never persist the partial text as a complete answer. - 7
stop_reasonexplains a successful response:end_turn(done),max_tokens(truncated: raise the limit or continue),stop_sequence,tool_use(run the tool and call again),pause_turn(server-tool loop paused: send the content back),refusal(declined; HTTP 200, not an error; seestop_details), andmodel_context_window_exceeded(treat as truncated). - 8
Check
stop_reasonbefore parsing. JSON parse failures paired withmax_tokensare truncation, not formatting, and a blank UI reply withtool_usemeans the app never completed the tool loop. - 9
Empty
end_turnresponses right after tool results are often caused by the harness adding text immediately after thetool_resultblocks. Send only the tool results. - 10
Trace analysis: log each request's full prompt, tool inputs and results,
stop_reason, usage andrequest_id. Agents are non-deterministic, so diagnose from production traces rather than hoping to reproduce locally. - 11
Isolate the origin by walking the trace to the first wrong step. If the tool input is correct but the tool result is wrong, the fault is in the integration layer (wrapper mapping, truncation, auth). If a malformed input is forwarded faithfully, it is a model-output problem, fixed with clearer tool definitions and validation.
- 12
Return tool failures as
is_error: truewith a specific cause and remedy. Unhelpful error strings cause repeated identical calls, so pair this with turn caps in the harness. - 13
Use the SDKs' typed exception classes rather than string-matching messages, and include the
request_idwhen escalating to support.
Test yourself on Debugging and Error Handling
Ten questions, with the answer and explanation after each one.