Claude API Mechanics
Know how the Messages API behaves in practice: request and response structure, the tool-use loop, streaming, vision and PDF input, thinking, prompt caching, cloud-platform access, data access patterns, and when to choose the Batches API over realtime calls.
Key points
- 1
The Messages API is stateless: every request must carry the full conversation in
messages(alternatinguser/assistantturns), with instructions in the top-levelsystemparameter. There is no server-side conversation ID. - 2
Always branch on
stop_reason:end_turn(finished),max_tokens(truncated, so raise the limit or continue),stop_sequence(one of yourstop_sequencesfired; thestop_sequencefield says which),tool_use(run your tools),pause_turn(a server-tool turn was paused, so send the content back unchanged to resume),refusal. Truncation still returns HTTP 200. - 3
Tool loop: when
stop_reasonistool_use, execute everytool_useblock and reply with one user message holding all matchingtool_resultblocks (bytool_use_id), placed before any text. Report failures withis_error: true; never drop them. - 4
Server tools such as web search run on Anthropic's side, and their results arrive in the same response. Only client tools need your code to execute them and return
tool_result. - 5
Streaming uses SSE events:
message_start, then per blockcontent_block_start/content_block_delta/content_block_stop, thenmessage_deltaandmessage_stop, withpingevents anywhere and possibleerrorevents (e.g.overloaded_error). Tool arguments stream asinput_json_deltafragments; join them and parse atcontent_block_stop. Stream long or large-max_tokensrequests to avoid HTTP timeouts. - 6
Images and PDFs are sent as
image/documentcontent blocks (base64, URL or Files APIfile_idsource), ideally placed before the question. PDF pages are read as both text and images. Requests with more than 20 images get a stricter per-image size limit, and the standard request size limit is 32 MB, so downscale images and prune old ones in image-heavy agents. - 7
Thinking: current models use adaptive thinking (
thinking: {type: "adaptive"}), with depth tuned byoutput_config.effort. When returning tool results, pass the assistant's thinking blocks back complete and unmodified; simply appendresponse.content. - 8
Prompt caching is a prefix match over
tools→system→messagesup to acache_controlbreakpoint (at most 4 breakpoints). Any byte change earlier in the prefix, such as a timestamp in the system prompt, invalidates everything after it. Default TTL is 5 minutes (1 hour is available at extra cost); cache reads are about 10% of the base input price and 5-minute writes cost 1.25×. Checkusage.cache_read_input_tokens. - 9
Token counting (
count_tokens) takes the same system, tools and messages as a real request and returns an estimate for that model. It is free but rate-limited. Tokenizers differ between models and vendors, so count against the model you will call. - 10
Citations: set
citations: {enabled: true}on all or none of the document blocks. PDFs cite page ranges, plain text cites character ranges, andcited_textis guaranteed to come from your document. Citations cannot be combined with structured outputs (output_config.format); doing so returns a 400. - 11
For RAG with your own sources, return
search_resultblocks (with requiredsourceandtitle, plus text content) from a tool or in a user message. With citations enabled, Claude cites them the way it cites web search results. - 12
Files API: upload once, then reference by
file_idin later requests instead of resending bytes. Files are scoped to the workspace, not to a user, so never acceptfile_idvalues from end users. - 13
Message Batches API: asynchronous, 50% of standard price, up to 100,000 requests or 256 MB per batch; most finish within 1 hour but the guarantee is only within 24 hours. Results come back in any order, so match them by
custom_id, and resubmiterrored/expireditems. Results are available for 29 days. Use realtime calls for anything user-facing or deadline-bound. - 14
Cloud platforms (Amazon Bedrock, Google Vertex AI, Microsoft Foundry) use their own SDK clients, auth (AWS IAM, Google Cloud credentials, Azure / Entra ID) and model IDs, and each supports a different set of features. For example, the Files API and Message Batches are not available on Bedrock or Vertex AI, and Batches is not on Foundry. Check every feature your app uses before migrating.
Read the source
Test yourself on Claude API Mechanics
Ten questions, with the answer and explanation after each one.