Study notes · 3.4% of the exam

4.5 Design efficient batch processing strategies

Choose the Message Batches API for latency-tolerant, non-blocking bulk work, derive submission cadence from the SLA, handle failures selectively by `custom_id`, and refine prompts on a sample before committing large volumes.

Key points

  1. 1

    Message Batches API facts: 50% lower cost than synchronous calls, results within up to 24 hours (most batches finish much sooner), no guaranteed latency SLA, and each request in a batch is processed independently.

  2. 2

    Appropriate: overnight technical-debt reports, weekly compliance audits, nightly test generation, large evaluations. Inappropriate: anything a person or pipeline is blocked on, such as a pre-merge check or an interactive "explain this failure" request. Match the API to the workflow's latency requirement.

  3. 3

    "Batches are usually fast" is not a guarantee; do not put blocking work on the batch API with a timeout fallback. Keep synchronous calls for blocking steps and batch the rest.

  4. 4

    Every batch request needs a unique custom_id. Results can come back in any order, so always correlate by custom_id, never by position.

  5. 5

    Result types per request: succeeded, errored (for example an invalid request or a document over the context limit), expired (the 24-hour window passed before the request ran) and canceled. Errored, expired and canceled requests are not billed, but they are also not processed.

  6. 6

    Handle failures selectively: resubmit only the failed custom_id values. Resubmit expired requests unchanged; modify requests that errored for a repeatable reason first (chunk or summarise oversized documents, then encode the parent document in the new custom_id values so pieces can be reassembled). Never resubmit the whole batch.

  7. 7

    Cadence from SLA: longest submission interval = deadline minus maximum batch processing time minus any margin. With a 30-hour deadline, 24-hour batches and a 2-hour margin, submit at least every 4 hours. Daily submission can take 48 hours worst case.

  8. 8

    Requests that expire close to a deadline should be resubmitted synchronously, not into the next batch, because a new batch may take another 24 hours. Only the urgent few go synchronous; the bulk stays batched.

  9. 9

    A batch request can include tools and may return a tool_use block, but the batch cannot execute the tool and continue the same request. Multi-turn tool loops (extract, look up vendor via MCP, finalise) must be split into stages with application-side tool execution between batches.

  10. 10

    A batch cannot run a validate-and-retry loop inside itself either; failures come back as failures and need another submission round, which costs money and up to a day. Refine the prompt synchronously on a stratified sample of a few hundred documents until the first-pass success rate is acceptable, then submit the large batch.

  11. 11

    Streaming, fast mode and cache pre-warming (max_tokens: 0) are not available inside batches. Prompt caching works on a best-effort basis; consider the 1-hour cache duration when many requests share a large prefix.

  12. 12

    Dry-run a single request shape through the synchronous Messages API before batching; batch validation errors are only reported when the whole batch has ended.

  13. 13

    Distractors the exam likes: switching an entire workload to synchronous calls after a handful of failures (doubles cost), submitting several prompt variants against the full volume instead of a sample, and treating unbilled expired requests as complete.

Test yourself on 4.5 Design efficient batch processing strategies

Ten questions, with the answer and explanation after each one.