4.5 Design efficient batch processing strategies
Choose the Message Batches API for latency-tolerant, non-blocking bulk work, derive submission cadence from the SLA, handle failures selectively by `custom_id`, and refine prompts on a sample before committing large volumes.
Key points
- 1
Message Batches API facts: 50% lower cost than synchronous calls, results within up to 24 hours (most batches finish much sooner), no guaranteed latency SLA, and each request in a batch is processed independently.
- 2
Appropriate: overnight technical-debt reports, weekly compliance audits, nightly test generation, large evaluations. Inappropriate: anything a person or pipeline is blocked on, such as a pre-merge check or an interactive "explain this failure" request. Match the API to the workflow's latency requirement.
- 3
"Batches are usually fast" is not a guarantee; do not put blocking work on the batch API with a timeout fallback. Keep synchronous calls for blocking steps and batch the rest.
- 4
Every batch request needs a unique
custom_id. Results can come back in any order, so always correlate bycustom_id, never by position. - 5
Result types per request:
succeeded,errored(for example an invalid request or a document over the context limit),expired(the 24-hour window passed before the request ran) andcanceled. Errored, expired and canceled requests are not billed, but they are also not processed. - 6
Handle failures selectively: resubmit only the failed
custom_idvalues. Resubmit expired requests unchanged; modify requests that errored for a repeatable reason first (chunk or summarise oversized documents, then encode the parent document in the newcustom_idvalues so pieces can be reassembled). Never resubmit the whole batch. - 7
Cadence from SLA: longest submission interval = deadline minus maximum batch processing time minus any margin. With a 30-hour deadline, 24-hour batches and a 2-hour margin, submit at least every 4 hours. Daily submission can take 48 hours worst case.
- 8
Requests that expire close to a deadline should be resubmitted synchronously, not into the next batch, because a new batch may take another 24 hours. Only the urgent few go synchronous; the bulk stays batched.
- 9
A batch request can include tools and may return a
tool_useblock, but the batch cannot execute the tool and continue the same request. Multi-turn tool loops (extract, look up vendor via MCP, finalise) must be split into stages with application-side tool execution between batches. - 10
A batch cannot run a validate-and-retry loop inside itself either; failures come back as failures and need another submission round, which costs money and up to a day. Refine the prompt synchronously on a stratified sample of a few hundred documents until the first-pass success rate is acceptable, then submit the large batch.
- 11
Streaming, fast mode and cache pre-warming (
max_tokens: 0) are not available inside batches. Prompt caching works on a best-effort basis; consider the 1-hour cache duration when many requests share a large prefix. - 12
Dry-run a single request shape through the synchronous Messages API before batching; batch validation errors are only reported when the whole batch has ended.
- 13
Distractors the exam likes: switching an entire workload to synchronous calls after a handful of failures (doubles cost), submitting several prompt variants against the full volume instead of a sample, and treating unbilled expired requests as complete.
Read the source
Test yourself on 4.5 Design efficient batch processing strategies
Ten questions, with the answer and explanation after each one.