Evaluate accuracy-latency trade-offs and justify configuration decisions
Know which configuration levers buy accuracy with latency and which buy latency without changing answers, and justify each choice with measured evals against the path's SLA.
Key points
- 1
Levers that raise accuracy at a latency and cost price: a more capable model, extended or adaptive thinking at higher
effort, deeper retrieval with reranking, more context, and agentic loops that take more turns. Anthropic's guidance: agentic systems trade latency and cost for better task performance, so decide when that trade makes sense. - 2
Levers that cut latency or cost without changing what the model concludes: streaming (perceived latency), prompt caching (repeated prefixes, lower time-to-first-token and cost), parallel tool calls, batch processing for asynchronous work, and routing simple inputs to a smaller or lower-effort configuration.
- 3
Streaming does not shorten total generation time; it delivers tokens as they are produced so time-to-first-token becomes the wait the user feels. It helps chat and long outputs, not a single classification label, and the SDKs require it for very large
max_tokens. - 4
effort(low,medium,high,xhigh,max) is the primary control for reasoning depth on adaptive-thinking models; lower effort means fewer thinking tokens, fewer and terser tool calls, and lower latency, with some capability reduction. Set it per workload and confirm on evals. - 5
Changing the top-level effort or thinking configuration between requests invalidates prompt-cache breakpoints, so hold it constant within a cached conversation; only models with per-message effort can change it mid-conversation cache-safely.
- 6
Route rather than configure globally: when evals show reasoning helps only a subset of inputs (for example multi-order messages), send the majority through the fast configuration and only that subset through the slow one.
- 7
Parallel tool use collapses independent lookups into one turn: the model emits several
tool_useblocks, your code runs them concurrently, and alltool_resultblocks return in one user message.disable_parallel_tool_useforces serial calls and is the opposite lever. - 8
Match configuration to each path's SLA: interactive paths take the most accurate configuration that fits the latency budget; asynchronous or audited paths with no latency limit can take the slowest, most accurate configuration, often through the Batch API at lower cost.
- 9
Justify with numbers: state the measured accuracy and p95 latency of each candidate configuration against the SLA, and record the accepted accuracy gap. A decision that changes the SLA instead of designing to it is not a justification.
- 10
Truncating output (
max_tokens), truncating policy documents, or "be concise" instructions reduce latency by sacrificing correctness and are common distractors; so is "switch to a bigger model" without evidence of the latency or accuracy effect. - 11
Thinking with very large budgets produces long-running requests; for the heaviest reasoning use batch processing rather than an interactive request that may hit timeouts.
- 12
Prompt caching addresses input-side latency and cost (place stable content first, dynamic content after the breakpoint); it does nothing for output-generation time or accuracy.
Read the source
Test yourself on Evaluate accuracy-latency trade-offs and justify configuration decisions
Ten questions, with the answer and explanation after each one.