Study notes · 2.6% of the exam

Optimize token usage, latency, and cost-performance trade-offs

Reduce token usage, latency and cost with caching, batching, routing, right-sized thinking budgets and context trimming, and measure the quality impact of every optimization on the eval.

Key points

  1. 1

    Prompt caching reuses an exact prefix. The prefix is built in the order tools, then system, then messages; a change anywhere invalidates that level and everything after it. Put static content (tool definitions, system instructions, documents, examples) first and dynamic content (user message, timestamps) after the breakpoint.

  2. 2

    Cache invalidators to remember: changed or reordered tool definitions, edited system prompt, added or removed images, changed thinking parameters (including budget_tokens), changed effort setting, changed tool_choice, and any per-request value placed inside the prefix.

  3. 3

    Cache economics: writes cost more than base input, reads cost a small fraction of it; the default lifetime is 5 minutes with a 1-hour option for longer reuse windows. Prompts shorter than the model's minimum cacheable length are not cached even if marked.

  4. 4

    Verify caching with the usage fields: cache_creation_input_tokens (written), cache_read_input_tokens (read) and input_tokens (uncached). High creation with near-zero reads means the prefix is changing between requests.

  5. 5

    The Message Batches API processes large volumes asynchronously at a substantial discount; most batches finish within an hour and may take up to 24 hours. Use it for evaluations, bulk classification, moderation and nightly jobs, never for interactive traffic. Caching also works inside batches.

  6. 6

    Extended thinking: budget_tokens is a target, not a hard cap, must be at least 1,024 and less than max_tokens; larger budgets improve hard tasks with diminishing returns and higher latency. Start near the minimum for simple tasks (larger for complex ones) and tune incrementally against the eval; keep the budget stable within a cached conversation; use batch processing for budgets above 32k.

  7. 7

    Routing: classify requests and send simple ones to a smaller, faster model, keeping the capable model for hard steps; validate on the eval slice for each route before rolling out.

  8. 8

    Context trimming: send only what the task needs, summarize or compact long histories, use retrieval instead of whole corpora when questions are local, and limit output length by instruction. Blanket max_tokens caps truncate mid-answer; removing grounding context saves tokens by sacrificing correctness.

  9. 9

    Latency levers: streaming for perceived responsiveness, caching for time-to-first-token on long prefixes, smaller models or lower thinking budgets for simple work, parallel tool calls, and avoiding unnecessary turns.

  10. 10

    Every optimization is a hypothesis about quality: measure accuracy, groundedness and safety on the eval before and after, and keep cost per successful task and p95 latency on the dashboard.

  11. 11

    Exam-style anti-patterns: truncating required documents, switching everything to the smallest model regardless of task fit, batching interactive traffic, extending cache TTL to fix a changing prefix, moving the breakpoint onto per-request content.

Test yourself on Optimize token usage, latency, and cost-performance trade-offs

Ten questions, with the answer and explanation after each one.