Study notes · 2.8% of the exam

Cost and Token Management

Track and forecast token usage and cost, and cut spend with prompt caching (breakpoints, TTLs, hit accounting), batching, and output control.

Key points

  1. 1

    Every response has a usage object: input_tokens (uncached input after the last breakpoint), output_tokens, cache_creation_input_tokens, and cache_read_input_tokens. Log it per request and tag it by feature.

  2. 2

    With caching, total input is input_tokens + cache_creation_input_tokens + cache_read_input_tokens, and each is priced differently. Price each field at its own rate.

  3. 3

    Use the token counting endpoint (POST /v1/messages/count_tokens) with the target model to estimate input before sending. It is free but rate-limited, and its counts are estimates. Do not use other vendors' tokenizers.

  4. 4

    Tokenizers can change between models, and thinking adds billed output. Re-baseline with count_tokens and a pilot when forecasting a migration.

  5. 5

    Output tokens cost several times more than input tokens. When output dominates spend, constrain length and format in the prompt (and tune effort) rather than truncating with max_tokens.

  6. 6

    Prompt caching is an exact prefix match in the order tools, then system, then messages. Any change before a breakpoint (a timestamp, reordered JSON, or a different tool set) invalidates everything after it.

  7. 7

    Keep stable content first and volatile content after the last cache_control breakpoint. A request can have up to 4 breakpoints. Top-level automatic caching places the breakpoint on the last cacheable block for you.

  8. 8

    For multi-turn chat, put the breakpoint on the last block of the latest turn (or use automatic caching) so each request reads the prior conversation from cache.

  9. 9

    The default TTL is 5 minutes. ttl: "1h" is available, and a cache read refreshes the entry. Cache writes cost 1.25x base input (5-minute) or 2x (1-hour), and reads cost about 0.1x.

  10. 10

    Prefixes below a model-dependent minimum length are not cached, and there is no error. Verify caching with cache_read_input_tokens: zero across repeated identical-prefix requests means something is invalidating the cache.

  11. 11

    Changing tools or switching models mid-conversation invalidates the whole cache (caches are model-scoped). Keep one stable tool list and pass modes in message content.

  12. 12

    The Message Batches API bills at 50% of standard prices, and most batches finish within an hour (up to 24 hours). Results can return in any order, so match them by custom_id.

  13. 13

    Caching and batch discounts can stack, but cache hits in batches are best effort. Use the 1-hour TTL for shared batch context and forecast a partial hit rate.

Test yourself on Cost and Token Management

Ten questions, with the answer and explanation after each one.