Cost and Token Management
Track and forecast token usage and cost, and cut spend with prompt caching (breakpoints, TTLs, hit accounting), batching, and output control.
Key points
- 1
Every response has a
usageobject:input_tokens(uncached input after the last breakpoint),output_tokens,cache_creation_input_tokens, andcache_read_input_tokens. Log it per request and tag it by feature. - 2
With caching, total input is
input_tokens+cache_creation_input_tokens+cache_read_input_tokens, and each is priced differently. Price each field at its own rate. - 3
Use the token counting endpoint (
POST /v1/messages/count_tokens) with the target model to estimate input before sending. It is free but rate-limited, and its counts are estimates. Do not use other vendors' tokenizers. - 4
Tokenizers can change between models, and thinking adds billed output. Re-baseline with
count_tokensand a pilot when forecasting a migration. - 5
Output tokens cost several times more than input tokens. When output dominates spend, constrain length and format in the prompt (and tune effort) rather than truncating with
max_tokens. - 6
Prompt caching is an exact prefix match in the order tools, then system, then messages. Any change before a breakpoint (a timestamp, reordered JSON, or a different tool set) invalidates everything after it.
- 7
Keep stable content first and volatile content after the last
cache_controlbreakpoint. A request can have up to 4 breakpoints. Top-level automatic caching places the breakpoint on the last cacheable block for you. - 8
For multi-turn chat, put the breakpoint on the last block of the latest turn (or use automatic caching) so each request reads the prior conversation from cache.
- 9
The default TTL is 5 minutes.
ttl: "1h"is available, and a cache read refreshes the entry. Cache writes cost 1.25x base input (5-minute) or 2x (1-hour), and reads cost about 0.1x. - 10
Prefixes below a model-dependent minimum length are not cached, and there is no error. Verify caching with
cache_read_input_tokens: zero across repeated identical-prefix requests means something is invalidating the cache. - 11
Changing tools or switching models mid-conversation invalidates the whole cache (caches are model-scoped). Keep one stable tool list and pass modes in message content.
- 12
The Message Batches API bills at 50% of standard prices, and most batches finish within an hour (up to 24 hours). Results can return in any order, so match them by
custom_id. - 13
Caching and batch discounts can stack, but cache hits in batches are best effort. Use the 1-hour TTL for shared batch context and forecast a partial hit rate.
Read the source
Test yourself on Cost and Token Management
Ten questions, with the answer and explanation after each one.