Select appropriate Claude models based on trade-offs
Choose the Claude model for a workload from measured quality, latency and cost trade-offs: the smallest tier that meets the quality bar, routed by task, and upgraded only where evals show the need.
Key points
- 1
Think in tiers and trade-offs, not names or prices. The current families are Claude 5 (Fable 5.1 as the most capable model for the hardest reasoning and long-horizon agentic work, Opus 5.5 in the Opus tier, Sonnet 5 as the balanced tier) plus Haiku 4.5 as the small, fast, low-cost tier. Higher tiers buy capability at the cost of latency and per-token price.
- 2
Start from a quality bar measured on your own representative data, then pick the smallest model that meets it. A quality gap on a labelled sample is evidence; 'the text is nuanced' is not.
- 3
The small tier fits high-volume, low-stakes, latency-bound work: classification, routing, extraction with clear rules, pre-screening of inputs, and worker or subagent roles under a stronger orchestrator.
- 4
Escalate to a higher tier when an eval gap cannot be closed with prompt tuning and examples, when the task needs reasoning across several documents, or when it is a long-horizon agentic loop where small per-step errors compound into failed runs.
- 5
Route by task for mixed workloads: a lightweight classifier or rules sends simple requests to the fast tier and hard ones to a capable tier. One model for everything either overspends on the easy majority or underserves the hard minority.
- 6
Symptoms of model mismatch after the prompt has been tuned: looping, re-reading the same files, stopping short on multi-step work, missing conditions that interact across documents. More turns, more emphasis or a reviewer of the same tier do not fix these.
- 7
'Switch to a bigger model' is the exam's favourite non-answer for problems that live elsewhere: vague prompts, missing programmatic guardrails, broken retrieval, bloated tool sets, or a need for determinism. A larger model does not provide reproducibility or authorisation.
- 8
Before building a two-model cascade, measure the more capable model at a lower
effortsetting on the same tasks; it often meets the bar at comparable cost with far less machinery. A single model means one prompt cache namespace, one thing to monitor and no confidence threshold to tune. - 9
Judge cost per completed task, not per request. A cheaper model that needs more retries, turns or human corrections is not cheaper.
- 10
Current models use adaptive thinking with
effortcontrolling depth; the fixed thinking budget belongs to older models such as Haiku 4.5. Do not build keys on which sampling parameters a specific model accepts. - 11
Context window size is a per-model property, not a tier ranking; check the model's specification rather than assuming a larger tier always has a larger window.
- 12
Pin the model per route or per conversation. Switching models mid-conversation invalidates the prompt cache and changes behaviour, and prompts tuned for one model may over- or under-trigger tools on another, so any model change is re-evaluated.
- 13
When presenting the choice to stakeholders, state the measured trade-off: quality on the eval, time-to-first-token and total latency, cost per task, and any feature or platform availability constraints.
Test yourself on Select appropriate Claude models based on trade-offs
Ten questions, with the answer and explanation after each one.