Study notes · 2.7% of the exam

Model Selection and Tradeoffs

Choose between Opus, Sonnet, and Haiku by weighing quality, latency, and cost for the task, account for feature support such as adaptive thinking, and migrate safely across model releases.

Key points

  1. 1

    Opus is the most capable tier, for complex reasoning and long-horizon agentic or coding work. Sonnet balances capability, speed, and cost. Haiku is the fastest and most cost-efficient, for simple, high-volume, latency-sensitive tasks.

  2. 2

    Start from requirements and measure on your own eval set. Public leaderboards do not reflect your workload's quality, latency, and cost constraints.

  3. 3

    Judge cost per completed task, not per token. A cheaper model that fails and retries on hard tasks can cost more overall.

  4. 4

    When one tier is too weak and another too costly, evaluate the middle tier. When stages differ, choose per stage, such as a small tier for bulk extraction and a top tier for rare, high-stakes synthesis.

  5. 5

    Tuning effort within one model is often a better lever than switching models. Measure lower effort on the current model before migrating or building cascades.

  6. 6

    Feature support differs by model. For example, the smaller Haiku model uses the older thinking configuration rather than adaptive thinking, and fast mode is offered on supported Opus models only. Build request parameters per model from documented capabilities or the Models API.

  7. 7

    Pin exact model IDs in production so behavior does not change underneath you, and move to new releases deliberately.

  8. 8

    New releases can bring breaking API changes. Recent models reject assistant prefill, and the newest reject fixed thinking budget_tokens and sampling parameters, which show up as 400 errors. Read the migration guide before switching.

  9. 9

    Releases also shift behavior without any API change. For example, newer models follow system prompts more closely, so emphatic "CRITICAL: you MUST" wording written for older models can cause tool overtriggering. Dial it back.

  10. 10

    Safe migration means reading the migration guide, updating request code, re-running evals, re-tuning prompts and effort, and rolling out gradually.

  11. 11

    Deprecated models still work but have a retirement date. Requests to retired models fail, so track deprecation notices.

  12. 12

    Trap: a bigger model is not automatically "safer", and a smaller model with more retries does not fix a capability gap.

Test yourself on Model Selection and Tradeoffs

Ten questions, with the answer and explanation after each one.