Study notes · 5.2% of the exam

LLM Fundamentals

Explain how LLMs generate text (tokens, context windows, sampling, non-determinism) and pick the right model options (thinking, effort, fast mode) and prompting technique (zero-, single-, multi-shot) for a scenario.

Key points

  1. 1

    Claude generates output one token at a time, each conditioned on the prompt and all earlier tokens. Reasoning written after an answer cannot change that answer, so let the model reason (or think) before it commits.

  2. 2

    Tokens are subword units, not words or characters, and tokenizers can differ between models. Limits, pricing, and max_tokens are all measured in tokens.

  3. 3

    The context window holds everything in the request (system prompt, tool definitions, messages, tool results) plus the output generated. The API is stateless, so history is resent and grows each turn.

  4. 4

    max_tokens caps generated output only. Reaching it stops output mid-thought with stop_reason: "max_tokens". It does not enlarge the context window.

  5. 5

    Output is sampled, so identical requests can return different text. Test with property checks and evals over many samples, not exact-string matches, and do not count on low temperature for reproducibility.

  6. 6

    Adaptive thinking (thinking: {type: "adaptive"}) lets Claude decide when and how much to reason, and it is the thinking mode on current models. Older fixed budget_tokens thinking is deprecated or rejected on the newest models.

  7. 7

    Thinking tokens are billed as output tokens and count toward max_tokens, even when the thinking text is hidden. On many current models the thinking display defaults to "omitted"; request "summarized" to show reasoning summaries.

  8. 8

    Effort (output_config.effort: low, medium, high, and higher levels on some models) trades thoroughness for tokens, latency, and cost. It affects all output, including tool calls, and high is the default on most models.

  9. 9

    Use lower effort for simple, high-volume, or latency-sensitive routes and higher effort for complex reasoning and agentic coding. Tune effort per route and confirm with evals.

  10. 10

    Fast mode (research preview, supported Opus models only) runs the same model at up to 2.5x higher output speed for premium pricing. It is a speed-for-money trade, not a quality change, and it is not available with the Batch API.

  11. 11

    Zero-shot uses instructions only; single-shot adds one example; multi-shot (few-shot) adds several. Docs recommend 3 to 5 relevant, diverse examples wrapped in <example> tags.

  12. 12

    Examples teach everything they share. If all examples have the same length or sections, outputs will copy that, so vary them and include edge cases.

  13. 13

    Trap: temperature and max_tokens are not quality levers for reasoning. Use thinking and effort instead.

Test yourself on LLM Fundamentals

Ten questions, with the answer and explanation after each one.