LLM Fundamentals
Explain how LLMs generate text (tokens, context windows, sampling, non-determinism) and pick the right model options (thinking, effort, fast mode) and prompting technique (zero-, single-, multi-shot) for a scenario.
Key points
- 1
Claude generates output one token at a time, each conditioned on the prompt and all earlier tokens. Reasoning written after an answer cannot change that answer, so let the model reason (or think) before it commits.
- 2
Tokens are subword units, not words or characters, and tokenizers can differ between models. Limits, pricing, and
max_tokensare all measured in tokens. - 3
The context window holds everything in the request (system prompt, tool definitions, messages, tool results) plus the output generated. The API is stateless, so history is resent and grows each turn.
- 4
max_tokenscaps generated output only. Reaching it stops output mid-thought withstop_reason: "max_tokens". It does not enlarge the context window. - 5
Output is sampled, so identical requests can return different text. Test with property checks and evals over many samples, not exact-string matches, and do not count on low temperature for reproducibility.
- 6
Adaptive thinking (
thinking: {type: "adaptive"}) lets Claude decide when and how much to reason, and it is the thinking mode on current models. Older fixedbudget_tokensthinking is deprecated or rejected on the newest models. - 7
Thinking tokens are billed as output tokens and count toward
max_tokens, even when the thinking text is hidden. On many current models the thinkingdisplaydefaults to"omitted"; request"summarized"to show reasoning summaries. - 8
Effort (
output_config.effort:low,medium,high, and higher levels on some models) trades thoroughness for tokens, latency, and cost. It affects all output, including tool calls, andhighis the default on most models. - 9
Use lower effort for simple, high-volume, or latency-sensitive routes and higher effort for complex reasoning and agentic coding. Tune effort per route and confirm with evals.
- 10
Fast mode (research preview, supported Opus models only) runs the same model at up to 2.5x higher output speed for premium pricing. It is a speed-for-money trade, not a quality change, and it is not available with the Batch API.
- 11
Zero-shot uses instructions only; single-shot adds one example; multi-shot (few-shot) adds several. Docs recommend 3 to 5 relevant, diverse examples wrapped in
<example>tags. - 12
Examples teach everything they share. If all examples have the same length or sections, outputs will copy that, so vary them and include edge cases.
- 13
Trap: temperature and
max_tokensare not quality levers for reasoning. Use thinking and effort instead.
Read the source
Test yourself on LLM Fundamentals
Ten questions, with the answer and explanation after each one.