Study notes · 2.6% of the exam

Implement prompt reuse strategies (caching, modular prompts, Skills)

Reuse prompt content efficiently with prompt caching for stable prefixes, modular versioned templates for shared sections, and Agent Skills for procedural knowledge loaded on demand.

Key points

  1. 1

    Prompt caching is an exact prefix match. The prompt renders as tools, then system, then messages; a cache_control breakpoint caches everything up to that point, and any byte change earlier in the prefix invalidates everything after it.

  2. 2

    Put static content first (tools, system prompt, reference documents, examples) and dynamic content after the last breakpoint or in the user turn. Timestamps, request IDs, user names and unsorted JSON at the front of the system prompt silently defeat caching.

  3. 3

    Up to four breakpoints per request; place them at stability boundaries (end of tools and system, end of shared reference material, last turn of a growing conversation). A top-level automatic option places one breakpoint on the last cacheable block for simple multi-turn cases.

  4. 4

    The default cache lifetime is five minutes, refreshed each time the prefix is read; a one-hour option exists for prefixes reused across longer gaps. Cache writes cost more than normal input and reads cost much less, so the longer lifetime must earn its higher write cost; entries with the longer lifetime must precede shorter ones.

  5. 5

    Invalidators: changing the model (caches are model-scoped), adding, removing or reordering tools, changing thinking or effort settings, and editing anything in the prefix. Changes to later messages only affect the messages tier.

  6. 6

    There is a model-dependent minimum cacheable prefix length; shorter prefixes silently do not cache. Verify with usage: cache_read_input_tokens at zero on identical prefixes means an invalidator; cache_creation_input_tokens near full size every request means the prefix is being rewritten or the model is alternating.

  7. 7

    Caching cuts input cost and time-to-first-token for repeated content; the Batch API cuts cost for asynchronous work but increases latency. They address different problems and can be combined.

  8. 8

    Modular prompts: compose each prompt from versioned modules (shared compliance or policy block, product block, task block) rendered in a fixed order with shared static modules first. One source of truth for shared text, and the stable prefix caches across all variants.

  9. 9

    Agent Skills package procedural knowledge as a directory with a SKILL.md whose YAML frontmatter has a required name and description; the description must say what the Skill does and when to use it because it is what triggers the Skill.

  10. 10

    Progressive disclosure: level 1 (name and description) is always in context at a small cost per Skill; level 2 (the SKILL.md body) loads only when triggered; level 3 (bundled reference files and scripts) loads only when referenced, and scripts run via bash so only their output enters context.

  11. 11

    Skills fit knowledge that is sometimes needed and procedural (workflows, formatting rules, house style, tooling steps) in an environment with a filesystem and code execution: the API with the code execution tool, Claude Code (.claude/skills/ or ~/.claude/skills/) and claude.ai. Content needed on every request belongs in a cached system prompt, not a Skill.

  12. 12

    Custom Skills do not sync across surfaces (claude.ai, API, Claude Code) and have different sharing scopes; they are installed software, so use only trusted sources and audit bundled scripts.

  13. 13

    Common exam distractors: truncating or summarising required content to save tokens, moving shared text into the user turn, adding breakpoints to 'extend' a cache lifetime, and packaging always-needed content as a Skill.

Test yourself on Implement prompt reuse strategies (caching, modular prompts, Skills)

Ten questions, with the answer and explanation after each one.