Evaluate progressive discovery vs. monolithic context strategy
Decide when to load tool definitions, skills and documents on demand (progressive discovery) versus up front (monolithic context), trading context cost and attention against latency and simplicity.
Key points
- 1
Context is a finite attention budget: as tokens grow, precision for retrieval and long-range reasoning degrades ("context rot"). Every tool definition, skill body and document loaded up front competes with the task for attention, even when the window has room.
- 2
Progressive discovery means loading definitions, instructions or documents only when they become relevant, using lightweight identifiers (names, descriptions, file paths, search) to find them just in time. Monolithic context means everything is present from the first token.
- 3
Tool search / deferred loading on the Claude API: include a tool search tool in
tools, mark rarely used toolsdefer_loading: true, and Claude searches the catalog (regex or BM25 variant) and loads only the 3–5 tools it needs. A five-server MCP setup can cost around 55k tokens of definitions; tool search typically cuts this by over 85% and restores selection accuracy, which degrades beyond roughly 30–50 tools. - 4
Use tool search when there are 10 or more tools, definitions exceed about 10k tokens, tool selection accuracy is dropping, or several MCP servers are aggregated. Standard loading is better with fewer than 10 tools that are used on most requests. For MCP servers via the connector, set
defer_loadingon themcp_toolsetdefault config. - 5
Hybrid is the recommended balance: keep the 3–5 most frequently used tools non-deferred so common requests need no search round trip, and defer the long tail. Fully on-demand designs pay a discovery latency on every conversation for the same hot tools.
- 6
Agent Skills implement progressive disclosure in three levels: name and description (about 100 tokens per skill) always in the system prompt; the SKILL.md body (under about 5k tokens) when the skill is triggered; bundled reference files only when read, and scripts run through bash with only their output entering context. Many skills can be installed without a context penalty.
- 7
Content that must always apply (coding standards, compliance rules) belongs in always-loaded context such as
CLAUDE.mdor the system prompt, because a skill or a retrieval step is triggered probabilistically. Large, occasionally needed procedures belong in Skills or on-demand documents. - 8
Just-in-time retrieval by the agent (grep, glob, file paths, search tools) avoids stale pre-computed indexes but is slower than pre-retrieved data; the Claude Code pattern is hybrid:
CLAUDE.mdup front plus tools to explore at runtime. - 9
Monolithic wins when the material is small, stable and needed on nearly every request: it can be placed as a cached prefix, costs one discovery round trip fewer, and is simpler to reason about. Progressive discovery wins when the candidate set is large and only a request-dependent subset is relevant.
- 10
Respect prompt caching: tool definitions are part of the cached prefix, so rewriting, re-sorting or trimming the
toolsarray mid-conversation invalidates the cache and can cost more than the definitions saved.defer_loadingkeeps the prefix untouched and expands discovered tools inline astool_referenceblocks; a custom discovery layer must keep the array stable or append after the cache breakpoint. - 11
Related context-control patterns: programmatic tool calling (Claude writes code that orchestrates tools in a sandbox so intermediate results never enter context), subagents that return condensed summaries, and compaction or tool-result clearing for long sessions.
- 12
MCP client guidance: switch from loading every tool to progressive discovery once definitions consume a meaningful share of the window (a 1–5% threshold is suggested), cache fetched definitions host-side, re-index on
tools/list_changed, and connect or disconnect whole servers on demand for general-purpose agents. - 13
Distractors the exam likes: switching to a larger context window instead of reducing what is loaded, shortening every description to one line (destroys selection quality), splitting servers into separate agents to hide the problem, and going fully on-demand for a small hot set that every request needs.
Read the source
Test yourself on Evaluate progressive discovery vs. monolithic context strategy
Ten questions, with the answer and explanation after each one.