Context Engineering
Keep an agent's context focused and within limits: manage the window, prevent drift and bloat (tool output pruning, clearing, compaction), use memory, and isolate context with subagents or multi-step workflows.
Key points
- 1
Everything in the request (system prompt, tool definitions, messages, tool results) plus the output shares the context window. Large tool sets and verbose tool results are common hidden consumers.
- 2
Quality drops as irrelevant tokens pile up, well before the hard limit. Design tools to return focused excerpts or IDs, and load details just in time instead of preloading everything.
- 3
Context editing clears content: old tool results (
clear_tool_uses_20250919) or thinking blocks (clear_thinking_20251015). It prunes rather than summarizes. - 4
Every clear changes the prompt prefix and invalidates the cache from that point. Set a high trigger and clear in larger batches so clears are infrequent.
- 5
Compaction replaces older turns with a server-written summary so long conversations and agent tasks can continue. With server-side compaction, append the full
response.content(including compaction blocks), not just the text. - 6
Anything a summary omits is gone. Customize the summarization instructions to keep identifiers, open tasks, and exact errors, or keep recent turns verbatim when they must reach Claude word for word.
- 7
The memory tool lets Claude store and retrieve files across conversations, but it runs client-side. Your app executes each command against storage you control and must reject paths outside the memory directory.
- 8
For work spanning multiple context windows, keep state outside the window (progress notes, a structured test-status file, git history), and start each fresh window by telling the agent exactly which files to read.
- 9
Context drift (forgotten constraints, repeated work) is prevented by durable state plus summaries that preserve goals, not by keeping every tool output forever.
- 10
Subagents isolate context: each does heavy reading in its own window and returns a condensed result for the orchestrator to synthesize. Use them for parallel, independent, or reading-heavy work.
- 11
Subagents add latency and tokens. If an agent over-delegates (for example for a single grep), add guidance on when to delegate and when to work directly.
- 12
Multi-step workflows isolate context too. Process each item in its own request to a fixed schema, then run a synthesis step over the compact results.
- 13
Trap: a bigger context window postpones limits but does not fix distraction from irrelevant content.
Read the source
Test yourself on Context Engineering
Ten questions, with the answer and explanation after each one.