Optimize context windows and manage token usage
Manage the context window as a finite attention budget: measure token usage, keep only high-signal content, and use retrieval, compaction, context editing, notes and subagents to stay effective on long tasks.
Key points
- 1
The context window is shared between input, thinking and output. Count tokens before sending, reserve headroom for the answer, and give oversized inputs an explicit path (chunked retrieval, summarise-then-answer) instead of silently truncating.
- 2
max_tokenscaps output only; raising it does not make more input fit, and a larger context window does not cure noisy context. - 3
Context rot: accuracy degrades as low-signal tokens accumulate, well before the window is full. The goal of context engineering is the smallest possible set of high-signal tokens that maximises the desired outcome.
- 4
Tool results are the largest and least controlled contributor to agent context. Trim them at the source: return only needed fields, paginate long lists, and keep error messages structured and short.
- 5
Prefer just-in-time retrieval to pre-loading: keep lightweight references (titles, paths, links, IDs) in context and load content when a step needs it. A hybrid that pre-loads a small stable core and retrieves the rest balances speed and flexibility.
- 6
Compaction summarises a conversation nearing its limit while preserving decisions, unresolved issues and implementation details and discarding redundant outputs. Age-based truncation and restarting lose exactly the state that matters. Server-side compaction returns a compaction block that must be passed back with the full response content.
- 7
Context editing clears old tool results (and optionally thinking) from the conversation before the model sees it; it is a clearing strategy, distinct from compaction, which summarises.
- 8
Structured note-taking keeps progress, facts and decisions in persistent memory outside the context window (a notes file or memory tool), so an agent can re-read what it needs rather than re-reading whole sources.
- 9
Sub-agent architectures give each focused task a clean context window and return condensed summaries to the lead agent, isolating verbose exploration.
- 10
In RAG, fewer high-relevance chunks beat many marginal ones; improve retrieval quality with reranking rather than compensating with volume.
- 11
Read the token profile before choosing a lever: if tool results and repeated reads dominate, clearing and notes win; lowering thinking effort only helps when thinking is a large share.
- 12
Tokenizers differ between model generations, so re-measure token counts with the token-counting endpoint when migrating, rather than assuming counts carry over.
Read the source
Test yourself on Optimize context windows and manage token usage
Ten questions, with the answer and explanation after each one.