Design a RAG pipeline with appropriate chunking and indexing strategies
Design a retrieval-augmented generation pipeline whose chunking, indexing, metadata and refresh process fit the documents and the questions, and diagnose retrieval regressions before blaming the model.
Key points
- 1
Chunk along document structure (sections, clauses, headings, table rows) and at the granularity of the questions being asked, not at a fixed token count. Fixed-size chunks that cut through a clause lose meaning at the boundary; add a small overlap and carry the section heading into each chunk.
- 2
Contextual Retrieval (Anthropic, 2024): before embedding and before building the BM25 index, prepend a short Claude-generated context (50–100 tokens) that situates the chunk within its document. Contextual embeddings alone cut top-20 retrieval failures by 35% (5.7% to 3.7%); adding contextual BM25 cut them by 49% (to 2.9%); adding reranking cut them by 67% (to 1.9%). Prompt caching makes generating the context cheap.
- 3
Index with both dense embeddings (semantic recall) and a lexical BM25 index (exact terms, codes, names), then fuse the results. Hybrid search is the default recommendation, not an optimization for later.
- 4
Store structured metadata with every chunk (document id, version, effective date, jurisdiction, product, tenant, permissions) so retrieval can filter before ranking. Similarity cannot express constraints such as "in force last March" or "only documents this user may see".
- 5
Match chunk size to query granularity: small chunks for precise factual matching, larger structural units for broad questions. When one corpus serves both, index small chunks but return the parent section (small-to-big retrieval) instead of doubling retrieval calls or compromising on one size.
- 6
If the whole knowledge base is under roughly 200,000 tokens (about 500 pages), skip RAG: put it all in the prompt as a cached prefix. RAG exists for corpora that do not fit.
- 7
Refresh and index hygiene: re-index by upserting chunks keyed by document id and version, delete superseded chunks, and pin the embedding model version. Appending new chunks while old ones remain, or re-embedding the corpus with a different model than the query encoder, are the classic post-refresh regressions.
- 8
Exam heuristic (official sample 3): confident-but-wrong answers immediately after a document refresh, with model and latency unchanged, point at the retrieval or indexing step returning stale or irrelevant chunks. Model weights, temperature and context-window size are not triggered by a refresh.
- 9
Diagnose retrieval by isolating branches: run BM25-only and dense-only test queries. Lexical healthy plus dense broken after a re-index means embedding mismatch; both broken with new documents missing means the index job failed; both fine but answers wrong means the problem is in the prompt or model.
- 10
Gate every refresh with an automated retrieval regression suite (known questions with expected source chunks, measuring recall@k) so a broken re-index is caught before users see it.
- 11
Return retrieved passages to Claude as
search_resultcontent blocks (from a tool result or as top-level content) withcitationsenabled so answers cite the source and title you supplied; for custom control of citation granularity use custom-content documents, since plain-text and PDF documents are chunked into sentences automatically. - 12
Use the Files API to upload large PDFs or text once and reference them by
file_idacross requests; file ids are workspace-scoped, so never accept afile_idfrom an end user and give each tenant its own workspace when isolation matters. - 13
Distractors the exam likes: increasing top-k to compensate for weak retrieval, switching to a bigger model or longer context window to fix a retrieval fault, telling Claude in the prompt to prefer newer or matching chunks, and embedding whole documents as one vector.
Test yourself on Design a RAG pipeline with appropriate chunking and indexing strategies
Ten questions, with the answer and explanation after each one.