Apply retrieval strategies matched to data shape and query pattern
Choose the retrieval mechanism from the shape of the data and the pattern of the queries: semantic search for prose, lexical search for identifiers, tool or SQL calls for structured records, and plain in-context text for small corpora.
Key points
- 1
Data shape decides the mechanism. Prose and FAQs: semantic (embedding) retrieval. Codes, SKUs, error numbers, names and exact phrases: lexical (BM25 or exact match). Tables and live systems with filter-and-aggregate questions: a parameterized SQL or API tool call. Small, stable corpora (under about 200k tokens): include everything in a cached prompt.
- 2
Embedding the rows of a database table and answering counts or sums by similarity search is an anti-pattern; aggregation is a query, not a retrieval. Give Claude a tool that runs a parameterized query and returns the computed result.
- 3
Exact-identifier queries fail on dense embeddings because vectors capture meaning, not character-level identity (E-4471 and E-4417 look alike). Add a lexical index and fuse results, and extract identifiers into filterable metadata at index time.
- 4
Hybrid search (semantic plus BM25, fused with rank fusion) is the default for mixed corpora; Anthropic's contextual BM25 results show the lexical signal materially reduces retrieval failures even for natural-language questions.
- 5
Reranking: retrieve a broad candidate set (Anthropic used top-150), score each candidate against the query with a reranker, and pass the top 20 to the model. It adds a small amount of latency and compute and is the targeted fix when the right chunk is retrievable but ranked outside the window sent to the model.
- 6
Pass around 20 chunks to the model rather than 5 or 10 when latency allows; Anthropic found top-20 most effective in its evaluation. Passing 100+ chunks is a sign of retrieval that cannot discriminate and degrades answers through noise.
- 7
Live or fast-changing records (orders, tickets, balances) should be fetched through an authorized tool call at query time, keyed by the exact identifier, never from a nightly embedded snapshot. A missing record should be reported as not found, not inferred from similar records.
- 8
Route by query pattern when one product has several data shapes: a lightweight classifier or the model's own tool choice sends policy questions to retrieval, id lookups to a tool, and glossary or reference questions to in-context text.
- 9
When answers must show their sources, return retrieved passages as
search_resultblocks with citations enabled so Claude cites the source and title; full-text database search or plain string concatenation gives no passage-level attribution. - 10
Authorization is part of retrieval: filter by the user's permissions before content enters the context, and call live systems with the user's delegated identity. Prompt-level instructions are not access control.
- 11
Cost and latency levers: cache the static prefix, keep top-k modest, rerank a wide candidate set instead of sending it all, and avoid multiple retrieval calls per question unless the budget allows.
- 12
Distractors the exam likes: one embedding index for every data type, larger embedding dimensions to fix exact-match failures, instructing Claude to double-check codes instead of fixing retrieval, and building a full RAG stack for a handbook that fits in the prompt.
Read the source
Test yourself on Apply retrieval strategies matched to data shape and query pattern
Ten questions, with the answer and explanation after each one.