Study notes · 2.5% of the exam

5.3 Implement error propagation strategies across multi-agent systems

Return structured error context from subagents (failure type, what was attempted, partial results, alternatives), distinguish access failures from valid empty results, recover locally where possible, and carry known gaps through to synthesis.

Key points

  1. 1

    Structured error context has four parts: failure type, the attempted query or URL, any partial results gathered before the failure, and potential alternative approaches. This is what lets a coordinator choose between retrying with a modified query, switching source, or proceeding with partial results.

  2. 2

    Generic error statuses ("search unavailable") hide the context the coordinator needs; there is nothing to protect by withholding details from another component of the same system.

  3. 3

    Two anti-patterns: silently suppressing errors (returning an empty result marked as success) and terminating the entire workflow on a single failure. The first produces reports with invisible gaps; the second throws away every other subagent's completed work.

  4. 4

    Distinguish an access failure (timeout, auth error, rate limit: a retry decision is needed) from a valid empty result (a successful query with zero matches: a finding, not an error). If both are reported as "no results", the coordinator either retries genuine gaps or gives up on transient failures.

  5. 5

    Subagents should implement local recovery for transient failures (retry with backoff, alternative source) and propagate only errors they cannot resolve, including what was attempted and partial results. This keeps the coordinator informed without burdening it with every hiccup and prevents it from duplicating retries the subagent already made.

  6. 6

    Logging a failure for humans to review later is not propagation; the coordinator needs the information at the moment it decides.

  7. 7

    Automatically broadening a query until something returns changes the question asked and hides both failures and genuine gaps.

  8. 8

    Synthesis output should carry coverage annotations: which findings are well supported and which topic areas have gaps because sources were unavailable, and why. A polished report that reads as complete when a key source was inaccessible is a propagation failure at the last step.

  9. 9

    Do not let synthesis fill gaps from general model knowledge or silently drop a section; both make the gap harder to detect.

  10. 10

    Anthropic's research system notes that agentic errors compound: one failed step can send agents down a different trajectory, so systems should resume from where the error occurred rather than restart, and surface failures rather than mask them.

Test yourself on 5.3 Implement error propagation strategies across multi-agent systems

Ten questions, with the answer and explanation after each one.