Study notes · 4.5% of the exam

Agent Architecture

Choose between a workflow and an agent for a given requirement, pick the right workflow pattern, and design supervisor/subagent hierarchies that justify their cost.

Key points

  1. 1

    Workflows orchestrate LLM calls and tools through predefined code paths. Agents let the model dynamically direct its own steps and tool use. Start with the simplest solution; often a single well-prompted call with retrieval is enough.

  2. 2

    Pick a workflow when the steps are known in advance and you need predictable cost, latency and auditability. Pick an agent when the task is open-ended and the number or order of steps cannot be hard-coded.

  3. 3

    Bound agent autonomy with stopping conditions (turn or budget limits) and human checkpoints. Agents trade higher cost and compounding errors for flexibility.

  4. 4

    Prompt chaining: a fixed sequence of calls, with programmatic "gates" between steps to catch errors early. It trades latency for accuracy.

  5. 5

    Routing: classify the input, then send it to a specialized prompt, tool set or model (for example, easy questions to a smaller model). Use it when inputs fall into distinct, reliably classifiable categories.

  6. 6

    Parallelization comes in two forms. Sectioning runs independent subtasks at the same time; a guardrail screening call in parallel with the answering call is a classic example, and it beats one call doing both. Voting runs the same task several times for higher confidence.

  7. 7

    Orchestrator-workers: a central model decides the subtasks for each input, delegates them, and synthesizes the results. It differs from sectioning because the subtasks are not predefined.

  8. 8

    Evaluator-optimizer: one call generates and another critiques, in a loop. It fits when clear evaluation criteria exist and feedback measurably improves the output.

  9. 9

    Supervisor/manager hierarchy: one orchestrator owns planning, delegation, termination and the final answer, and workers report back to it. Free peer-to-peer messaging between agents leads to loops and unclear ownership.

  10. 10

    Subagents work in their own context windows and return condensed results. Their benefits are parallelism, context isolation and specialization (scoped prompts and tools).

  11. 11

    Multi-agent cost trap: in Anthropic's data, agents use about 4 times the tokens of chat, and multi-agent systems about 15 times. They fit valuable, highly parallel work or work that exceeds one context window, and fit poorly with tightly coupled tasks that need shared context (most coding).

  12. 12

    Delegate clearly: give each worker an objective, scope boundaries, sources or tools, and an output format. Embed effort-scaling rules so the orchestrator does not spawn many subagents for simple queries.

  13. 13

    Avoid the "game of telephone": have workers return structured summaries, and store full artifacts externally so the lead receives references rather than paraphrasing everything through its own context.

Test yourself on Agent Architecture

Ten questions, with the answer and explanation after each one.