Study notes · 2.8% of the exam

Design end-to-end architectures (input → processing → output → feedback loops)

Design the full flow, input → processing → output → feedback, choosing the processing mode per workload constraint, gating output deterministically, and closing the loop so human corrections and evaluation results drive iteration.

Key points

  1. 1

    Input stage: normalise every channel (email, web form, scans, API) into a canonical record at ingestion so downstream prompts see one shape and every request is reproducible.

  2. 2

    Processing stage: match the mode to the constraint. Interactive paths use streaming and a compact context to protect time-to-first-token; bulk, non-urgent work (nightly catalogues, transcript review within 24 hours) goes through the Message Batches API at lower cost; urgent slices get their own real-time path.

  3. 3

    Output stage: validate model output in code before it reaches a downstream system. Use structured outputs or a JSON schema, deterministic cross-field checks (line items summing to totals), and write through a validated, parameterised path rather than letting the model write directly.

  4. 4

    Feedback stage: capture the accountable human's corrections (reviewer edits, nurse overrides, merchandiser rejections, accept/edit/discard) as labelled pairs and feed them into an evaluation set that runs before any prompt, model or retrieval change.

  5. 5

    Add stratified sampled review on top of overrides to catch silent errors the humans did not notice; disagreements from sampling go into the same evaluation set.

  6. 6

    Self-reported confidence, model racing, temperature changes and untested prompt patches are not feedback mechanisms; the model does not learn from its own production outputs.

  7. 7

    Provenance: return citations or source locations with extracted values so verifiers can check quickly and regulators get evidence per figure. Citations plus deterministic checks plus sampled verification beat universal manual re-keying.

  8. 8

    Reproducibility is architectural: version prompts, pin model IDs, persist the exact context and response per case. A database transaction log or prompt history alone cannot explain a specific output months later.

  9. 9

    Do not put a full review call inline on a latency-critical path when the requirement is asynchronous review; lightweight inline guardrails and heavy asynchronous review are different components.

  10. 10

    Cost dashboards and token metrics are operational telemetry, not a feedback loop; a bigger model is not a substitute for measuring what reviewers fix.

Test yourself on Design end-to-end architectures (input → processing → output → feedback loops)

Ten questions, with the answer and explanation after each one.