Design end-to-end architectures (input → processing → output → feedback loops)
Design the full flow, input → processing → output → feedback, choosing the processing mode per workload constraint, gating output deterministically, and closing the loop so human corrections and evaluation results drive iteration.
Key points
- 1
Input stage: normalise every channel (email, web form, scans, API) into a canonical record at ingestion so downstream prompts see one shape and every request is reproducible.
- 2
Processing stage: match the mode to the constraint. Interactive paths use streaming and a compact context to protect time-to-first-token; bulk, non-urgent work (nightly catalogues, transcript review within 24 hours) goes through the Message Batches API at lower cost; urgent slices get their own real-time path.
- 3
Output stage: validate model output in code before it reaches a downstream system. Use structured outputs or a JSON schema, deterministic cross-field checks (line items summing to totals), and write through a validated, parameterised path rather than letting the model write directly.
- 4
Feedback stage: capture the accountable human's corrections (reviewer edits, nurse overrides, merchandiser rejections, accept/edit/discard) as labelled pairs and feed them into an evaluation set that runs before any prompt, model or retrieval change.
- 5
Add stratified sampled review on top of overrides to catch silent errors the humans did not notice; disagreements from sampling go into the same evaluation set.
- 6
Self-reported confidence, model racing, temperature changes and untested prompt patches are not feedback mechanisms; the model does not learn from its own production outputs.
- 7
Provenance: return citations or source locations with extracted values so verifiers can check quickly and regulators get evidence per figure. Citations plus deterministic checks plus sampled verification beat universal manual re-keying.
- 8
Reproducibility is architectural: version prompts, pin model IDs, persist the exact context and response per case. A database transaction log or prompt history alone cannot explain a specific output months later.
- 9
Do not put a full review call inline on a latency-critical path when the requirement is asynchronous review; lightweight inline guardrails and heavy asynchronous review are different components.
- 10
Cost dashboards and token metrics are operational telemetry, not a feedback loop; a bigger model is not a substitute for measuring what reviewers fix.
Read the source
Test yourself on Design end-to-end architectures (input → processing → output → feedback loops)
Ten questions, with the answer and explanation after each one.