Study notes · 2.8% of the exam

Apply human-in-the-loop validation strategies

Place human review where risk and uncertainty justify it: approval gates before irreversible or high-impact actions, stratified sampling for volume work, explicit escalation triggers with context handoff, and audit trails that make decisions reconstructable.

Key points

  1. 1

    Calibrate by risk: irreversible or high-value actions get an approval gate before execution; reversible, low-value or already-reviewed steps run automatically with logging and sampling. Approving everything adds delay without reducing risk.

  2. 2

    Anthropic's Usage Policy requires, for high-risk uses affecting individuals (legal, healthcare, insurance underwriting and coverage, credit and lending, employment, housing, academic testing), that a qualified professional review the decision before it is finalized, and that users are told AI is involved. The requirement follows the type of decision, not its dollar value.

  3. 3

    Sampling review: use stratified random samples across the segments that matter (customer, document layout, language, confidence band) and report error rate per stratum; review 100% of low-confidence and high-value items. Convenience samples, complaint-driven review and model self-flagging leave blind spots.

  4. 4

    Escalation triggers should be explicit and testable: request outside policy scope, failed or conflicting tool results, ambiguity, explicit request for a person, actions above a threshold. Sentiment, self-reported confidence and turn count are weak proxies.

  5. 5

    Handoffs must carry context: a structured summary of the conversation, actions already taken and their results, so the human does not start from zero.

  6. 6

    Audit trails capture what was true at decision time: prompt and model versions, retrieved evidence, tool calls, the recommendation with its stated factors, the human decision and who made it. Re-running a later model on appeal does not reconstruct the original decision.

  7. 7

    Agent SDK mechanics: canUseTool prompts a human only for calls no earlier step resolved; allow rules and permissive modes auto-approve before it. A PreToolUse hook can return ask (or deny/defer) to force a decision on the risky subset, and PostToolUse hooks give an audit log of every executed call.

  8. 8

    Permission modes: default prompts through the callback; acceptEdits auto-approves file operations; plan never auto-approves edits; dontAsk denies instead of prompting; bypassPermissions approves everything except deny rules, ask rules and hook denies. Subagents inherit the parent's mode.

  9. 9

    Watch the override rate: a review step where humans accept 99% of outputs in seconds has become rubber-stamping; redesign the review, do not just raise model accuracy.

  10. 10

    Contestability complements but does not replace prior review: an appeal path is necessary for decisions affecting people, but auto-deciding and letting people appeal is not human-in-the-loop.

  11. 11

    Design review capacity as a constraint: route by risk (denials, low confidence, high value) to 100% review and sample the rest so total effort stays within budget while coverage is where harm concentrates.

Test yourself on Apply human-in-the-loop validation strategies

Ten questions, with the answer and explanation after each one.