Apply human-in-the-loop validation strategies
Place human review where risk and uncertainty justify it: approval gates before irreversible or high-impact actions, stratified sampling for volume work, explicit escalation triggers with context handoff, and audit trails that make decisions reconstructable.
Key points
- 1
Calibrate by risk: irreversible or high-value actions get an approval gate before execution; reversible, low-value or already-reviewed steps run automatically with logging and sampling. Approving everything adds delay without reducing risk.
- 2
Anthropic's Usage Policy requires, for high-risk uses affecting individuals (legal, healthcare, insurance underwriting and coverage, credit and lending, employment, housing, academic testing), that a qualified professional review the decision before it is finalized, and that users are told AI is involved. The requirement follows the type of decision, not its dollar value.
- 3
Sampling review: use stratified random samples across the segments that matter (customer, document layout, language, confidence band) and report error rate per stratum; review 100% of low-confidence and high-value items. Convenience samples, complaint-driven review and model self-flagging leave blind spots.
- 4
Escalation triggers should be explicit and testable: request outside policy scope, failed or conflicting tool results, ambiguity, explicit request for a person, actions above a threshold. Sentiment, self-reported confidence and turn count are weak proxies.
- 5
Handoffs must carry context: a structured summary of the conversation, actions already taken and their results, so the human does not start from zero.
- 6
Audit trails capture what was true at decision time: prompt and model versions, retrieved evidence, tool calls, the recommendation with its stated factors, the human decision and who made it. Re-running a later model on appeal does not reconstruct the original decision.
- 7
Agent SDK mechanics:
canUseToolprompts a human only for calls no earlier step resolved; allow rules and permissive modes auto-approve before it. APreToolUsehook can returnask(ordeny/defer) to force a decision on the risky subset, andPostToolUsehooks give an audit log of every executed call. - 8
Permission modes:
defaultprompts through the callback;acceptEditsauto-approves file operations;plannever auto-approves edits;dontAskdenies instead of prompting;bypassPermissionsapproves everything except deny rules, ask rules and hook denies. Subagents inherit the parent's mode. - 9
Watch the override rate: a review step where humans accept 99% of outputs in seconds has become rubber-stamping; redesign the review, do not just raise model accuracy.
- 10
Contestability complements but does not replace prior review: an appeal path is necessary for decisions affecting people, but auto-deciding and letting people appeal is not human-in-the-loop.
- 11
Design review capacity as a constraint: route by risk (denials, low confidence, high value) to 100% review and sample the rest so total effort stays within budget while coverage is where harm concentrates.
Read the source
Test yourself on Apply human-in-the-loop validation strategies
Ten questions, with the answer and explanation after each one.