Study notes · 2.5% of the exam

5.5 Design human review workflows and confidence calibration

Design review workflows that validate accuracy per document type and field, calibrate field-level confidence on labelled data, route low-confidence and ambiguous cases to limited reviewers, and keep sampling automated output to catch drift and novel errors.

Key points

  1. 1

    Aggregate accuracy ("97% overall") can mask poor performance on a specific document type or field (for example 80% on delivery notes, or on total_amount). Automation is applied per segment, so validate accuracy by document type and by field before reducing human review.

  2. 2

    JSON schema validity is not accuracy: a wrong value in the right type passes validation every time. Schema checks belong in the pipeline but are not evidence of correctness.

  3. 3

    Have the model output field-level confidence scores, then calibrate review thresholds against a labelled validation set: measure the actual error rate at each confidence level, per field, and choose thresholds that meet the target error rate. Thresholds picked because they "sound reasonable", or set from reviewer capacity, are uncalibrated.

  4. 4

    Calibration is only valid for the segments in the validation set. A new supplier, layout or document type is a new segment: route it to human review until it has its own labelled data, then recalibrate per segment. Raising a global threshold floods review with correct extractions from well-calibrated segments.

  5. 5

    Route to human review: extractions with low calibrated confidence on important fields, and documents whose source content is ambiguous or contradictory (two different totals, an illegible date), where no single correct extraction exists for the model to be confident about.

  6. 6

    Prioritise limited reviewer capacity by error likelihood, not by document length, customer tier, queue position or the model's uncalibrated overall self-assessment.

  7. 7

    Once high-confidence extractions are auto-accepted they are otherwise unobserved. Implement ongoing stratified random sampling of that stream across document types and sources, reviewed by humans, to measure the true high-confidence error rate and detect novel error patterns such as a changed invoice layout.

  8. 8

    Low-confidence review cannot see high-confidence errors; downstream complaints are slow and incomplete; a second-model comparison is a supplement that does not measure the human-verified error rate and can share the same misreading.

  9. 9

    Before cutting review from 100% to a fraction: per-segment accuracy analysis, calibrated field-level thresholds, and a stratified sampling programme for the auto-accepted output. Prompt tweaks such as "be accurate" or a bigger model are not gates.

Test yourself on 5.5 Design human review workflows and confidence calibration

Ten questions, with the answer and explanation after each one.