Study notes · 3.3% of the exam

4.4 Implement validation, retry, and feedback loops for extraction quality

Validate extracted data in code, retry with the document plus the specific error when the failure is correctable, recognise when retries cannot help, and instrument findings so false-positive patterns can be analysed.

Key points

  1. 1

    Retry-with-error-feedback: the follow-up request must include the original document, the failed extraction, and the specific validation error ("line items sum to 1,240.00 but total is 1,420.00"). Resending the same request, raising temperature, or appending "be more careful" gives the model nothing to correct with.

  2. 2

    Never retry with only the failed JSON and the error message. Without the source document the model can only edit numbers until the rule passes, which is fabrication.

  3. 3

    Distinguish semantic validation errors (values do not sum, wrong field placement, date before order date) from schema syntax errors. Tool use with a strict schema eliminates the second class; the first class still needs code validation.

  4. 4

    Retries fix format and structural errors (a date not in ISO format, an array returned as a comma-joined string) because the information is in the document. Retries cannot recover information that is absent from the source, for example a policy number that lives only in an attachment that was never sent.

  5. 5

    When a field is repeatedly null because the source lacks it, stop retrying: supply the missing input where available, otherwise make the field nullable and route the record to a human queue. Raising the retry limit, strengthening the error message, or asking the model to infer the value are the wrong responses.

  6. 6

    Bound the loop: a fixed retry budget (for example two retries), a stop condition when the model reports the value is not present, and a human review route for unresolved records. Loops that retry until valid, regardless of cause, waste money and eventually drop documents.

  7. 7

    Self-correction flows: extract calculated_total (the model's own sum of the rows) alongside stated_total (the figure printed in the document) so validation can tell an extraction mistake from a source document that disagrees with itself.

  8. 8

    Add a conflict_detected boolean (with a detail string) that the model sets when the source gives inconsistent values for the same figure. Route those records to review instead of retrying them or silently replacing the stated value with the computed one.

  9. 9

    Never instruct the model to overwrite stated figures so validation passes; that destroys the evidence of a discrepancy and reports numbers the document never contained.

  10. 10

    Feedback loops for review quality: add a detected_pattern field to every structured finding naming the code construct that triggered it. Aggregating developer dismissals per pattern shows exactly which criteria produce false positives, which is far more actionable than free-text dismissal reasons or raw logs.

  11. 11

    Self-reported confidence is not a substitute for pattern tracking or validation; filtering findings by confidence hides them without explaining them.

  12. 12

    Common distractors: larger context window (irrelevant when the page is not in the request), bigger model, more retries, or skipping code validation because the schema is strict.

Test yourself on 4.4 Implement validation, retry, and feedback loops for extraction quality

Ten questions, with the answer and explanation after each one.