Evaluate Claude-generated outputs for accuracy and completeness
Judge a Claude output on two axes before using it: completeness against the original request and accuracy against the primary source, without being swayed by fluency, formatting or Claude's own assessment.
Key points
- 1
Evaluate on two separate axes. Completeness: did the output cover everything the request asked for (every section, every metric, every entity)? Accuracy: does each claim match the source document or the real world? A polished output can fail either one.
- 2
Turn the original request into a checklist before you read the output (8 business units x 2 metrics = 16 cells; three policy requirements; ten obligations). Tick each item off against the source. This catches silent omissions and silent substitutions (6 months quietly becoming 12).
- 3
A claim that data is "not reported" or "not in the document" is itself a claim to verify. In long uploads Claude may miss a page rather than the data being absent; check the source before accepting a gap.
- 4
Fluency, consistent formatting, plausible-looking numbers and a sensible length are not evidence of correctness. They are what makes a wrong output dangerous.
- 5
Claude's own review of its output in the same conversation ("Is anything missing?", "Are you sure?", "Rate your confidence") is not an independent check; it tends to confirm what it just wrote. The exam guide is explicit that self-rated confidence is not an accuracy signal.
- 6
Spot checks estimate the error rate of what is present; they cannot detect what was left out. When the output is short enough to verify fully and the consequences are material, verify every item and scan the source (headings, table of contents) for topics the output never mentions.
- 7
Look for internal inconsistencies: percentages that cannot add up given how the data was collected, a respondent count that changes between paragraphs, a table that disagrees with the narrative. Any contradiction means at least one statement is wrong and the method is suspect.
- 8
For numbers, make the computation visible: ask Claude which rows, columns and years it used and how it calculated the result, or have it compute with its data analysis tool so counts are calculated rather than estimated. Then reconcile discrepancies before any figure is used.
- 9
Extra details that were not in the request (accrual rules, notice periods, industry norms) often come from Claude's general knowledge rather than your document. Ask Claude to point to the passage each one came from and remove anything it cannot source.
- 10
Re-running the request in a fresh chat checks consistency, not completeness: two runs can share the same omission or the same misreading. Use it as a supplement, never as the check against the source.
- 11
Anti-patterns the exam uses as distractors: accept because it reads well; accept because a sample of three was correct; ask Claude to list what it missed; check a total reconciles and infer the components are right; rescale or patch numbers whose derivation is unknown.
Read the source
Test yourself on Evaluate Claude-generated outputs for accuracy and completeness
Ten questions, with the answer and explanation after each one.