Study notes · 2.6% of the exam

Output Handling

Produce, validate, and consume Claude output safely: structured outputs and strict tools, response validation, defensive parsing, and appropriate skepticism toward confident output.

Key points

  1. 1

    Structured outputs constrain the response to a JSON schema through output_config.format. SDK helpers such as messages.parse() validate the result against the schema.

  2. 2

    Strict tool use (strict: true on a tool definition) guarantees tool inputs match the tool's schema. JSON outputs cover the response, strict tools cover tool arguments, and the two can be combined in one request.

  3. 3

    For classification, constrain the label with an enum in structured outputs or a tool, rather than parsing free text.

  4. 4

    Schemas guarantee shape, not truth. Validate business rules in code (totals, date order, allowed combinations) before acting.

  5. 5

    When validation fails, re-prompt with the specific error, cap the retries, and escalate persistent failures to human review. Never silently "fix" or accept bad data.

  6. 6

    Always check stop_reason before parsing. With max_tokens the output may be truncated, and with refusal it may not match the schema, even with structured outputs.

  7. 7

    Defensive parsing: wrap parsing in error handling, retry or route failures, and parse tool inputs with a JSON parser, never raw string matching.

  8. 8

    Some JSON Schema features, such as numeric minimum/maximum and string length limits, are not supported by structured outputs. Enforce those rules in code.

  9. 9

    Structured outputs cannot be combined with citations or with assistant prefill. Recent models reject prefill anyway, so use structured outputs instead.

  10. 10

    Fluent, confident output is not evidence of correctness. Self-reported confidence scores and lower temperature do not prevent hallucinations.

  11. 11

    To reduce hallucinations: allow Claude to say "I don't know", ground answers in provided documents with direct quotes, and verify each claim against its source. Best-of-N comparisons can flag inconsistencies.

  12. 12

    Where a model's claim can be checked (tests passed, file written, value present in source), have the harness verify it independently before taking automated action.

Test yourself on Output Handling

Ten questions, with the answer and explanation after each one.