Study notes · 3.3% of the exam

4.3 Enforce structured output using tool use and JSON schemas

Get schema-compliant structured output through tool use with JSON schemas (or structured outputs), choose the right `tool_choice`, and design schemas (nullable fields, extensible enums, normalisation rules) so the model is never forced to fabricate.

Key points

  1. 1

    Tool use with a JSON schema is the most reliable way to get structured output: define an extraction tool whose input_schema is the record schema, force the call, and read the record from the tool_use block's input. This eliminates JSON syntax errors, preambles and markdown fences that break JSON.parse on text responses.

  2. 2

    tool_choice has three relevant settings: {"type": "auto"} lets the model answer in text instead of calling a tool; {"type": "any"} forces it to call one of the provided tools but lets it choose which; {"type": "tool", "name": "extract_metadata"} forces that specific tool. (none prevents tool calls.)

  3. 3

    Use any when several extraction schemas exist and the document type is unknown but structured output is mandatory. Use a forced named tool when a particular extraction must run first, before a dependent enrichment step, and let the application sequence the follow-up.

  4. 4

    Prompt instructions such as "always call extract_metadata first" make an ordering likely; a forced tool_choice makes it certain. When the ordering or the structured response must always hold, use the API-level control.

  5. 5

    Strict schemas (strict: true on the tool, or output_config.format structured outputs) guarantee syntax and types, not semantics. Line items that do not sum to the total, or a unit price placed in the quantity field, are valid JSON and pass the schema; catch them with semantic validation in code.

  6. 6

    Numeric range constraints (minimum, maximum), string length limits and recursive schemas are not supported by strict structured output, and even in an ordinary validator they cannot express cross-field rules such as "rows must sum to total".

  7. 7

    Required fields force the model to produce a value even when the source has none, which is where fabricated PO numbers and invented dates come from. Make a field nullable (anyOf string/null) whenever the source may legitimately lack it, and describe when null is expected.

  8. 8

    Do not make everything optional: fields that are always present should stay required so validation still catches missing data. A prompt line saying "never guess" cannot override a schema that demands a value.

  9. 9

    For categorical fields, use an enum extended with "other" plus a detail string (and "unclear" where ambiguity is common) so unusual cases are captured without breaking the controlled vocabulary. Free-text categories destroy downstream consistency; a closed enum forces wrong assignments.

  10. 10

    Include format normalisation rules in the prompt alongside the schema: locale for numeric dates (day/month/year), the output format (ISO 8601), units, and the reference date for relative expressions. The schema constrains the shape of the value; the rules govern its interpretation.

  11. 11

    Carry provenance for ambiguous values: a nullable source-text field and an ambiguity flag let downstream validation route uncertain records to review instead of accepting a confident but arbitrary reading.

  12. 12

    A self-attestation field (is_valid: true) is filled by the same model that made the mistake and is not validation. Switching between tool use and output_config format changes nothing about semantic guarantees.

  13. 13

    Newer models may reject forced tool selection (any or a named tool); the documented alternative is auto with strict: true tools or structured outputs. The exam tests the auto / any / named-tool distinction as described in the exam guide.

Test yourself on 4.3 Enforce structured output using tool use and JSON schemas

Ten questions, with the answer and explanation after each one.