4.2 Apply few-shot prompting to improve output consistency and quality
Use a small number of targeted examples to make output consistently formatted, to demonstrate how ambiguous cases are decided, and to teach a distinction the model can generalise, especially where detailed instructions alone have failed.
Key points
- 1
Few-shot examples are the most effective technique for consistently formatted, actionable output when detailed instructions still produce inconsistent results. Show the exact shape (location, issue, severity, suggested fix) rather than describing it.
- 2
Two to four targeted examples usually beat a long list. Pick cases at the decision boundary: ambiguous tool selection, a borderline severity, an acceptable pattern that resembles a bug, a real bug that resembles an acceptable pattern.
- 3
Include the reasoning. An example that shows why one action was chosen over a plausible alternative lets the model generalise the judgement to novel inputs; a bare input/output pair teaches only that one case.
- 4
For ambiguous tool selection ("my order was damaged and I was charged twice"), examples with reasoning fix what improved descriptions alone leave inconsistent. Forcing
tool_choice: anyguarantees a tool call but not the right one. - 5
For branch-level test coverage gaps, worked examples (function, existing tests, the unexercised branch, the reasoning that found it, the closing test) generalise where an enumerated list of branch patterns catches only what it names.
- 6
To reduce false positives on sanctioned patterns, pair an acceptable pattern annotated as acceptable with a superficially similar genuine issue annotated as a real finding. Contrasting examples teach the boundary; allowlists and skip rules only cover listed cases.
- 7
Examples must be in the same format the output must take. Prose examples for a structured output make consistency worse, not better.
- 8
Only negative examples ("do not flag this") push the model toward under-reporting; always include positive examples of correct findings too.
- 9
In extraction, few-shot examples reduce hallucination and empty required fields when documents vary in structure: show correct extraction from a paper whose methods are embedded in the results, from inline citations versus a bibliography, and from informal measurements such as a sample size in a caption.
- 10
Diversity beats volume. Fifteen examples from the same template the model already handles add tokens and teach nothing; one example of each failing structure does.
- 11
Forbidding null ("this field must never be null") does not teach the model where to look; it invites fabrication. Demonstrate the correct extraction instead, and keep genuinely absent values nullable.
- 12
Common traps offered as distractors: adding "MUST", raising
max_tokens, changing temperature, switching model tier, or regex post-processing of inconsistent output. None of them demonstrates the desired behaviour. - 13
Wrap examples in clear delimiters (for example
<example>tags) and keep them relevant to the real task so they are not mistaken for instructions or for the input.
Read the source
Test yourself on 4.2 Apply few-shot prompting to improve output consistency and quality
Ten questions, with the answer and explanation after each one.