Design evaluation datasets and test frameworks using mixed methodologies
Build and maintain evaluation datasets and test frameworks that mirror real traffic, cover edge and adversarial cases, and mix code-based, LLM-based and human grading appropriately.
Key points
- 1
Eval design principles from the docs: be task-specific (mirror the real-world task distribution, including edge cases), automate when possible (structure tasks so they can be graded by code, string match or an LLM), and prioritize volume over quality (more automatically graded cases beat a handful of hand-graded ones).
- 2
Start early and small: 20 to 50 tasks drawn from real failures, bug reports and support tickets is a strong first eval. Grow it from production traffic and add synthetic edge cases and adversarial inputs over time.
- 3
A good task is one where two domain experts would independently reach the same pass/fail verdict. Ambiguous cases produce noisy scores.
- 4
Vocabulary from the evals post: a task (input plus success criteria), a trial (one attempt), a grader (the scoring logic, possibly several assertions), a transcript (the full record of the trial) and the outcome (the final state of the environment).
- 5
Grader trade-offs: code-based grading (exact match, regex, unit tests, state checks) is fastest, cheapest and most reliable but brittle to valid variation and lacks nuance; LLM-based grading (rubric scores, natural-language assertions, pairwise comparison) is fast, flexible and scalable but non-deterministic and must be tested for reliability; human grading is the most flexible and highest quality but slow and expensive, so use it for calibration and spot checks rather than the whole suite.
- 6
Tips for LLM-based grading: give the judge a detailed, clear rubric; ask for empirical or specific outputs ("correct"/"incorrect", a defined 1 to 5 scale) rather than free-form opinions; encourage the judge to reason before deciding; use a different model from the one being evaluated; and calibrate against human labels on a sample.
- 7
Mix methods by output: deterministic checks for fixed answers and schema validity, rubric-based LLM grading for free-text quality, human review of a stratified sample for calibration.
- 8
Balance the set: include cases where the behaviour should occur and where it should not (refusals, escalations, out-of-scope questions), avoid class imbalance that hides rare high-consequence failures, and report pass rate per slice with slice-level thresholds.
- 9
Eval hygiene: hold out a test split that is never used while iterating on the prompt, version the dataset alongside the prompt so scores are comparable across time, keep passing cases as the regression suite, and read transcripts regularly to check the graders are fair.
- 10
Capability evals target hard tasks and start with low pass rates; regression evals should sit near 100% and catch backsliding. Both are needed.
- 11
Non-determinism: run multiple trials per task. pass@k (at least one of k succeeds) suits "shots on goal"; pass^k (all k succeed) measures the consistency customer-facing agents need. Isolate trials in clean environments so shared state does not contaminate results.
- 12
For agents, grade the outcome (did the reservation get created, does the test pass) and only the transcript assertions that matter (no destructive action without confirmation); rigid step-sequence checks penalize valid alternative paths.
- 13
Watch for saturation (the agent passes nearly everything solvable, so the eval gives no signal) and overfitting (scores rise while users see no improvement, or graders reward loopholes rather than the task).
- 14
Tooling: the Console Evaluation tool lets you add test cases by hand, generate them with Claude, import a CSV, run a prompt with
{{variables}}across the set, and compare prompt versions side by side; it is a convenient place to start before building a code-based harness.
Read the source
Test yourself on Design evaluation datasets and test frameworks using mixed methodologies
Ten questions, with the answer and explanation after each one.