Study notes · 2.7% of the exam

Conduct A/B testing and iterative improvements

Compare prompts, models and configurations rigorously, offline on a shared eval set and online on a controlled traffic split, and run a disciplined iterate-measure loop with guardrails.

Key points

  1. 1

    Offline first: run the candidate and the control on the same versioned eval set, with the same model and the same metrics, so only the change under test differs. Scores from different datasets or different dates are not comparable.

  2. 2

    The Console Evaluation tool's side-by-side comparison is designed for this kind of prompt-versus-prompt check; re-run the suite after each prompt edit.

  3. 3

    Online A/B tests: assign randomly and keep the assignment stable per user or session (never per message), fix the sample size, duration and stop rule before starting, name one primary metric, and define guardrail metrics (p95 latency, cost per task, escalation or error rate, safety rate) with limits.

  4. 4

    Common invalid designs: time-of-day or day-of-week splits (confounded by traffic mix), stopping as soon as one arm pulls ahead (inflates false positives), tiny samples that cannot distinguish a small lift from noise, week-over-week comparisons after a full rollout.

  5. 5

    A breached guardrail is a failed test even if the primary metric won. Do not raise the threshold after seeing the result; diagnose the cause and re-test a revised candidate.

  6. 6

    Analyse by slice (language, customer segment, document type). A pre-registered slice guardrail stops an aggregate win from shipping harm to a subgroup; when routing by slice is reliable and permitted, ship where the gain was validated and keep the control elsewhere.

  7. 7

    Never ship an untested fix on the strength of a hypothesis; every revision goes through the same offline eval and, where needed, the same traffic test.

  8. 8

    Iterative improvement discipline: run the full regression suite on every change, gate on no net regression, and diagnose why a fix broke other cases rather than stacking more rules onto the prompt.

  9. 9

    Keep capability evals (what to improve next) separate from regression evals (what must not break); a change is finished only when both are satisfied.

  10. 10

    Evals shape how fast you can adopt new models: with a suite in place a model swap can be validated in days, without one it takes weeks of manual testing.

  11. 11

    Production monitoring, user feedback, A/B testing, transcript review and human evaluation are complementary layers; no single layer catches every issue.

Test yourself on Conduct A/B testing and iterative improvements

Ten questions, with the answer and explanation after each one.