Study notes · 2.7% of the exam

Define evaluation metrics (accuracy, latency, cost, safety, security)

Define success criteria for a Claude solution as specific, measurable, achievable and relevant metrics across accuracy, latency, cost, safety and security, each with a pass threshold that can gate a release.

Key points

  1. 1

    Good criteria are specific ("accurate sentiment classification", not "good performance"), measurable (a number on a defined dataset or a consistently applied qualitative scale), achievable (grounded in benchmarks, prior experiments or expert knowledge) and relevant (aligned with the application's purpose and users).

  2. 2

    Most use cases need multidimensional evaluation. The docs' example: on a held-out set of 10,000 posts, F1 at least 0.85, 99.5% of outputs non-toxic, 90% of errors merely inconvenient, and 95% of responses under 200 ms. Common criteria: task fidelity, consistency, relevance and coherence, tone and style, privacy preservation, context utilization, latency, price.

  3. 3

    Express latency as percentiles with thresholds (p50 for the typical request, p95/p99 for the tail) measured under realistic load. A mean or median can meet its target while a minority of requests are unusably slow.

  4. 4

    Define cost per unit of business outcome (cost per successfully completed task), including every call, retry, verification step and tool invocation. Cost per API call or per attempt can improve while total spend and true efficiency get worse.

  5. 5

    Safety and security are quantified like everything else: a defined test set, a countable event and a threshold. Examples: share of responses leaking another customer's data (target 0), fraction of injected documents whose embedded instructions the agent follows, toxicity-flag rate, policy-violation rate on a red-team set.

  6. 6

    Raw refusal counts, self-reported model confidence, reviewer impressions and "control is enabled" flags are not metrics of outcomes; they are common distractors.

  7. 7

    For RAG and document tasks add a groundedness or citation-support metric (every claim traces to a retrieved passage), separate from answer accuracy: an answer can match the reference and still add an unsupported condition.

  8. 8

    Gate releases on each metric independently with its own threshold. Use composite scores only to rank candidates that already pass every gate; a weighted average lets a gain on one axis hide an SLA breach on another.

  9. 9

    Write the criteria before building the eval set; they determine what the dataset must contain and which grader fits each criterion.

  10. 10

    Set thresholds from the business context: what error rate is acceptable, what latency the user experience tolerates, what the SLA promises, what the budget per task is.

  11. 11

    Report metrics by slice (language, customer segment, document type, severity) as well as in aggregate, so a subgroup regression is visible.

Test yourself on Define evaluation metrics (accuracy, latency, cost, safety, security)

Ten questions, with the answer and explanation after each one.