Study notes · 2.8% of the exam

Support lifecycle phases (discovery, design, handoff, monitoring, iteration)

Support each lifecycle phase (discovery, design, pilot, handoff, monitoring, iteration) with the right evidence and gates, and treat model deprecation and prompt changes as managed change.

Key points

  1. 1

    The lifecycle runs discovery, design, pilot, handoff, monitoring, iteration. Each transition should be gated on evidence appropriate to the phase, not on enthusiasm, feature counts or self-reported confidence.

  2. 2

    Design to pilot gate: agreed success criteria, an evaluation set, and a named pilot cohort. Pilot to handoff gate: the criteria met on real traffic, runbooks, and named owners. Monitoring to iteration: explicit measured thresholds (accuracy drift, escalation-rate rise, cost departure) that open an iteration cycle.

  3. 3

    A demo shows feasibility on curated inputs; a pilot measures the agreed success criteria on real traffic with a feedback loop before launch. Skipping the pilot because the demo was convincing (treating a demo as a launch) is the most common lifecycle failure the exam tests.

  4. 4

    Model deprecation is change management, not a string swap: run the existing evaluation suite against the successor model, compare, record the decision in an ADR, roll out in stages with monitoring and a rollback path, and communicate to stakeholders. Do not swap on retirement day assuming prompts behave identically, and do not refuse to migrate.

  5. 5

    Prompt versioning is change management too: prompts are versioned, gated by the evaluation suite before deployment, rolled out in stages, and rollable; direct edits to the production prompt are prohibited. A tone tweak that silently drops extraction accuracy by 8 points is what unmanaged prompt edits look like.

  6. 6

    Monitoring feeds iteration through sampled human review on a fixed cadence compared against the accepted baseline, plus explicit thresholds that trigger work. Add production failures to the evaluation set so each iteration is tested against what actually went wrong.

  7. 7

    Weak monitoring signals: average output length as a quality proxy, number of prompt edits as a health indicator, and opening an iteration whenever the vendor releases a new model. A new model is a candidate to evaluate, not a trigger.

  8. 8

    Iterate from the error taxonomy: address the largest measured cause first (an ambiguous policy is a stakeholder and data fix, not a model fix), re-measure, and only then consider more complex architecture. Start simple and add complexity only when it demonstrably improves outcomes; fixing the cheapest category first optimises for visible activity, not impact.

  9. 9

    A larger model tier is never a change-control mechanism and rarely a substitute for fixing the process, data or policy that produced the errors.

  10. 10

    Handoff requires ownership. Without a named owner for prompts, the evaluation suite and deprecation response, a system that passed its pilot degrades quietly.

Test yourself on Support lifecycle phases (discovery, design, handoff, monitoring, iteration)

Ten questions, with the answer and explanation after each one.