Systems Life Cycle
Apply life-cycle discipline to LLM systems: define success criteria and evals during design, gate releases on evals, roll out in stages with monitoring and cheap rollback, and maintain the system through input drift and model deprecations.
Key points
- 1
The phases are requirements → design → implementation → test/eval → deployment → operation/monitoring → maintenance. For LLM systems, measurable success criteria and an eval set are outputs of the design phase.
- 2
Without agreed success criteria, "is the new prompt better?" is just opinion. With an eval set, you can compare every change against a baseline.
- 3
Prompt, model and parameter changes go through the same test gate as code. Run the evals, compare with the current version, and promote only if nothing regresses.
- 4
Deployment strategies get progressively riskier. Shadow testing runs the candidate on mirrored live traffic without serving its output, so no users are exposed. A canary sends a small share of traffic to it and compares metrics with a control. A full rollout comes last.
- 5
Judge a canary against all requirements: quality, error rate, latency and cost. If quality holds but
output_tokens, latency or cost jump, pause and fix the cause before widening. - 6
Make rollback a configuration change. Keep the model ID and prompt version in versioned runtime configuration, and keep the last known-good version available.
- 7
Operation means monitoring live traffic: latency, error rates (429 rate_limit_error, 529 overloaded_error, 5xx), token spend, and sampled output quality. Log the
request-idof failures. - 8
Read error codes correctly. A 429 means your organization hit its own limits. A 529 is a temporary overload across the API, so use backoff and graceful degradation rather than rolling back a healthy release. A 400 means the request itself is wrong.
- 9
A model ID refers to a fixed snapshot whose weights never change, although the serving infrastructure can change slightly. Quality drift on an unchanged ID usually comes from prompt, retrieval or input changes, so diagnose it from traces against a known-good baseline.
- 10
Evals drift too. Refresh eval sets regularly from sampled production traffic, including new topics and products, and add every reproduced production failure to the regression set.
- 11
The model lifecycle runs Active → Legacy → Deprecated → Retired. Requests to retired models fail, there is no automatic rerouting, and Anthropic gives at least 60 days' notice for publicly released models.
- 12
Migration plan: list where the model is used (the Console usage export breaks usage down by API key and model), evaluate the recommended replacement on real workloads, including downstream format and parser checks, then roll out in stages well before retirement.
Read the source
Test yourself on Systems Life Cycle
Ten questions, with the answer and explanation after each one.