Identify risks, limitations, and failure modes of LLM systems
Recognise the characteristic ways LLM systems fail (hallucination, injection through tool results, over-reliance, compounding errors in loops, non-determinism, silent drift after model or data changes) and pick mitigations that address the mechanism rather than the symptom.
Key points
- 1
Hallucination: reduce it by allowing "I don't know", grounding in direct quotes from long documents, requiring citations and retracting uncited claims, restricting the model to provided knowledge, and using chain-of-thought or best-of-N comparison to reveal inconsistencies. These reduce but never eliminate hallucination; validate critical information.
- 2
Confident-but-wrong answers that start right after a document refresh, with model and latency unchanged, point to the retrieval or indexing layer, not the model (CCAR-P sample 3).
- 3
Compounding errors: in an N-step agent loop per-step error multiplies (twelve steps at 95% gives about 54% end to end). Mitigate with programmatic validation between steps, checkpoints, short loops, and turn limits; more steps or a bigger model do not change the arithmetic.
- 4
Non-determinism: outputs vary run to run even at low temperature. Test with labelled evaluation sets and statistical pass thresholds, assert on structured fields rather than prose, and never rely on single-sample exact-match tests, retries until green, or cached fixtures.
- 5
Silent drift: model upgrades, prompt edits, retrieval changes and data refreshes change behaviour without errors. Pin versions, gate every change behind a regression eval, and monitor quality metrics; error rate and latency dashboards stay flat during a quality regression.
- 6
Over-reliance (automation bias): when reviewers approve almost everything in seconds the human check has stopped working. Require active verification on a stratified sample plus all low-confidence or high-value items and track override rate as a health metric.
- 7
Prompt injection via tool results and retrieved content is a failure mode of the architecture, not of the model alone; contain it with least privilege and gated actions so a fooled model cannot cause irreversible harm.
- 8
Irreversible actions (payments, deletions, outbound messages) must be gated by deterministic checks such as resolving entities to canonical IDs and validating against a system of record before the action executes.
- 9
Sampling parameters are not safety controls: temperature affects repeatability, not groundedness or resistance to manipulation.
- 10
Model self-reported confidence is not calibrated; do not use it as the sole trigger for escalation or as a substitute for evaluation.
- 11
Treat model and data changes like code deploys: change-triggered evaluation, canaries and rollback plans, with the eval suite covering the segments that matter to the business.
- 12
Common distractors: switch to a larger model, lower temperature, ask the model to be careful, add more context, or retry; each fails to address the underlying mechanism.
Read the source
Test yourself on Identify risks, limitations, and failure modes of LLM systems
Ten questions, with the answer and explanation after each one.