Study notes · 2.8% of the exam

Identify risks, limitations, and failure modes of LLM systems

Recognise the characteristic ways LLM systems fail (hallucination, injection through tool results, over-reliance, compounding errors in loops, non-determinism, silent drift after model or data changes) and pick mitigations that address the mechanism rather than the symptom.

Key points

  1. 1

    Hallucination: reduce it by allowing "I don't know", grounding in direct quotes from long documents, requiring citations and retracting uncited claims, restricting the model to provided knowledge, and using chain-of-thought or best-of-N comparison to reveal inconsistencies. These reduce but never eliminate hallucination; validate critical information.

  2. 2

    Confident-but-wrong answers that start right after a document refresh, with model and latency unchanged, point to the retrieval or indexing layer, not the model (CCAR-P sample 3).

  3. 3

    Compounding errors: in an N-step agent loop per-step error multiplies (twelve steps at 95% gives about 54% end to end). Mitigate with programmatic validation between steps, checkpoints, short loops, and turn limits; more steps or a bigger model do not change the arithmetic.

  4. 4

    Non-determinism: outputs vary run to run even at low temperature. Test with labelled evaluation sets and statistical pass thresholds, assert on structured fields rather than prose, and never rely on single-sample exact-match tests, retries until green, or cached fixtures.

  5. 5

    Silent drift: model upgrades, prompt edits, retrieval changes and data refreshes change behaviour without errors. Pin versions, gate every change behind a regression eval, and monitor quality metrics; error rate and latency dashboards stay flat during a quality regression.

  6. 6

    Over-reliance (automation bias): when reviewers approve almost everything in seconds the human check has stopped working. Require active verification on a stratified sample plus all low-confidence or high-value items and track override rate as a health metric.

  7. 7

    Prompt injection via tool results and retrieved content is a failure mode of the architecture, not of the model alone; contain it with least privilege and gated actions so a fooled model cannot cause irreversible harm.

  8. 8

    Irreversible actions (payments, deletions, outbound messages) must be gated by deterministic checks such as resolving entities to canonical IDs and validating against a system of record before the action executes.

  9. 9

    Sampling parameters are not safety controls: temperature affects repeatability, not groundedness or resistance to manipulation.

  10. 10

    Model self-reported confidence is not calibrated; do not use it as the sole trigger for escalation or as a substitute for evaluation.

  11. 11

    Treat model and data changes like code deploys: change-triggered evaluation, canaries and rollback plans, with the eval suite covering the segments that matter to the business.

  12. 12

    Common distractors: switch to a larger model, lower temperature, ask the model to be careful, add more context, or retry; each fails to address the underlying mechanism.

Test yourself on Identify risks, limitations, and failure modes of LLM systems

Ten questions, with the answer and explanation after each one.