Identify hallucinations, inconsistencies, and biases in responses
Recognize the signatures of hallucinated content, internal inconsistency and skewed framing in a Claude response, and know that specificity, confidence and consistency are not evidence of truth.
Key points
- 1
A hallucination is content that is factually wrong or not grounded in the provided context but presented as if it were. The help centre warns that Claude can "display quotes that may look authoritative or sound convincing, but are not grounded in fact" and can "write things that might look correct but are very mistaken".
- 2
The most convincing hallucinations are the most specific: named court cases with reporter citations, statute section numbers, report titles, statistics to one decimal place, quotations attributed to named executives. Precision is a red flag, not a credential.
- 3
Claude's training data has a cutoff, and without web search or an uploaded document it cannot know the contents of a recent report, regulation or announcement. Confident specifics about recent events produced with no source available are presumptively invented.
- 4
Claude can also hallucinate its own capabilities: it may produce links that do not work, or claim to have sent an email or created an external file when it has no such integration. Treat claims of external actions the same way as factual claims: verify.
- 5
Internal inconsistency is a hallucination signal: the market grows 12% in one paragraph and contracts in another; a table says 6 months and the conclusion says 4; percentages that cannot sum to 100% for multi-select data. At least one statement is wrong, and the whole output's method is suspect.
- 6
Skewed framing shows up as evaluative language that the numbers do not support ("aggressive" versus "competitive" pricing when the cheaper vendor is the one called aggressive), or as an output that rewards attributes not in the stated criteria (prestige of university or employer in a ranking that was supposed to be about frontline experience).
- 7
Bias survives anonymization. Removing names does not remove proxies such as institution prestige, employer size, postcode or writing style. Diagnose bias by comparing what the output rewards with what was asked for.
- 8
The remedy for skewed outputs is explicit criteria, traceable evidence and a human decision: re-run with a rubric limited to the stated criteria, require Claude to quote the evidence behind each score, and keep the shortlist or recommendation decision with a person. Anthropic does not endorse using models to make automated high-stakes decisions about people.
- 9
Consistency is not correctness. A repeatable ranking shows the skew is systematic; agreement between two runs shows the same error can be made twice; majority voting across three runs does not verify a single figure against the source.
- 10
Signs of a well-grounded output are the opposite of red flags: hedged language, a stated "gaps and limitations" section, a cell left blank with a note that the data could not be found, and an explicit "I don't have enough information". Anthropic's guidance encourages giving Claude permission to say it does not know.
- 11
Asking Claude whether it hallucinated or was biased, and accepting a no, is not detection. The model cannot reliably audit its own weighting, and a reassuring answer is not evidence.
- 12
When you find a hallucination or bias, use the thumbs-down feedback on the response; on Team and Enterprise plans feedback settings are managed by the admin. Feedback helps Anthropic but does not fix the output in front of you.
Read the source
- Claude is providing incorrect or misleading responses. What's going on?
- Claude is producing links that don't work and falsely claiming that it has sent emails or produced external documents
- How up to date is Claude's training data?
- Evaluating and mitigating discrimination in language model decisions
- Reduce hallucinations
Test yourself on Identify hallucinations, inconsistencies, and biases in responses
Ten questions, with the answer and explanation after each one.