4.6 Design multi-instance and multi-pass review architectures
Design review architectures that use an independent Claude instance rather than self-review, split large reviews into per-file passes plus a cross-file integration pass, and use verification passes with calibrated confidence to route findings.
Key points
- 1
Self-review limitation: a model that generated code or a report retains the reasoning that produced it, so when asked to review its own output in the same session it tends to confirm its decisions. Sterner instructions, repetition, or extended thinking in that session do not remove the bias.
- 2
Use a second, independent instance to review generated code: start it fresh with only the code (and any spec) and none of the generation conversation. Independent review catches subtle issues that self-review instructions or extended thinking miss.
- 3
The same applies to claim verification in research systems: a separate verification subagent that receives only the draft report and the sources, with no synthesis history, catches citations that do not support their sentences.
- 4
Attention dilution: a single pass over many files gives detailed feedback on the first files and superficial feedback on the rest, misses obvious bugs, and produces contradictory findings (flagging a pattern in one file and approving it in another). A larger context window does not fix attention quality.
- 5
Multi-pass review: run focused per-file passes for local issues, then a separate integration pass over the full diff plus call sites for cross-file data flow, signature changes and inconsistent handling. Either half alone leaves a gap.
- 6
Per-file passes cannot see other files, so instructing them to "consider how this file interacts with the rest of the PR" asks for analysis they cannot perform; that is the integration pass's job.
- 7
Requiring developers to split PRs, or to list impacted callers by hand, shifts the burden to people without improving the system.
- 8
Consensus voting across several full-PR runs suppresses real issues that are only caught intermittently, multiplies cost, and does not fix systematic misjudgements that every run shares.
- 9
Verification passes: after a first pass produces findings, an independent instance re-examines each finding against the code and reports a confidence value with it. Use the confidence for calibrated routing: post confident findings directly, send uncertain ones to a summary or a human triage queue, and never silently discard them.
- 10
Confidence reported by the same conversation that produced the findings is inflated (nearly everything comes back at 0.9+ and nothing is retracted). Moving the threshold on an uncalibrated signal does not help; move verification to an independent instance and check its confidence against labelled accept/dismiss outcomes before trusting thresholds.
- 11
Confidence is not severity: how sure the verifier is that a finding is real is a different quantity from how bad the finding is.
- 12
Do not skip verification when the first pass claims high confidence in itself, and do not replace per-finding verification with pattern-level dismissal history;
detected_patternstatistics are for prompt improvement, not for deciding whether a specific finding is real.
Read the source
Test yourself on 4.6 Design multi-instance and multi-pass review architectures
Ten questions, with the answer and explanation after each one.