Implement guardrails and safety controls
Design layered, mostly deterministic guardrails for a Claude system: screen inputs, bound the system prompt, handle untrusted content structurally, validate outputs in code, and enforce least privilege on tools so that no single probabilistic control is load-bearing.
Key points
- 1
Two threat models: in jailbreaks and direct prompt injection the user is the adversary; in indirect prompt injection the user is trusted but third-party content (emails, web pages, OCR text, tool results) carries the attack. Different mitigations apply to each.
- 2
Against jailbreaks: a harmlessness screen (a lightweight model such as a Haiku-tier classifier, constrained with structured outputs to a boolean), input validation, a system prompt that states boundaries and how to refuse, and throttling or banning repeat offenders.
- 3
Against indirect injection: put untrusted content only inside
tool_resultblocks, label what it is and where it came from, JSON-encode it so it cannot break out of its delimiters, state in the system prompt that tool content is data, and screen tool outputs with a small classifier before Claude acts on them. - 4
Do not place your own instructions inside tool results; Claude treats that content skeptically. Send follow-up instructions in a user turn instead.
- 5
Least privilege is the strongest guardrail: an agent that reads untrusted content should not also hold secrets or an unconstrained high-impact action. Remove capabilities the role does not need rather than logging or guarding them (CCAR-P sample 1).
- 6
Output validation belongs in code: constrain output shape with structured outputs, validate machine-checkable facts (approved product lists, IDs, amounts) against a source of truth, and use a cheap classifier only for the fuzzy parts. Self-certification tokens and majority voting are not guarantees.
- 7
Agent SDK enforcement order: hooks, deny rules, ask rules, permission mode, allow rules, then
canUseTool. APreToolUsehook deny and scoped deny rules apply even inbypassPermissions;allowedToolsdoes not constrain that mode and auto-approved calls never reachcanUseTool. - 8
Prompt leak: try monitoring and post-processing first (output screening, regex or an LLM filter), remove details the model does not need (move endpoints and secrets behind tools), and add leak-resistant wording only when necessary because it can degrade task quality.
- 9
Keeping Claude in character: define the role in the system prompt with detail, list expected scenarios and canned responses for edge cases, and use retrieval for a fixed knowledge set. Prefill is not available on newer models; use structured outputs or system instructions.
- 10
Content moderation with Claude: define categories with explicit definitions, return structured JSON (violation, categories, explanation, or a risk level), batch when latency allows, track precision and recall, and iterate the policy text as needs change. Built-in safety behaviour may refuse some content regardless of prompt.
- 11
Classic distractors: capital-letter MUST instructions as the only control, regex blocklists as the sole input filter, switching to a bigger model, lowering temperature, keyword stripping, encoding data instead of removing it, and after-the-fact log review presented as prevention.
- 12
Red-team before launch with documents, emails and tool outputs that contain injection attempts, then monitor outputs continuously and feed findings back into prompts, validation and filters.
Read the source
Test yourself on Implement guardrails and safety controls
Ten questions, with the answer and explanation after each one.