Study notes · 2.3% of the exam

Guardrails and Safe Deployment

Design layered, enforceable guardrails and deploy Claude applications securely by design: content policy, defense in depth, least privilege, identity and access management, and human approval for high-impact actions.

Key points

  1. 1

    Guardrail layering means defense in depth: input screening, system-prompt policy, tool restrictions, output checks and human approval, so one bypass is caught by another layer. Repeating one instruction in several places is still a single layer.

  2. 2

    Prompt instructions (system prompt, CLAUDE.md, Skills) are requests, not guarantees. Anything that must always hold belongs in code, permissions, hooks or infrastructure.

  3. 3

    A content policy stricter than general defaults must be stated explicitly and enforced by an independent check, such as a lightweight classifier on outputs. Built-in safety training does not know your application's policy.

  4. 4

    Tune model-based classifiers with precise category definitions and examples of content that should and should not be flagged. Track both false positives and false negatives (precision and recall) on a labeled set rather than removing the layer.

  5. 5

    Least privilege: give each tool and agent a dedicated, narrowly scoped credential (for example a read-only role on reporting views). Do not reuse broad service accounts, and do not grant access for hypothetical future needs.

  6. 6

    Human approval is meaningful only when the agent lacks the privilege to act without it and the approval comes from a separate, authenticated identity. A "yes" typed in the chat is model input, not an approval.

  7. 7

    Approve by risk tier: run read-only tools automatically and hold irreversible or side-effecting actions (payments, deletes, sends, production changes) for approval. Approving everything causes approval fatigue.

  8. 8

    Secure by design: threat-model up front, isolate tenants, minimize data, give each component its own least-privilege identity, and keep audit logs. Pen tests, provider safeguards and prompt rules are supplements, not substitutes.

  9. 9

    Claude Code permission rules are evaluated deny, then ask, then allow, and the first match wins. An allow rule cannot carve an exception out of a deny rule.

  10. 10

    Managed settings (file, MDM policy or server-managed) have the highest precedence and cannot be overridden by user, project or command-line settings. Use them for org-wide policy such as deny rules.

  11. 11

    Bash permission rules match command text, so they are not a security boundary around a program (for example sh -c 'curl ...'). For filesystem and network limits that hold no matter how a command is written, use the sandboxed Bash tool (OS-enforced) or a PreToolUse hook.

  12. 12

    bypassPermissions mode is for isolated containers or VMs only. For unattended CI, prefer dontAsk mode with exact allow rules, and remove credentials the job does not need.

Test yourself on Guardrails and Safe Deployment

Ten questions, with the answer and explanation after each one.