AI Application Security
Recognize jailbreaks and direct or indirect prompt injection, and choose structural mitigations: least privilege, untrusted-data handling, authorization outside the model, data minimization and PII protection.
Key points
- 1
Two threat models: in jailbreaks and direct prompt injection the *user* is the adversary; in indirect prompt injection the user is trusted but third-party content (web pages, emails, documents, tool results) carries the attack.
- 2
For jailbreaks, pre-screen user input with a lightweight model (for example a Haiku-tier classifier with a structured yes/no output), add input validation, and write a system prompt that states boundaries and how to refuse. Throttle or ban repeat offenders.
- 3
For indirect injection, deliver untrusted content only inside
tool_resultblocks (never in thesystemprompt or plain user text), say what it is and where it came from, and state in the system prompt that tool content is data that must not override instructions. - 4
JSON-encoding untrusted strings inside a tool result gives clear delimiters, so an attacker cannot close a tag or quote to break into an instruction context.
- 5
Don't put your own instructions inside tool results, because Claude treats that content skeptically. Send instructions in a following user turn.
- 6
The strongest defense is structural least privilege: an agent that reads untrusted content should not also hold sensitive data and an outbound channel (send email, fetch arbitrary URLs). Split capabilities or gate the high-impact action behind human approval.
- 7
Common traps: raising temperature, switching to a "smarter" model, asking users nicely, keyword blocklists, and repeating warnings in the prompt are not effective mitigations on their own.
- 8
Authorization belongs in application code. Tool handlers derive identity from the authenticated session, or call downstream systems with the user's delegated credentials. Never trust a user ID or role that arrives through the model or the conversation.
- 9
Data leakage: anything in the context window can end up in a response. Filter retrieval by the user's permissions before content reaches the model, and scope conversation history and caches per user or tenant.
- 10
Reduce prompt leak by leaving out details the model does not need. Keep secrets such as discount codes or keys behind tools and backend logic. Output screening and post-processing can catch leaks, and "never reveal" wording is not a control.
- 11
PII: redact or tokenize before the API call, keep any re-identification mapping on your side, and keep raw PII out of logs. Base64 is encoding, not protection.
- 12
Treat model output as untrusted too. Escape it before rendering (to prevent XSS), and validate proposed actions against a schema or allowlist before your own code applies them with parameterized queries.
- 13
Encryption at rest protects stored data from infrastructure compromise. It does not stop the application from serving data to the wrong user.
- 14
For data handling on the Claude API, Anthropic offers zero data retention (ZDR) arrangements per organization, and on Bedrock or Google Cloud the cloud provider is the data processor. Check a feature's eligibility before relying on ZDR.
Read the source
Test yourself on AI Application Security
Ten questions, with the answer and explanation after each one.