Understanding Requirements
Turn business goals and constraints (latency, volume, cost, residency, compliance, human oversight) into concrete functional and infrastructure requirements, and pick the Claude surface and deployment that satisfies all of them at once.
Key points
- 1
Separate functional requirements (what the system does: languages, citations, refusals) from infrastructure or non-functional ones (concurrency, latency percentiles, throughput, region, retention). The second group sizes the deployment.
- 2
Make vague goals measurable. "Feels instant" becomes a time-to-first-token target met with streaming, and "100% accurate" becomes per-field accuracy targets checked by evals, plus human review for costly cases.
- 3
Latency-sensitive, user-facing features use the Messages API with streaming. Large, latency-tolerant jobs fit the Message Batches API, which costs 50% less; most batches finish within 1 hour, and a batch expires if it has not finished within 24 hours.
- 4
Match the workload to the surface. Interactive product features use the Messages API. Agentic work inside a codebase uses Claude Code (interactive,
claude -pheadless, or GitHub Actions). An agent embedded in your own product with built-in tools and approval callbacks uses the Claude Agent SDK. - 5
Messages API rate limits are set per model class as requests, input tokens and output tokens per minute (RPM, ITPM, OTPM). Any one of them can be the binding limit, and going over it returns a 429 with a
retry-afterheader. - 6
Limits use a continuously refilling token bucket and may be enforced over shorter windows than a minute, so short bursts can hit 429 even when the per-minute average is under the limit. Plan for a client-side queue with backoff.
- 7
For most models, only uncached input tokens count toward ITPM, and cache reads do not. Caching a large shared prefix therefore raises the throughput you can actually get.
- 8
Isolate competing workloads. The Message Batches API has its own rate limits, and per-workspace limits can cap one workspace so it cannot starve another. Organization-wide limits still apply on top.
- 9
Data residency is set in infrastructure, never in the prompt. On the Claude API,
inference_geo("global" or "us") controls where inference runs, per request or as a workspace default. Workspace geo controls where data is stored at rest. - 10
Buying through a cloud provider decides the platform. Claude on Vertex AI or Amazon Bedrock is billed and governed through that cloud. On Vertex AI, regional and multi-region endpoints keep traffic in a geography, while the global endpoint does not.
- 11
Retention and location are separate requirements. Zero data retention (ZDR) covers only eligible features: Message Batches are not ZDR-eligible because results are kept for 29 days, and ZDR does not block non-eligible features.
- 12
Where human oversight is required, put approval gates exactly where policy says. Enforce the thresholds deterministically in code rather than leaving them to model judgment, size the review queue to reviewer capacity, and audit a sample of auto-approved cases.
- 13
Auditability requirements mean logging the inputs, prompt version, model ID, output and approver for each decision. Keeping only the final outcome is not enough to reconstruct how it was reached.
Test yourself on Understanding Requirements
Ten questions, with the answer and explanation after each one.