CCAR-P sample questions with answers
10 questions from the Claude Certified Architect – Professional practice bank, spread across its domains. Pick your answer, then open the explanation to see why each option is right or wrong.
- Question 1Integration
A defense contractor deploys Claude Code to 600 engineers. Policy says the agent must never run
curlorwgetand must never edit anything under/etc/secrets. Several teams have already addedBash(curl *)to their project.claude/settings.jsonallow lists, and individuals can edit.claude/settings.local.jsonfreely. The security team needs the block to hold regardless of what any developer or team configures. Where should the rules go?- A
In each repository's
.claude/settings.jsonunderpermissions.deny, committed to version control so every team member inherits the same deny rules on checkout. - B
In every developer's
~/.claude/settings.json, pushed once by IT during machine setup. - C
In managed settings (
managed-settings.jsonvia MDM) aspermissions.denyrules forBash(curl *),Bash(wget *)andEdit(//etc/secrets/**). - D
In a
CLAUDE.mdat the organization root instructing Claude never to use curl or wget and never to touch/etc/secrets.
Show the answer and explanation
Answer: C
Guaranteed enforcement across a fleet needs a deterministic control at the layer nothing below can override: managed settings deployed by the organization. Deny rules there take precedence over any allow rule from project, local or user files, and the
//pathform anchors the edit rule to the absolute filesystem path.Why the other options are wrong
A. Project settings are owned by the teams that have already been adding curl allows; anyone with commit access can remove the deny rule again.
B. User settings are fully editable by the developer and sit below project and managed settings in precedence, so the block would not hold.
D. Memory files shape behaviour but are not enforced; a prompt-level instruction can be ignored or overridden and gives no guarantee.
- A
- Question 2Solution Design & Architecture
A SaaS vendor's CEO says: "Customers churn because onboarding is hard; build an AI assistant." The architect's discovery finds that 60% of churned accounts never completed integration setup, and support tickets raised during setup cluster around three recurring configuration errors. Which solution translation makes the strongest business case?
- A
A guided-setup assistant that checks the customer's integration configuration and walks them through the three recurring fixes, measured by 14-day integration completion against a control cohort.
- B
A general product assistant with retrieval over the full documentation set and every help-centre article, available on every page of the product, measured by customer satisfaction with the assistant's answers.
- C
An agent with credentials to each customer's environment that detects configuration problems and applies the fixes automatically without asking, measured by the number of fixes applied per week.
- D
A six-month discovery programme interviewing churned customers to confirm the cause before committing to any build, measured by interviews completed.
Show the answer and explanation
Answer: A
Translate the business problem into the narrowest solution that attacks the evidenced root cause, and define success with a metric that moves with the business outcome and can be compared against a control. Breadth, autonomy and further research all weaken the case here.
Why the other options are wrong
B. A broad assistant dilutes effort across the whole product, and its satisfaction metric does not tell the CEO whether onboarding completion or churn moved.
C. Autonomous writes into customer environments carry blast-radius and trust risks that are unjustified for a first solution when guided fixes address the same errors.
D. Discovery already surfaced a specific, measurable cause; further open-ended research delays value without changing the diagnosis.
- A
- Question 3Evaluation, Testing & Optimization
An insurer migrated its claim-summary service to a newer Claude model. Eval accuracy is unchanged, but summaries are now about 40% longer, breaching the 150-word guideline that adjusters rely on, and the tone is more conversational. The prompt was written for the previous model and never states a length or tone explicitly. What is the most likely cause and the right fix?
- A
The new model is hallucinating; add citation requirements to the prompt.
- B
Cap
max_tokensat roughly 150 words' worth so the model cannot exceed the limit. - C
The prompt relied on the old model's implicit defaults; state length and tone explicitly and re-run the eval with a length check.
- D
Roll back to the previous model permanently, since a model that changes summary length and tone without being asked is not fit for this task.
Show the answer and explanation
Answer: C
A behaviour change that appears only after a model swap, with accuracy intact, is model mismatch with the prompt's unstated assumptions. Make the requirements explicit, add the missing checks to the eval, and validate the migration rather than blaming the model or truncating output.
Why the other options are wrong
A. Accuracy is unchanged and the symptom is length and tone, not invented facts; a hallucination remedy addresses the wrong failure.
B. A hard token cap truncates mid-sentence and can cut off the conclusion; it enforces a ceiling, not a well-formed short summary.
D. A style difference caused by an under-specified prompt is not evidence of model unfitness; a permanent rollback forfeits the migration without diagnosing anything.
- A
- Question 4Governance, Safety & Risk Management
An HR platform uses Claude to rank applicants for interview. On a held-out set the ranking agrees with recruiters 91% of the time, and the team considers it ready. A compliance reviewer says the evaluation is incomplete. What is missing?
- A
Per-segment evaluation comparing selection and error rates across applicant groups and proxies such as name, school and address, plus qualified human review of decisions.
- B
A larger held-out test set drawn from the same applicant pool so the 91% agreement figure carries a tighter confidence interval before launch.
- C
Removing applicant names from the prompt before ranking, which eliminates the main channel through which demographic bias could enter the model's decision.
- D
A system prompt instruction telling the model to be fair and unbiased in its rankings.
Show the answer and explanation
Answer: A
Bias testing means evaluating outcomes per segment, not just overall accuracy, and looking for proxies for protected characteristics. Employment is a high-risk domain in Anthropic's Usage Policy, so a qualified human must review decisions as well.
Why the other options are wrong
B. More of the same aggregate metric does not reveal whether errors are distributed unequally.
C. Names are one proxy among many; addresses, schools, employment gaps and phrasing can carry the same signal.
D. An instruction is not a measurement; it gives no evidence about outcomes across groups.
- A
- Question 5Stakeholder Communication & Lifecycle Management
An architect is preparing a steering-committee presentation recommending a Claude-based product-description generator for a retailer's 40,000-SKU catalogue. Which two elements should the presentation include? (Choose two.)
Choose 2.
- A
The options considered and the trade-off each makes between accuracy, cost, latency and risk, expressed in business terms.
- B
Only the recommended option, to keep the committee focused on approving the budget.
- C
A statement that the system prompt prevents the model from producing incorrect product claims.
- D
The residual risks after mitigation, such as the expected rate of descriptions needing correction, and who is accountable for them.
- E
Demo outputs from a curated set of ten products, presented as representative of production quality.
- F
The full system prompt, so the committee can review the wording line by line.
Show the answer and explanation
Answer: A and D
Communicating an architectural decision means showing the options and their trade-offs in business terms, and being explicit about residual risk and ownership. Hiding alternatives, promising that prompts prevent errors and passing off curated demos as production quality are the common ways such presentations mislead.
Why the other options are wrong
B. Suppressing the alternatives hides the trade-offs the committee is being asked to accept; it may speed approval but undermines trust when the trade-offs surface later.
C. A prompt reduces but cannot eliminate errors; presenting it as a guarantee over-promises and sets the committee up for surprise when descriptions need correction.
E. Ten hand-picked outputs show feasibility, not quality across 40,000 SKUs; presenting them as representative misleads the committee about the error rate.
F. Prompt wording is an implementation detail the committee cannot evaluate; it displaces the business trade-offs the presentation should be about.
- A
- Question 6Claude Models, Prompting & Context Engineering
A coding agent is working on a database migration over several hours. As the conversation approaches its context limit, the team must keep the architectural decisions made so far and the list of open bugs, but not the hundreds of file listings and test outputs already dealt with. Which technique fits?
- A
Drop the oldest messages once the limit is near, keeping only the most recent forty turns, since the recent tool outputs and test results are what the next edit depends on.
- B
Start a fresh session and paste in the original task description.
- C
Increase
max_tokensso the model can summarise its own history in the reply. - D
Compact the conversation: summarise the history into a note that keeps decisions and open items, drop redundant tool outputs, and continue from the summary.
Show the answer and explanation
Answer: D
Compaction summarises a conversation nearing its limit while preserving what matters (decisions, unresolved issues, implementation details) and discarding redundant outputs. Age-based truncation and restarts lose the state the agent needs.
Why the other options are wrong
A. The earliest messages often contain the task framing and key decisions; discarding them by age loses exactly what must be kept.
B. This loses every decision and bug discovered during the run; the agent would redo work and repeat mistakes.
C.
max_tokensgoverns output length; it neither shrinks the context nor creates a mechanism for carrying a summary forward.
- A
- Question 7Developer Productivity & Operational Enablement
A retail engineering team commits a
.mcp.jsonso everyone gets the same Jira MCP server. During review, the lead notices a developer pasted a personal Jira API token into theheadersblock of the committed file. The team wants to keep one shared server definition in the repository and keep the credential out of git. What should they do?- A
Reference the token as
${JIRA_TOKEN}in.mcp.jsonand have each developer set that variable in their own environment. - B
Move the token into the
envblock of the committed.claude/settings.jsonso it is separated from the server definition. - C
Add
.mcp.jsonto.gitignoreand ask each developer to recreate the server withclaude mcp add. - D
Keep the token in the file but restrict the repository to internal contributors only.
Show the answer and explanation
Answer: A
Project-scoped MCP servers are shared through a committed
.mcp.json, and credentials are injected with environment-variable expansion so the file never contains the secret. Personal or secret-bearing servers can instead be added at local or user scope, which is stored in~/.claude.jsonrather than in the repository.Why the other options are wrong
B. The shared settings file is also committed, so the credential would still be in version control; separating it from the server entry changes nothing about exposure.
C. This removes the secret but also throws away the shared, reviewed configuration the team wanted, and every developer's setup will drift.
D. A secret in git history is copied to every clone, backup and CI runner; repository visibility is not a substitute for keeping credentials out of the file.
- A
- Question 8Integration
An HR team wants an assistant over its 90,000-token employee handbook, which changes about once a week. The assistant will handle roughly 2,000 questions a day. The engineering team is about to build a vector database, chunking pipeline and reranker. Which approach should the architect recommend?
- A
Place the entire handbook in the prompt as a static prefix with prompt caching enabled, and refresh the cached prefix when the handbook changes.
- B
Build the RAG pipeline as planned so that only the handful of relevant chunks is sent on each request, keeping per-request token usage minimal as the handbook grows.
- C
Have Claude summarize the handbook to 5,000 tokens once and answer from the summary.
- D
Split the handbook across three specialist agents (benefits, leave, conduct) and route questions between them.
Show the answer and explanation
Answer: A
Anthropic's guidance is that if the knowledge base is smaller than roughly 200,000 tokens (about 500 pages), the simplest and most reliable option is to include it all in the prompt and rely on prompt caching, which cuts latency and cost on the repeated prefix. RAG is for corpora that do not fit.
Why the other options are wrong
B. RAG introduces chunking and retrieval failure modes and operational cost that are unnecessary for a corpus that fits comfortably in context.
C. A summary drops the specific policy details that HR questions depend on; the answers will be plausible but unreliable.
D. A multi-agent split adds orchestration, latency and cost to solve a problem that a single cached prompt already solves.
- A
- Question 9Solution Design & Architecture
A hospital network must compute 12 quality indicators (for example, documented medication reconciliation within 24 hours) across 2,000 patient records for a quarterly report. Records are independent, indicators require reading free-text notes, and the report needs network-wide totals per indicator. Which two decomposition decisions are correct? (Choose two.)
Choose 2.
- A
Decompose by indicator: for each of the 12 indicators, send all 2,000 records in one request and ask for that indicator's total.
- B
Decompose by data boundary: process each patient record independently, extracting all 12 indicators for that record as structured output.
- C
Let a single agent read records in whatever order it chooses until it is confident of the network-wide totals.
- D
Compute the network-wide totals in code from the per-record structured outputs rather than asking the model to aggregate.
- E
Process records in prompts of 500 records at a time so the model can spot trends across patients while it extracts.
Show the answer and explanation
Answer: B and D
Partition along the natural data boundary (one record per call), extract structured facts per unit, and leave counting and aggregation to deterministic code. Model-side aggregation over thousands of units is neither reliable nor auditable.
Why the other options are wrong
A. Two thousand records do not fit in one request, and mixing many patients' notes into one context invites cross-patient confusion and privacy exposure.
C. Unbounded autonomy over 2,000 records is slow, unreproducible and provides no guarantee that every record was examined.
E. Large mixed batches degrade per-record accuracy and blur patient boundaries; trends belong in the aggregation layer, not the extraction call.
- A
- Question 10Evaluation, Testing & Optimization
An education company's essay-feedback assistant has an eval set that the prompt engineer iterates against daily. Scores have climbed from 78% to 97% in a month, but teachers report no improvement. The architect suspects the eval has stopped measuring real quality. Which two practices should be introduced? (Choose two.)
Choose 2.
- A
Hold out a test split that is never inspected while iterating on the prompt, and report it separately from the development split.
- B
Version the dataset alongside the prompt so every reported score is tied to a specific dataset revision.
- C
Retire cases the model passes consistently to keep the set small and fast.
- D
Let the prompt engineer read every failing test-split case and edit the prompt until it passes.
- E
Regenerate the entire dataset with the latest model before each release so it stays current.
Show the answer and explanation
Answer: A and B
Eval hygiene: keep a held-out split that is not used for tuning, version the dataset with the prompt so scores are comparable, and keep passing cases as a regression suite. Rising scores with no user-visible improvement are the classic sign of overfitting to the eval.
Why the other options are wrong
C. Consistently passing cases are the regression suite; removing them means future changes can silently break behaviour that used to work.
D. This is exactly the leakage that inflates the score: the test split becomes another development set and stops predicting real-world performance.
E. Regenerating everything destroys the baseline; scores across releases become incomparable and the set inherits whatever the generating model gets wrong.
- A
Practise all 231 CCAR-P questions
Start with the free 15-question diagnostic. It shows where to focus, and your results carry over if you sign up.