AI Cert Prep
Type to search documentation.

CCAR-P Practice Exam

A full-length, 63-item timed practice exam for the Claude Certified Architect – Professional credential, distributed by blueprint weight with answer key and score interpretation.

This is a full-length, 63-item practice exam for CCAR-P. The questions are new (not reused from the domain pages) and distributed to match the blueprint. Sit it under exam conditions before you rely on the score.

Instructions

  • Time: 120 minutes. Set a timer and do not pause it.
  • Format: multiple-choice (select one) and multiple-response (select two); each stem states how many to select.
  • Scoring: the real exam is scaled 100–1000 with a pass at 720. There is no guessing penalty — answer everything. As a raw proxy, aim for ≥ 80% (≈ 50/63) before booking.
  • Method: read the whole stem, identify the binding constraint, eliminate the constraint-blind / over-engineered / prompt-as-enforcement / self-report / silent-failure / recall-only / aggregate-metric distractors, then choose.

Domain distribution

#DomainWeightItems here
1Solution Design & Architecture17%11
2Claude Models, Prompting & Context Engineering13%8
3Integration (incl. RAG)19%12
4Evaluation, Testing & Optimization16%10
5Governance, Safety & Risk Management14%9
6Stakeholder Communication & Lifecycle Management14%9
7Developer Productivity & Operational Enablement7%4
Total100%63


Score interpretation

Raw score (of 63)PercentReading
57–6390–100%Exam-ready; strong across all domains
50–5680–89%Likely pass; shore up your weakest domain
45–4971–79%Borderline; drill D3/D1/D4 and re-sit
38–4460–70%Not ready; systematic gaps — restudy by domain
< 38< 60%Restudy the full course before re-attempting

The real exam scales 100–1000 with a pass at 720; treat ≥ 80% raw (≈ 50/63) here as your go/no-go line, and confirm you have no single domain far below the rest.

Take the practice exam

Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.

Interactive mode

Take the practice exam

63 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.

All questions (review mode)

Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.

  1. Q1D1 · Solution Design and ArchitectureSelect one

    A logistics firm wants to automate shipment-status replies: 500/hour, one internal API, p95 under 4 s, tight cost ceiling. Which pattern should the architect propose?

    • A. A coordinator with five specialised subagents.
    • B. An augmented-LLM / simple workflow (classify → status lookup → structured reply) because the steps are known and the SLA and cost are binding.
    • C. An open-ended agentic loop with a large tool catalogue.
    • D. Fine-tuning before building any pipeline.
    Show answer

    Answer: B.

    Known steps plus binding latency/cost favour the simplest pattern. Multi-agent (A) and broad agentic loops (C) over-engineer; fine-tuning (D) is premature.

  2. Q2D1 · Solution Design and ArchitectureSelect one

    A German bank requires all data to stay in the EU and already runs on AWS. Which placement is BEST?

    • A. Direct Anthropic API.
    • B. Amazon Bedrock in an EU region, inheriting existing AWS IAM and residency.
    • C. Google Vertex AI.
    • D. Whichever model is newest.
    Show answer

    Answer: B.

    EU residency plus an existing AWS footprint point to Bedrock EU. Direct API (A) risks residency; Vertex (C) ignores the AWS investment; novelty (D) can't override residency.

  3. Q3D1 · Solution Design and ArchitectureSelect one

    An agentic loop occasionally runs forever. Which is the CORRECT primary control plus backstop?

    • A. Parse the model's text for the word 'finished'.
    • B. Terminate on stop_reason (end_turn / no further tool_use), with an iteration cap only as a safety backstop.
    • C. A fixed 5-iteration cap as the sole control.
    • D. Lower temperature until it stops.
    Show answer

    Answer: B.

    Control flow keys off stop_reason; a cap is a backstop, not the mechanism. Text parsing (A) is brittle; a cap alone (C) is an anti-pattern; temperature (D) doesn't govern termination.

  4. Q4D1 · Solution Design and ArchitectureSelect two

    Peak load is 25 req/s at ~10k input tokens each, and the team hits ITPM limits. Which TWO are sound?

    • A. Shift latency-tolerant work to the Message Batches API and add backoff with jitter honouring retry-after.
    • B. Retry immediately in a tight loop.
    • C. Request a higher tier and/or route overflow to a second model.
    • D. Return empty results as success on 429.
    • E. Send everything to Opus 5.
    Show answer

    Answer: A and C.

    Batch for tolerant work plus backoff, and raising the tier / spillover routing, are the correct capacity levers. Tight-loop retry (B) worsens it; swallowing 429s (D) is silent failure; Opus for all (E) raises cost without fixing limits.

  5. Q5D1 · Solution Design and ArchitectureSelect one

    What makes an end-to-end reference architecture complete at Professional level?

    • A. Input and processing stages only.
    • B. Input → processing → output → a feedback loop that routes production signals back into prompts/models/retrieval.
    • C. A single model call.
    • D. Only a load balancer and the model.
    Show answer

    Answer: B.

    The feedback loop is mandatory; a design with no path from production signals to improvement is incomplete. The others omit it.

  6. Q6D1 · Solution Design and ArchitectureSelect one

    A stem gives a hard $8,000/month budget and a modest accuracy bar for high-volume classification. Which posture is BEST?

    • A. Opus 5 with xhigh effort.
    • B. Haiku 4.5 (or Sonnet 5) with prompt caching and few tool round-trips, treating cost/volume as the binding pillars.
    • C. Fable 5.1 for maximum quality.
    • D. A multi-agent pipeline.
    Show answer

    Answer: B.

    The binding pillar is cost at volume with a modest quality need — the cheapest adequate model with caching. Opus/Fable (A, C) overspend; multi-agent (D) adds cost.

  7. Q7D1 · Solution Design and ArchitectureSelect one

    When is multi-agent orchestration genuinely justified?

    • A. A three-step sequential pipeline sharing context.
    • B. Independent parallel investigations with isolated context that fan in to a synthesis step.
    • C. A single API lookup with a tight SLA.
    • D. Always, to be safe.
    Show answer

    Answer: B.

    Parallel, context-isolated subtasks with fan-in are the multi-agent signal. Sequential shared work (A) is a workflow; a single lookup (C) is augmented-LLM; 'always' (D) is over-engineering.

  8. Q8D1 · Solution Design and ArchitectureSelect one

    The primary model returns sustained 529s under a spike. Which fallback is BEST?

    • A. Error out to all users until recovery.
    • B. Retry with backoff; if still failing, fall back to a secondary model/provider and log the request as served in degraded mode.
    • C. Serve stale cached answers as if fresh.
    • D. Fail an active Fable 5.1 thinking session over to a much older model.
    Show answer

    Answer: B.

    Retry-then-fallback with degraded-mode logging preserves availability and observability. Hard failure (A) is poor HA; stale-as-fresh (C) is silent failure; downgrading a Fable 5.1 session (D) drops thinking blocks.

  9. Q9D1 · Solution Design and ArchitectureSelect one

    The company lacks ops capacity, needs launch in a month, and the capability is commodity. Build or buy?

    • A. Build a custom Agent SDK runtime and sandbox.
    • B. Use managed agents (Anthropic hosts the loop/sandbox) to hit the deadline with minimal ops.
    • C. Delay until an ML platform team is hired.
    • D. Buy a product with no Claude support.
    Show answer

    Answer: B.

    No ops capacity + tight deadline + commodity capability → managed/buy. Building (A) contradicts the constraints; delaying (C) misses the deadline; (D) abandons the requirement.

  10. Q10D1 · Solution Design and ArchitectureSelect one

    Why capture a design as an ADR?

    • A. To meet a documentation quota.
    • B. To record context, rejected alternatives with reasons, and the accepted trade-off so the decision is reviewable and revisitable.
    • C. Because ADRs replace evaluation.
    • D. Because agentic systems require them technically.
    Show answer

    Answer: B.

    ADRs make reasoning and trade-offs explicit and reviewable. They are not a formality (A), don't replace evals (C), and aren't a technical prerequisite (D).

  11. Q11D1 · Solution Design and ArchitectureSelect one

    A stakeholder gives a vague goal and no numbers. What is the BEST first action?

    • A. Start building immediately.
    • B. Run discovery to define measurable per-segment success criteria and enumerate constraints before designing.
    • C. Default to Opus 5.
    • D. Assume a target and proceed.
    Show answer

    Answer: B.

    Without criteria/constraints the design is unanchored; discovery surfaces the binding constraint. Building (A), defaulting to a model (C), or inventing a target (D) skip anchoring.

  12. Q12D2 · Claude Models, Prompting and Context EngineeringSelect one

    A pipeline runs everything on Opus 5 at 4× budget with acceptable quality. Best cost fix preserving quality?

    • A. Move all traffic to Haiku 4.5.
    • B. Cascade cheap-first (Haiku/Sonnet), escalate to Opus 5 on validation-check failure, and cache the stable prefix.
    • C. Escalate on the model's self-reported confidence.
    • D. Raise effort to xhigh everywhere.
    Show answer

    Answer: B.

    Cheap-first with escalation on validation failure preserves hard-case quality; caching amortises the prefix. Blanket Haiku (A) sacrifices quality; self-report (C) is an anti-pattern; xhigh everywhere (D) raises cost.

  13. Q13D2 · Claude Models, Prompting and Context EngineeringSelect one

    After upgrading to Sonnet 5, requests setting budget_tokens fail with 400. What is the fix?

    • A. Add retries.
    • B. Use thinking: {"type":"adaptive"} with effort, since budget_tokens is removed on Sonnet 5.
    • C. Downgrade to Haiku 4.5 permanently.
    • D. Remove thinking entirely.
    Show answer

    Answer: B.

    budget_tokens is removed on Sonnet 5 (returns 400); adaptive thinking + effort replaces it. Retries (A) can't fix a 400; Haiku (C) is not equivalent; removing thinking (D) discards needed reasoning.

  14. Q14D2 · Claude Models, Prompting and Context EngineeringSelect one

    A cache hit rate is near zero because the volatile user turn is placed before the system prompt and documents. What is the fix?

    • A. Disable caching.
    • B. Reorder to stable-first (system → tools → docs), then the user turn, marking the stable boundary with cache_control.
    • C. Shorten the documents.
    • D. Switch to Haiku.
    Show answer

    Answer: B.

    Caching keys on a stable prefix; the volatile turn must come last. Disabling (A) forfeits savings; length (C) only matters for the minimum; model choice (D) is irrelevant.

  15. Q15D2 · Claude Models, Prompting and Context EngineeringSelect one

    In an active Fable 5.1 session the team must change the system instructions mid-run. Correct approach?

    • A. Edit the original system field.
    • B. Append a new role: "system" message and leave earlier turns untouched (append-only harness).
    • C. Delete the earliest turns.
    • D. Reorder messages so the new instruction is first.
    Show answer

    Answer: B.

    Fable 5.1 thinking blocks are invalidated by editing/reordering/removing earlier turns, so mid-session changes are appended. The other options invalidate downstream thinking.

  16. Q16D2 · Claude Models, Prompting and Context EngineeringSelect one

    A hard rule 'no discount above 25%' must be guaranteed. Where does it belong?

    • A. A strongly worded system-prompt sentence.
    • B. A programmatic tool-permission hook / validation that rejects any discount over 25% before execution.
    • C. A few-shot example.
    • D. The thinking budget.
    Show answer

    Answer: B.

    Hard rules require deterministic enforcement. Prompt wording (A) and few-shot (C) are prompt-as-enforcement; the thinking budget (D) is unrelated.

  17. Q17D2 · Claude Models, Prompting and Context EngineeringSelect one

    A workload mandates Zero Data Retention but the team wants Fable 5.1. Correct guidance?

    • A. Use Fable 5.1; ZDR is universal.
    • B. Fable 5.1 requires 30-day retention and is not ZDR-eligible; choose a ZDR-eligible model (Opus 5 / Sonnet 5).
    • C. Disable your logging to get ZDR on Fable 5.1.
    • D. Use Fable 5.1 and delete logs at 30 days.
    Show answer

    Answer: B.

    Fable 5.1's 30-day retention disqualifies it under a ZDR mandate. ZDR isn't universal (A); your logging (C) doesn't change provider retention; 30-day deletion (D) is what ZDR forbids.

  18. Q18D2 · Claude Models, Prompting and Context EngineeringSelect one

    Which statement about current-model thinking config is correct?

    • A. All current models require budget_tokens.
    • B. Current models use adaptive thinking with effort (low/medium/high/xhigh); only Haiku 4.5 still uses budget_tokens and lacks effort.
    • C. effort exists only on Haiku 4.5.
    • D. xhigh is the default everywhere.
    Show answer

    Answer: B.

    Adaptive thinking + effort is standard; Haiku 4.5 is the budget_tokens exception without effort. The others misstate the API.

  19. Q19D2 · Claude Models, Prompting and Context EngineeringSelect one

    Reusable capability blocks are pasted into every prompt, bloating context and breaking the cache prefix. Better pattern?

    • A. Package them as Skills loaded progressively on demand, keeping the cache prefix lean.
    • B. Duplicate them into each service prompt.
    • C. Move them into the volatile user turn.
    • D. Enlarge the context window.
    Show answer

    Answer: A.

    Skills load capability on demand and keep the prefix stable. Duplication (B) caused the bloat; the volatile turn (C) worsens caching; a bigger window (D) doesn't fix cost/caching.

  20. Q20D3 · Integration (incl. RAG)Select one

    A knowledge assistant keeps giving the OLD answer, confidently, after a document was updated overnight. What to investigate FIRST?

    • A. Rewrite the system prompt.
    • B. Inspect what retrieval returned; the updated doc was likely not re-chunked/re-embedded/re-indexed, so retrieval serves stale vectors.
    • C. Switch to Opus 5 xhigh.
    • D. Add few-shot examples.
    Show answer

    Answer: B.

    Confident-wrong-after-refresh points to stale retrieval/indexing; inspect retrieved chunks first. Prompt changes (A, D) and a bigger model (C) can't fix stale retrieval.

  21. Q21D3 · Integration (incl. RAG)Select one

    A support agent has 18 tools including delete_account it should never use. Correct remediation?

    • A. Keep it but log every call.
    • B. Remove the unneeded tools from the allowlist (least privilege) so it cannot invoke them.
    • C. Add an 'are you sure?' prompt.
    • D. Forbid it in the system prompt.
    Show answer

    Answer: B.

    Least privilege removes the capability. Logging (A) and confirmation (C) leave it reachable; a prompt rule (D) is bypassable.

  22. Q22D3 · Integration (incl. RAG)Select one

    Users search by exact SKU and by description; dense-only retrieval misses many SKUs. Best fix?

    • A. Increase k to 1000.
    • B. Hybrid retrieval (BM25 + dense with rank fusion) so exact identifiers and semantic matches both rank.
    • C. A bigger generation model.
    • D. Remove metadata filters.
    Show answer

    Answer: B.

    Dense is weak on exact tokens; sparse handles SKUs and hybrid covers both. Huge k (A) adds noise; a bigger model (C) doesn't change retrieval; removing filters (D) harms precision/security.

  23. Q23D3 · Integration (incl. RAG)Select one

    recall@10 is 0.94 but answers include facts not in the retrieved chunks. Which layer and fix?

    • A. Retrieval; lower k.
    • B. Generation/grounding; enforce answer-only-from-context, add citations, and rerank so the best passage is on top.
    • C. Embeddings; change the model.
    • D. Indexing; re-index everything.
    Show answer

    Answer: B.

    High recall means retrieval works; unsupported facts are a grounding failure. The retrieval-side fixes (A, C, D) target a healthy layer.

  24. Q24D3 · Integration (incl. RAG)Select one

    A 3M-document corpus changes daily and answers must cite the exact clause. Best approach?

    • A. Weekly fine-tuning on the corpus.
    • B. RAG with hybrid retrieval, reranking, citations, and a scheduled refresh pipeline.
    • C. Stuff the whole corpus into 1M context per query.
    • D. Long context + fine-tuning.
    Show answer

    Answer: B.

    Large, changing, citation-requiring corpora are canonical RAG. Fine-tuning (A) can't track daily facts; the corpus exceeds/overspends context and loses citations (C); (D) inherits both issues.

  25. Q25D3 · Integration (incl. RAG)Select two

    Contracts chunked at fixed 500 tokens produce answers citing the wrong sub-clause and losing context. Which TWO help most?

    • A. Structural/document-aware chunking on clause/section boundaries.
    • B. Parent–child chunking: match small children, return the larger parent clause.
    • C. Dense-only retrieval.
    • D. Higher temperature.
    • E. Removing citations.
    Show answer

    Answer: A and B.

    Structured docs need boundary-aware chunking and parent–child context. Dense-only (C), temperature (D), and removing citations (E) don't address chunking.

  26. Q26D3 · Integration (incl. RAG)Select one

    A remote MCP server over Streamable HTTP exposes powerful tools with no auth. What must be added?

    • A. Nothing; MCP is safe by default.
    • B. OAuth 2.1 on the remote MCP server plus per-user permission checks inside the tools.
    • C. A longer prompt.
    • D. A higher rate-limit tier.
    Show answer

    Answer: B.

    Remote MCP requires OAuth 2.1 and per-user authorization. MCP isn't authenticated by default (A); prompts (C) and tiers (D) don't address authz.

  27. Q27D3 · Integration (incl. RAG)Select one

    After enabling backoff retries, some customers are charged twice. Correct fix?

    • A. Disable retries.
    • B. Add idempotency keys to the charge action so retries de-duplicate with no duplicate side effect.
    • C. Lower temperature.
    • D. Reconcile duplicates weekly.
    Show answer

    Answer: B.

    Idempotency keys make retries safe. Disabling retries (A) harms resilience; temperature (C) is irrelevant; reconciling later (D) still double-charged.

  28. Q28D3 · Integration (incl. RAG)Select one

    Some questions require follow-up lookups combining multiple sources. Which design fits and its trade-off?

    • A. Agentic RAG (model decides when/what to retrieve, issues follow-ups), at the cost of more latency/tokens.
    • B. Stuff everything into context.
    • C. Fine-tune on the multi-hop set.
    • D. Remove reranking for speed.
    Show answer

    Answer: A.

    Multi-hop queries justify agentic RAG; the trade-off is cost/latency. Stuffing (B) doesn't scale; fine-tuning (C) can't hold changing facts; removing reranking (D) hurts precision.

  29. Q29D3 · Integration (incl. RAG)Select one

    A capability will be reused across many agents/clients under a standard protocol. Best integration mechanism?

    • A. A per-app CLI tool copied into each service.
    • B. An MCP server (OAuth 2.1 if remote) for standard, reusable access.
    • C. Agent-to-agent handoff only.
    • D. Pasting the logic into every prompt.
    Show answer

    Answer: B.

    Cross-client reuse under a standard protocol is MCP's purpose. Per-app tools (A) fragment; agent-to-agent (C) is for context isolation; prompt-pasting (D) is unmaintainable.

  30. Q30D3 · Integration (incl. RAG)Select two

    An agent loads all 40 tools and 30 docs each request: high cost, low cache hits, poor tool selection. Which TWO fix it?

    • A. Tool search with defer_loading: true to load only relevant tools.
    • B. Package reference material as Skills loaded on demand.
    • C. Enlarge the window to 1M and keep loading everything.
    • D. Escalate every request to Opus 5.
    • E. Disable prompt caching.
    Show answer

    Answer: A and B.

    Progressive discovery keeps context lean, restores the cache prefix, and improves tool selection. A bigger window (C) still pays for bloat; escalating (D) raises cost; disabling caching (E) is the opposite of the fix.

  31. Q31D3 · Integration (incl. RAG)Select two

    Which observability signals let you reconstruct a failed multi-step request end to end?

    • A. A correlation ID threaded through app → model → tools → downstream.
    • B. Per-span traces with token/cost telemetry, retrieved chunk IDs, and stop reasons.
    • C. Only the final HTTP status code.
    • D. Aggregate daily request counts only.
    • E. The model's self-assessment of the request.
    Show answer

    Answer: A and B.

    Correlation IDs plus per-span traces (with retrieval and cost detail) make an incident reconstructable. A status code (C) and daily counts (D) are too coarse; self-assessment (E) is unreliable.

  32. Q32D4 · Evaluation, Testing and OptimizationSelect one

    A bot reports 90% overall accuracy but only refund tickets fail. What is needed?

    • A. Nothing; 90% is fine.
    • B. Per-segment metrics that break accuracy down by ticket type, exposing the refund failure the aggregate hides.
    • C. A bigger same-mix golden set.
    • D. A higher-effort model for all tickets.
    Show answer

    Answer: B.

    Aggregate masks a segment failure; stratified metrics expose it. The aggregate (A) misleads; a same-mix set (C) still hides it; blanket escalation (D) overpays.

  33. Q33D4 · Evaluation, Testing and OptimizationSelect one

    Grading answers with the same model and session that produced them is unsound because…

    • A. it costs too much.
    • B. same-session self-review carries the producing bias; the judge must be a separate model/session and calibrated against humans.
    • C. the judge must always be larger.
    • D. LLM-as-judge is never valid.
    Show answer

    Answer: B.

    Independence plus human calibration is required. Cost (A) isn't the core issue; the judge needn't be larger (C); LLM-as-judge is valid when independent (D).

  34. Q34D4 · Evaluation, Testing and OptimizationSelect one

    Prompt B beats A on 7 of 12 hand-picked examples. Correct conclusion?

    • A. Ship B.
    • B. The sample is too small; run a paired A/B on a large per-segment set and test for statistical significance before promoting.
    • C. Ship A as incumbent.
    • D. Average the prompts.
    Show answer

    Answer: B.

    7/12 is noise; promotion needs a significant paired result with no per-segment regression. Eyeballing (A, C) and averaging prompts (D) are unsound.

  35. Q35D4 · Evaluation, Testing and OptimizationSelect one

    Mean latency is 2 s but users complain; p95 is 12 s. What should the SLA track?

    • A. Mean only.
    • B. Tail latency (p95/p99), which reflects the slow experiences the mean hides.
    • C. Request count.
    • D. Token count.
    Show answer

    Answer: B.

    Interactive SLAs track p95/p99. The mean (A) hides the tail; count (C) and tokens (D) aren't latency SLAs.

  36. Q36D4 · Evaluation, Testing and OptimizationSelect two

    An LLM judge rates longer answers higher regardless of correctness. What is it and the fix?

    • A. Length bias.
    • B. Calibrate against human labels and revise the rubric to score correctness/grounding, not length.
    • C. The judge is reliable.
    • D. Always prefer the longer answer.
    • E. Delete the eval.
    Show answer

    Answer: A and B.

    Length bias is fixed by calibration and a correctness-focused rubric. The judge isn't reliable (C), rewarding length (D) is the bug, and deleting evals (E) abandons measurement.

  37. Q37D4 · Evaluation, Testing and OptimizationSelect one

    How to prevent a fix for one segment from silently breaking another?

    • A. Manual spot-checks post-release.
    • B. A per-segment regression suite in CI that blocks merge on any segment regression.
    • C. Trust the author.
    • D. Test only the changed segment.
    Show answer

    Answer: B.

    A per-segment regression gate in CI is the guard. Spot-checks (A) and author trust (C) are unreliable; testing only the changed segment (D) is the risk.

  38. Q38D4 · Evaluation, Testing and OptimizationSelect one

    A system fails only on the hardest 5% of cases. Most likely cause and fix?

    • A. Retrieval; re-chunk everything.
    • B. Model mismatch (tier too small); escalate just those cases to a stronger model via a cascade.
    • C. Prompt failure; rewrite the whole prompt.
    • D. Infra; add replicas.
    Show answer

    Answer: B.

    'Only the hardest cases fail' signals model capability; a cascade escalates those cases. Re-chunking (A) and rewriting the prompt (C) target working layers; replicas (D) don't affect correctness.

  39. Q39D4 · Evaluation, Testing and OptimizationSelect one

    Cut cost 50% on a nightly bulk job with no latency need. Best lever and what must follow?

    • A. Lower effort blindly and ship.
    • B. Move to the Message Batches API (50% off, within 24 h) and re-run the per-segment regression suite to confirm no quality regression.
    • C. Switch interactive traffic to Haiku too.
    • D. Disable evaluation.
    Show answer

    Answer: B.

    Latency-tolerant bulk work fits Batch; every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net.

  40. Q40D4 · Evaluation, Testing and OptimizationSelect two

    Which two are the right metrics for a high-volume, cost-sensitive service with an interactive SLA?

    • A. Cost per task.
    • B. p95 latency.
    • C. Mean latency only.
    • D. Total fleet tokens.
    • E. Prompt-version count.
    Show answer

    Answer: A and B.

    Cost-sensitive + interactive → cost per task and p95. Mean (C) hides the tail; total tokens (D) and version count (E) aren't user-facing SLA metrics.

  41. Q41D4 · Evaluation, Testing and OptimizationSelect one

    A golden set is 85% English tickets; production fails on Spanish. Flaw and fix?

    • A. The set was too small.
    • B. It lacked per-segment coverage; stratify by language and report per segment.
    • C. The model is broken.
    • D. Online evals are unnecessary.
    Show answer

    Answer: B.

    An unstratified set makes a segment invisible offline. Size (A) isn't the issue; the model isn't broken (C); online evals (D) are still needed but the root cause is coverage.

  42. Q42D5 · Governance, Safety and Risk ManagementSelect one

    A payment agent must never move over $10,000 without approval. Correct design?

    • A. A firm system-prompt sentence.
    • B. A tool-permission hook / validation that blocks transfers over \$10,000 and routes them to a human approval gate.
    • C. Ask the model to report confidence before transfers.
    • D. Escalate only when the user sounds anxious.
    Show answer

    Answer: B.

    Hard limits need deterministic enforcement plus a human gate. Prompt wording (A), self-report (C), and sentiment (D) are anti-patterns.

  43. Q43D5 · Governance, Safety and Risk ManagementSelect one

    A US federal agency requires FedRAMP High. Compliant deployment?

    • A. Direct Anthropic API.
    • B. Claude via Bedrock or Vertex within the FedRAMP High boundary.
    • C. Any cloud; Claude is inherently compliant.
    • D. A personal laptop.
    Show answer

    Answer: B.

    FedRAMP High is provided through Bedrock/Vertex boundaries. Direct API (A) is outside it; compliance isn't inherent (C); a laptop (D) is unacceptable.

  44. Q44D5 · Governance, Safety and Risk ManagementSelect two

    An agent summarising web pages reads hidden text telling it to email data externally. What is it and the defence?

    • A. Indirect prompt injection.
    • B. Treat external content as untrusted, apply least privilege so it lacks an unrestricted email tool, and validate/deny such actions.
    • C. Trust it because it came via a tool.
    • D. Raise effort.
    • E. Add the instruction to the system prompt.
    Show answer

    Answer: A and B.

    Malicious instructions in fetched content are indirect injection; defence is untrusted-content handling plus least privilege and validation. Trusting tool output (C) is the vulnerability; effort (D) is irrelevant; adding it to the prompt (E) executes the attack.

  45. Q45D5 · Governance, Safety and Risk ManagementSelect two

    A clinic wants a Claude assistant reading patient records. First compliance requirements?

    • A. A signed BAA with the provider.
    • B. PHI-handling controls (minimum necessary, access controls, audit logs) and an appropriate retention/ZDR posture.
    • C. Nothing; it is internal.
    • D. Publishing patient data to improve the model.
    • E. Escalating on patient sentiment.
    Show answer

    Answer: A and B.

    PHI under HIPAA requires a BAA and PHI controls with auditing. 'Internal' (C) doesn't exempt PHI; publishing data (D) is a violation; sentiment escalation (E) is unrelated.

  46. Q46D5 · Governance, Safety and Risk ManagementSelect one

    Escalating to a human whenever the customer 'sounds angry' is flawed because…

    • A. it is optimal.
    • B. sentiment is not a proxy for complexity/risk; escalate on measured complexity or deterministic rules instead.
    • C. it escalates too rarely.
    • D. anger always means the model failed.
    Show answer

    Answer: B.

    Sentiment-based escalation conflates emotion with difficulty. It isn't optimal (A); frequency (C) isn't the core flaw; anger doesn't imply model failure (D).

  47. Q47D5 · Governance, Safety and Risk ManagementSelect one

    Which set best represents layered guardrails?

    • A. One very detailed system prompt.
    • B. Input classifier → prompt guidance → tool-permission hooks → output validation → human review for high-stakes actions.
    • C. Only human review at the end.
    • D. Only an output regex.
    Show answer

    Answer: B.

    Defence in depth stacks independent layers. A single prompt (A), review-only (C), or regex-only (D) are single points of failure.

  48. Q48D5 · Governance, Safety and Risk ManagementSelect two

    High-risk processing of EU residents' personal data — required before launch?

    • A. A DPIA.
    • B. EU residency and a data-subject erasure/deletion path.
    • C. Wait for a complaint.
    • D. Store all data indefinitely.
    • E. Escalate on model confidence.
    Show answer

    Answer: A and B.

    GDPR high-risk processing requires a DPIA plus residency and DSAR/erasure support. Waiting (C) and indefinite storage (D) violate GDPR; confidence escalation (E) is unrelated.

  49. Q49D5 · Governance, Safety and Risk ManagementSelect one

    A prompt-injection incident is detected in production. Sound first containment?

    • A. Delete all logs.
    • B. Roll back the prompt/model version and/or disable the exploited tool via its permission hook, scoping impact with correlation IDs and traces.
    • C. Ignore it until the next release.
    • D. Publicly post customer data for transparency.
    Show answer

    Answer: B.

    Containment is fast rollback and disabling the exploited capability, scoped via traces. Deleting logs (A) destroys evidence; ignoring (C) prolongs harm; disclosing data (D) is a second breach.

  50. Q50D5 · Governance, Safety and Risk ManagementSelect one

    Where should API keys and DB credentials for a Claude agent live?

    • A. In the system prompt.
    • B. In environment variables or a secret manager, never in prompts, CLAUDE.md, or logs.
    • C. In CLAUDE.md for convenience.
    • D. Hard-coded in tool source and printed to logs.
    Show answer

    Answer: B.

    Secrets belong in env/secret managers, never in model-visible config or logs. The others create exfiltration paths.

  51. Q51D5 · Governance, Safety and Risk ManagementSelect two

    A hiring-screening assistant makes consequential recommendations. Required controls?

    • A. Per-group fairness evaluation and symmetric treatment reviewed for bias.
    • B. A human decision-maker retains authority over the hiring decision.
    • C. Full automation to remove human bias.
    • D. Escalate only when a candidate sounds confident.
    • E. Store all applicant data indefinitely.
    Show answer

    Answer: A and B.

    Consequential people-decisions require fairness evaluation and human authority. Full automation (C) launders bias; sentiment escalation (D) and indefinite retention (E) are anti-patterns/GDPR issues.

  52. Q52D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Two stakeholders disagree and no success metric exists; the team wants to build. First action?

    • A. Build the most likely interpretation.
    • B. Run structured discovery to align on measurable per-segment criteria and constraints, and get sign-off before design.
    • C. Pick Opus 5 and start.
    • D. Escalate to the vendor.
    Show answer

    Answer: B.

    Without agreement or metrics the discovery gate isn't passed. Building (A), picking a model (C), or escalating (D) skip anchoring.

  53. Q53D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Presenting a workflow-vs-agentic decision to an executive sponsor AND engineers. Best approach?

    • A. Send both the full technical ADR.
    • B. Sponsor gets a one-page decision matrix with cost/risk/value; engineers get the ADR with rejected alternatives and consequences.
    • C. Both get a value-only slide.
    • D. Both get raw eval logs.
    Show answer

    Answer: B.

    Communication matches audience altitude. One document for both (A, C) or raw logs (D) mismatches an audience.

  54. Q54D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Turning 'faster support' into an SLA. Correct approach?

    • A. Promise 'much faster'.
    • B. Agree numeric per-segment targets (p95 latency, deflection, cost/ticket) and report on a cadence.
    • C. Track mean latency only.
    • D. Leave it undefined.
    Show answer

    Answer: B.

    SLAs must be measurable and per-segment with reporting. Vague promises (A) can't be verified; mean-only (C) hides the tail; undefined (D) invites disputes.

  55. Q55D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    Which C4 level best shows executives how the system fits its environment?

    • A. Level 4 (Code).
    • B. Level 1 (Context).
    • C. Level 3 (Component).
    • D. Exhaustive sequence diagrams.
    Show answer

    Answer: B.

    The Context level shows the system, users, and external systems for a broad audience. Code (A) and Component (C) are for engineers; exhaustive diagrams (D) overwhelm.

  56. Q56D6 · Stakeholder Communication and Lifecycle ManagementSelect two

    A model the service pins was retired. Correct migration approach?

    • A. Read the deprecation notice, identify breaking changes, and re-run the regression suite per segment on the target model, adjusting prompts.
    • B. Canary the new model with rollback, then ramp, and communicate timeline/differences.
    • C. Swap the ID in production and hope.
    • D. Keep calling the retired model.
    • E. Disable evals to migrate faster.
    Show answer

    Answer: A and B.

    Deprecation is a governed lifecycle event: assess, re-validate per segment, canary, communicate. A blind swap (C) risks regressions; calling a retired model (D) fails; disabling evals (E) removes the safety net.

  57. Q57D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    What deliverable is required to pass the handoff gate to operations?

    • A. A marketing deck.
    • B. A runbook (dashboards, alerts, on-call, rollback, dependencies) plus C4 docs so operators can run and recover it.
    • C. Nothing; they'll figure it out.
    • D. Only the source code.
    Show answer

    Answer: B.

    Handoff's exit criterion is that operators can run and recover the system. A deck (A), nothing (C), or code alone (D) don't enable safe operation.

  58. Q58D6 · Stakeholder Communication and Lifecycle ManagementSelect one

    When should a DPIA/BAA be addressed in the lifecycle?

    • A. After launch if a regulator asks.
    • B. At the design gate, so residency, retention, and PHI/PII handling shape the architecture before build.
    • C. Never, if internal.
    • D. Only during incident response.
    Show answer

    Answer: B.

    Compliance must shape design, so DPIA/BAA belong to the design exit criterion. Later (A, D) is too late; 'internal' (C) doesn't exempt regulated data.

  59. Q59D6 · Stakeholder Communication and Lifecycle ManagementSelect two

    A technically strong system sees low adoption. Which levers help most?

    • A. Training/enablement and clear communication of capabilities and limits.
    • B. Phased rollout with feedback capture and success-metric reporting.
    • C. Mandatory immediate use with no support.
    • D. Removing human-in-the-loop gates for speed.
    • E. Hiding limitations from users.
    Show answer

    Answer: A and B.

    Adoption is driven by enablement, phased rollout, feedback, and demonstrated value. Forced use (C), removing gates (D), and hiding limits (E) backfire.

  60. Q60D6 · Stakeholder Communication and Lifecycle ManagementSelect two

    The team plans to deploy to 100% of users with no runbook and no agreed metrics. What does the correct answer insert?

    • A. Agreed measurable success criteria before proceeding.
    • B. A runbook and a canary/phased rollout with rollback before full deployment.
    • C. Immediate full deployment to gather data.
    • D. Skipping monitoring to reduce overhead.
    • E. Per-engineer metric definitions.
    Show answer

    Answer: A and B.

    The missing gates are agreed criteria and a runbook plus canary with rollback. Full deployment (C) is high blast radius; skipping monitoring (D) blinds ops; per-engineer metrics (E) destroy a shared definition of success.

  61. Q61D7 · Developer Productivity and Operational EnablementSelect one

    A platform team wants a coding standard and denied destructive commands enforced across all repos with no override. Best mechanism?

    • A. Ask developers to add rules to CLAUDE.local.md.
    • B. Managed/enterprise policy plus checked-in ./CLAUDE.md and .claude/settings.json with permissions.deny.
    • C. A shared Slack message.
    • D. A quarterly email.
    Show answer

    Answer: B.

    Non-overridable, org-wide enforcement is what managed policy provides, with checked-in project config as the baseline. CLAUDE.local.md (A) is git-ignored/overridable; Slack (C) and email (D) aren't enforcement.

  62. Q62D7 · Developer Productivity and Operational EnablementSelect one

    Setting up automated diff review in CI. Correct approach?

    • A. Run interactive Claude and paste diffs by hand.
    • B. Headless mode: claude -p '…' --output-format json with restricted --allowedTools and an appropriate --permission-mode.
    • C. Give the CI agent all tools.
    • D. Disable permissions in CI.
    Show answer

    Answer: B.

    CI is non-interactive → headless with structured output and restricted tools. Interactive use (A) can't run in CI; all-tools (C) and disabled permissions (D) break least privilege.

  63. Q63D7 · Developer Productivity and Operational EnablementSelect one

    A manager wants to measure AI productivity by lines of code generated and suggestions accepted. Guidance?

    • A. Those are good primary metrics.
    • B. Measure delivery outcomes (cycle time, change-failure/rework rate, review quality); lines and acceptances are vanity metrics.
    • C. Measure tokens consumed.
    • D. Don't measure at all.
    Show answer

    Answer: B.

    Productivity ties to delivery outcomes, not volume. Lines/acceptances (A) and tokens (C) are vanity signals; not measuring (D) forfeits value demonstration.

Last updated Sep 18, 2026