CCAO-F Practice Exam
A 60-item, blueprint-weighted practice exam for the Claude Certified Associate – Foundations credential, with full explanations and an answer key.
This is a full-length, blueprint-weighted practice exam for CCAO-F. All 60 questions are new and do not repeat the domain-page items.
Instructions
- Time: 120 minutes, matching the real exam.
- Items: 60, multiple-choice and multiple-response. Each item states how many answers to select.
- Scoring: scaled 100–1000, pass 720/1000. There is no guessing penalty. As a rough guide, aim for ≥ 80% raw (≈ 48/60) before sitting the real exam.
- Rule: answer every question. For multiple-response items you must select all correct options and no incorrect ones.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Prompting and Task Execution | 14% | 8 |
| 2 | Output Evaluation and Validation | 21% | 13 |
| 3 | Product and Model Selection | 12% | 7 |
| 4 | Workflow Integration and Solution Design | 16% | 10 |
| 5 | Configuration and Knowledge Management | 12% | 7 |
| 6 | Governance, Risk, and Responsible Use | 15% | 9 |
| 7 | Troubleshooting and Optimization | 10% | 6 |
Total: 8 + 13 + 7 + 10 + 7 + 9 + 6 = 60 items.
Score interpretation
| Raw score (of 60) | Approx. scaled | Interpretation |
|---|---|---|
| 54–60 | 900–1000 | Exam-ready; strong across all domains |
| 48–53 | 800–890 | Solid pass zone; review weak domains |
| 43–47 | 720–790 | Marginal pass; targeted revision advised |
| 36–42 | 600–710 | Below pass; revisit heavy domains (D2, D4, D6) |
| < 36 | < 600 | Study all domains before re-attempting |
Pass mark is 720/1000 (≈ 43/60 raw), but aim for ≥ 48/60 in practice to leave a safety margin.
Take the practice exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A user prompts 'Write a job description' and gets a generic result. Which addition would help MOST?
Show answer
Answer: A.
The output is generic because the prompt names a topic, not a job; supplying role/responsibilities/skills/tone/length fixes the root cause. Length (B), model tier (C) and 'write it well' (D) do not add the missing specifics.
A request bundles 'research, outline, draft, and proofread a whitepaper' and quality is uneven. What is BEST?
Show answer
Answer: A.
Bundled, dependent deliverables should be decomposed with review between stages. Nagging (B) is fragile, thinking alone (C) does not split the work, and shortening (D) drops requirements.
Which follow-up is the BEST example of targeted iteration?
Show answer
Answer: B.
Good iteration says exactly what to keep and what to change. 'Make it better', 'try again' and 'regenerate' (A, C, D) invite random variation.
A subtle three-way classification is inconsistent under a zero-shot prompt. What is the BEST improvement?
Show answer
Answer: A.
Few-shot examples lock in subtle categories and format. 'Focus' (B) is not actionable; length (C) and more chats (D) do not fix category confusion.
Which prompt produces the MOST checkable, spreadsheet-ready output?
Show answer
Answer: B.
A defined schema, column order and missing-value rule makes the output structured and checkable. The others leave shape and completeness undefined.
To extract figures from a supplied report without fabrication, which TWO techniques are BEST?
Show answer
Answer: A and B.
Grounding to the document and a rule for missing values reduce fabrication. 'Be accurate' (C) is not actionable; model tier (D) and length (E) do not enforce grounding.
A request is ambiguous ('prepare the update'). What BEST prevents wasted rework?
Show answer
Answer: B.
Resolving ambiguity up front avoids generating on a wrong assumption. Generating first (A, C) risks solving the wrong problem; length (D) is irrelevant.
A brainstorming task asks for 'the one best idea' immediately. Why is this weak?
Show answer
Answer: A.
Divergent generation of many varied options precedes convergence in brainstorming. It is genuinely weak (B); JSON (C) and tighter length (D) work against divergence.
Claude returns five precise, unsourced statistics on a niche market. What is BEST?
Show answer
Answer: C.
Precise, unsourced numerics on niche topics are the hallucination signature; verify independently. Consistency (A) is not truth; same-session self-check (B) shares the bias; temperature (D) is about variance.
A summary of a supplied contract states a 90-day notice period. How BEST to validate?
Show answer
Answer: B.
Provenance test: a supplied-document claim is verified by locating the supporting text. Supplying the doc (A) does not guarantee correct extraction; two models agreeing (C) is weak; re-summarising (D) tests stability, not accuracy.
A customer email drafted by Claude names a third-party vendor as the cause of an outage. Which TWO are required before sending?
Show answer
Answer: A and B.
External, irreversible, third-party-naming content needs factual verification and the organisational review gate. Tone (C) is optional; sending now (D) skips review; temperature (E) is irrelevant.
A vendor comparison lists six strengths for A and six weaknesses for B. What is the issue and fix?
Show answer
Answer: B.
Asymmetric treatment is framing bias; a neutral symmetric structure fixes it. Sources (A) may help but do not fix asymmetry; it is not truncation (C) or a totals issue (D).
Which output should receive line-by-line human verification?
Show answer
Answer: A and C.
Regulated filings and medical content are high-stakes triggers. Speed, formatting and length (B, D, E) are irrelevant to stakes.
Claude computes a quarterly total from a pasted table and it 'looks right'. Best verification?
Show answer
Answer: B.
LLMs can err on arithmetic; recompute independently for financial figures. Arithmetic is not reliably deterministic for a language model (A); same-session doubt (C) and rounding (D) do not verify.
In research mode, 2 of 8 citations checked at random — one is fabricated. What now?
Show answer
Answer: C.
A fabricated citation invalidates sample-based trust; move to exhaustive verification. Partial trust (A, B) is unsafe; inventing a citation (D) compounds the problem.
A user notices Claude agrees with whatever view they state. What is this and the best mitigation?
Show answer
Answer: B.
Agreement with the user's position is sycophancy; neutral framing and independent critique counter it. It is not hallucination (A) or truncation (C); a longer prompt (D) does not help.
Interview-feedback summaries will feed hiring decisions. What is REQUIRED?
Show answer
Answer: B.
Hiring is regulated and fairness-sensitive; a human-review gate is mandatory. 'Internal' (A) does not exempt it; numeric ratings (C) can launder bias; regeneration (D) does not add review.
Which output format is BEST for 300 records with fixed fields destined for a CRM import?
Show answer
Answer: C.
Machine-consumed, fixed-field data needs a specified structured format. Prose (A), narrative (B) and slides (D) are not importable structured data.
A translated policy renders 'processor' three different ways. What lens and fix?
Show answer
Answer: B.
Inconsistent rendering of a defined term is an internal-consistency failure; a glossary and terminology table fix it. It is not an accuracy-sourcing (A), length (C) or bias (D) issue.
An executive brief will be built from a long analysis. Best output and check?
Show answer
Answer: B.
Audience fit (decision first, one page), format fit (iterable artifact) and a consistency + accuracy check. Pasting everything (A) ignores audience; JSON (C) misfits prose; more detail (D) is the wrong direction.
Why is 'the output was long and well-formatted, so send it' wrong?
Show answer
Answer: B.
Fluency is not accuracy; verification is still required. Length does not imply accuracy (A); formatting does not reduce it (C); and the statement is indeed wrong (D).
A summary table's row values do not sum to its stated total. What lens caught this and what is the cheapest check?
Show answer
Answer: B.
Totals not matching rows is an internal-consistency failure, checkable by recomputation without external sources. External sourcing (A) is unnecessary here; relevance (C) and bias (D) are unrelated.
What is the safest way to double-check a high-stakes Claude answer?
Show answer
Answer: B.
Independent verification avoids same-session self-review bias. Asking the same chat (A) and regenerating in it (C) carry the same bias; model size (D) does not verify.
The same brand context is reused across dozens of chats weekly. Best management?
Show answer
Answer: B.
Recurring, shared context is what Projects are for. Pasting each time (A) is inconsistent; model tier (C) is unrelated; and reliance on recall across new chats (D) is unreliable.
8,000 short messages must be tagged quickly and cheaply. Best model?
Show answer
Answer: D.
High-volume, low-complexity, latency-sensitive tagging is the archetypal Haiku task. The higher tiers over-buy cost and speed.
A complex, board-bound scenario analysis needs the best reasoning. Best choice?
Show answer
Answer: B.
Hard reasoning + high stakes justifies the top tier and extended thinking, plus verification. Haiku (A) under-buys; Sonnet with thinking off (C) under-serves; research mode (D) does not do the analysis.
A two-hour chat is now forgetting earlier decisions. Best action?
Show answer
Answer: B.
Long, drifting chats should be reset with a summarised brief. Continuing (A) compounds drift; model tier (C) and thinking (D) do not clear bloat.
Which needs a Team/Enterprise plan?
Show answer
Answer: B.
Org governance controls live in Team/Enterprise. The individual/casual uses (A, C, D) do not require org-level administration.
A current, sourced regulatory summary is needed. Best surface?
Show answer
Answer: B.
Current, sourced facts require research mode with citation verification. Built-in knowledge (A) may be stale; the code tool (C) computes; artifacts (D) are a container.
Which TWO tasks best fit the code execution / analysis tool?
Show answer
Answer: A and B.
Reliable arithmetic over many rows and chart generation are executed, not estimated. Emails, brainstorming and translation are ordinary language tasks.
In a mapped process, which step should stay FULLY human?
Show answer
Answer: B.
Approving a contract is a consequential, often-irreversible decision that stays human. Summarising, drafting and reformatting (A, C, D) are assisted, reversible steps.
A refund-email process is proposed to run fully automated on Claude's judgement. Best design?
Show answer
Answer: B.
Irreversible, financial actions require a human approval gate. Full automation (A) removes it; self-rated confidence (C) is not a valid gate; model tier (D) does not remove risk.
Connecting an email account for reply drafting — which TWO apply?
Show answer
Answer: A and B.
Least privilege and access-inheritance awareness apply. Connecting everything (C) over-exposes; connectors have governance implications (D); model size (E) is irrelevant to scope.
A nightly, unattended invoice-to-finance-system pipeline is needed. Correct response?
Show answer
Answer: B.
Automatic, system-to-system, unattended processing escalates to Developers/Architects. A Project (A) does not run automation; manual work (C) defeats the goal; research mode (D) is unrelated.
Which statement to leadership is MOST appropriate?
Show answer
Answer: B.
Concrete value, verification, human ownership and a stated limitation — honest and accurate. The others overpromise (A, C) or misstate accountability (D).
A pilot cut handling time; a leader asks if quality held. Best metric approach?
Show answer
Answer: B.
Per-segment quality metrics reveal categories an average hides. Overall satisfaction (A) and assuming quality (C) mask failures; sign-ups (D) measure adoption.
Best FIRST pilot candidate?
Show answer
Answer: B.
A reversible, internal, low-risk, language-heavy task is the ideal pilot. The others are regulated, irreversible or financial.
Two prerequisites BEFORE scaling a successful pilot org-wide?
Show answer
Answer: A and B.
Responsible scaling needs evidence of impact and governance scaffolding. Standardising on the priciest model (C) wastes cost; removing review (D) is unsafe; banning connectors (E) is not a prerequisite.
Where does Claude add MOST value in requirements analysis?
Show answer
Answer: B.
Synthesis into a structured draft is high-value assistance; humans decide. Investment (A), budget (C) and vendor commitments (D) are human decisions.
A recurring, human-driven task needs shared instructions/docs but no automation. Right solution and owner?
Show answer
Answer: B.
Recurring, human-driven work with shared context is a business-user Project/Skill owned by the Associate. It needs no developer escalation (A) or autonomous agent (C); manual pasting (D) is inefficient.
An HR assistant cites an outdated policy; both old and new handbooks are in knowledge. Best fix?
Show answer
Answer: B.
Conflicting versions are a curation problem; keep one authoritative source. A bigger model (A) cannot pick authority; a prompt (C) does not reliably override a stale source; new Projects (D) ignore the duplicated files.
To ensure an assistant NEVER exposes salary figures, what is MOST reliable?
Show answer
Answer: B.
Hard guarantees come from not making the data reachable (access control), not instruction text alone. An instruction (A) guides but does not enforce; model tier (C) is irrelevant; relying on users (D) is not a control.
A recurring 'standard weekly status report' with fixed sections and format needs to be reusable. Best mechanism?
Show answer
Answer: A.
A repeatable multi-step procedure is a Skill. A long prompt (B) is manual; model tier (C) is unrelated; Memory (D) stores context, not a procedure.
Two principles of good connector management?
Show answer
Answer: A and B.
Least privilege and periodic review/revocation are core. Whole-account connection (C) over-exposes; connectors have governance implications (D); ownership must be documented (E).
A company-wide code-of-conduct assistant must behave identically for all. Which scope?
Show answer
Answer: B.
A company-wide, consistent assistant belongs at org scope with central governance. Personal scopes (A, C, D) cannot guarantee consistency or apply org controls.
Likely consequence of dumping every document into knowledge 'for completeness'?
Show answer
Answer: B.
Over-stuffed knowledge reduces quality and raises conflict risk. It does not sharpen answers (A), affect speed (C) or replace connectors (D).
Two maintenance practices for a Project set up a year ago and never revisited?
Show answer
Answer: A and B.
Configurations need ownership and a review cadence with a change log. 'Never change it' (C) lets it rot; removing oversight (D) and deleting after use (E) are not maintenance.
An employee pastes a confidential roadmap into a personal free AI account to work faster. Best characterisation?
Show answer
Answer: B.
Confidential data in an unapproved account is shadow AI, bypassing controls regardless of output. Deletion (C) does not undo exposure; roadmaps are confidential even if not PII (D).
A clinic wants Claude to help with PHI. What matters MOST first?
Show answer
Answer: B.
PHI triggers HIPAA/policy questions to settle before use, defaulting to caution. Cost (A), format (C) and speed (D) are irrelevant to the legal/policy question.
A flawed email is blamed on 'the AI'. Which principle applies?
Show answer
Answer: B.
Human accountability is the anchor principle. 'The AI decided' is never a defence (A, D); tier is irrelevant (C).
Two reasons to use a commercial plan for business data?
Show answer
Answer: A and B.
The no-training guarantee and admin controls justify commercial tooling. Paid plans do not change intelligence (C); personal plans lack org governance (D); no plan removes human review (E).
Responsible use of Claude to screen candidates requires what?
Show answer
Answer: B.
Hiring is fairness-sensitive and regulated; humans stay accountable and review for bias. Auto-reject (A) launders bias and removes accountability; models are not inherently neutral (C); cost (D) is beside the point.
Unsure whether a dataset is restricted and policy-allowed. Best action?
Show answer
Answer: B.
Uncertainty about restricted data is resolved by escalation, defaulting to caution. Proceeding on a guess (A, D) risks a violation; Claude is not an authoritative policy source (C).
Before publishing a Claude-drafted external blog post, what must happen?
Show answer
Answer: B.
The publisher is accountable for accuracy and IP. Generated content is not automatically safe (A); spell check (C) is insufficient; publish-first (D) is irreversible and unsafe.
Which is clearly an INAPPROPRIATE use case?
Show answer
Answer: B.
Deceptive impersonation to mislead is prohibited. The others are ordinary appropriate tasks.
Company policy requires disclosing AI assistance in client deliverables; an Associate omits it to impress. Issue?
Show answer
Answer: B.
Where policy/norms require disclosure, omitting it is a violation. It is not optional (A), not limited to regulated industries (C), and broader than copyright (D).
A prompt returns a generic answer. What should happen FIRST?
Show answer
Answer: B.
Diagnosis precedes action; fix the identified cause in the prompt. A bigger model (A) does not add context; regenerating (C) varies wording; thinking (D) does not fix a vague ask.
A long report keeps getting cut off. Two best responses?
Show answer
Answer: A and B.
Truncation is handled by continuing or sectioning. Retrying (C) repeats the limit; model tier (D) and thinking (E) do not address length.
A long multi-topic chat now contradicts earlier decisions and ignores the format. Best fix?
Show answer
Answer: B.
Context bloat/drift is cured by capturing state and resetting. Correcting in place (A) fights the bloat; model tier (C) and thinking (D) do not clear it.
A prompt bundles five deliverables and consistently drops two. Cause and fix?
Show answer
Answer: B.
Dropping parts of a bundled ask is the overloaded-prompt pattern; decomposition fixes it. It is not truncation of one output (A), a refusal (C) or a model issue (D).
Last updated Sep 18, 2026