CCAO-F Practice Exam 2
A second full-length, blueprint-weighted practice exam for the Claude Certified Associate – Foundations credential, with new, harder items.
This is a second full-length, blueprint-weighted practice exam for CCAO-F. All 60 items are new – they do not repeat Practice Exam 1 or the domain-page questions. Exam 2 is deliberately slightly harder: more stems carry multiple constraints at once, and more use FIRST, MOST cost-effective, and TWO qualifiers, so you must weigh competing considerations rather than spot a single keyword. Distractors lean on the classic patterns – constraint-blind, over-engineered, prompt-as-enforcement, self-report reliance, silent failure, recall-only and aggregate-metric.
Instructions
- Time: 120 minutes, matching the real exam.
- Items: 60, multiple-choice and multiple-response. Each item states how many answers to select.
- Scoring: scaled 100–1000, pass 720/1000. There is no guessing penalty. Aim for ≥ 80% raw (≈ 48/60) before booking.
- Rule: answer every question. For multiple-response items you must select all correct options and no incorrect ones.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Weight | Items here |
|---|---|---|---|
| 1 | Prompting and Task Execution | 14% | 8 |
| 2 | Output Evaluation and Validation | 21% | 13 |
| 3 | Product and Model Selection | 12% | 7 |
| 4 | Workflow Integration and Solution Design | 16% | 10 |
| 5 | Configuration and Knowledge Management | 12% | 7 |
| 6 | Governance, Risk, and Responsible Use | 15% | 9 |
| 7 | Troubleshooting and Optimization | 10% | 6 |
Total: 8 + 13 + 7 + 10 + 7 + 9 + 6 = 60 items.
How to use both exams
- Sit Practice Exam 1 first, untimed if you are still learning, to surface weak domains.
- Study the domain pages for any domain where you scored below ~75%, paying attention to the “Common misconceptions” tables and scenario walkthroughs.
- Use Practice Exam 2 as your booking gate: sit it timed (120 minutes) under exam conditions. If you clear ≥ 48/60 here, with no domain badly lagging, you are ready to book.
- Compare per-domain results between the two exams. A domain that is strong on Exam 1 but weak on the harder Exam 2 is a shallow-understanding signal – revisit the decision tables and distractor patterns for that domain.
Score interpretation
| Raw score (of 60) | Approx. scaled | Interpretation |
|---|---|---|
| 54–60 | 900–1000 | Exam-ready; strong across all domains |
| 48–53 | 800–890 | Solid pass zone; review weak domains |
| 43–47 | 720–790 | Marginal pass; targeted revision advised |
| 36–42 | 600–710 | Below pass; revisit heavy domains (D2, D4, D6) |
| < 36 | < 600 | Study all domains before re-attempting |
Because Exam 2 runs harder than Exam 1, treat a comfortable pass here as a stronger readiness signal than the same score on Exam 1.
Take the practice exam
The interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
60 questions · one at a time · 120-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
An associate pastes a customer email and types in one run-on line: 'reply to this and also tell me whether we broke the SLA'. Claude drafts a warm reply but never mentions the SLA. What is the BEST FIRST fix?
Show answer
Answer: B.
The two asks were jumbled with the source text, so one was dropped; separating and labelling the parts and ordering the tasks fixes the root cause. A bigger model (A) still faces the jumble (constraint-blind); regenerating (C) is blind variation; 'be more thorough' (D) is a wish, not a specification (prompt-as-effort).
A four-part request (research competitors, write a brief, draft headlines, build a timeline) consistently returns a strong brief but a weak, generic timeline. Which approach is MOST appropriate?
Show answer
Answer: B.
Uneven quality on a bundled ask is the overloaded-prompt pattern; isolating the weak part as its own well-specified step fixes it. 'Make it better' (A) is a non-actionable wish; a bigger model (C) does not address bundling (over-engineered); more words (D) is the wrong lever.
An operations lead needs figures pulled from an attached 30-page report and wants to minimise both fabrication AND unusable formatting. Which TWO prompt techniques are BEST?
Show answer
Answer: A and B.
Grounding to the document with a missing-value rule reduces fabrication, and a defined schema fixes formatting. 'Be accurate and well-organised' (C) is not actionable; model tier (D) does not enforce grounding (over-engineered); maximum length (E) is the wrong lever.
A new joiner opens a brand-new chat and types 'continue the analysis from yesterday', and Claude has no idea what they mean. What is the correct explanation and fix?
Show answer
Answer: B.
A fresh chat has no memory of a prior one unless context is persisted in a Project or Memory; the fix is to supply or persist it. Model tier (A) does not restore missing context; length (C) is irrelevant; a self-check (D) is self-report reliance and cannot recover data that is not present.
A brainstorming prompt keeps returning three safe, similar event themes when the associate needs a wide range. Which single change helps MOST?
Show answer
Answer: B.
Brainstorming rewards explicit quantity and diversity before convergence; stating '15 distinct, different styles, no filtering' widens the divergent phase. 'Better' (A) is a non-actionable wish; model tier (C) is irrelevant; asking for one idea (D) collapses divergence — a task-type mismatch.
Despite a written description of the required layout, a recurring report is formatted slightly differently each time. What is the MOST reliable fix?
Show answer
Answer: B.
When a described format will not hold, one concrete example (few-shot) pins the pattern far more reliably. More words (A) and 'be consistent' (C) are not actionable; a higher temperature (D) increases variation, worsening the problem.
A request is ambiguous — 'prepare the quarterly update' — with no audience, scope, or format named. Which approach BEST prevents wasted rework?
Show answer
Answer: B.
Resolving ambiguity up front avoids generating on a wrong assumption. Generating first (A, C) risks solving the wrong problem and wastes iterations; maximum length (D) is unrelated to the missing specifications.
An extraction prompt asks Claude to 'pull all action items' from meeting notes, but the output omits owners and due dates. What is the BEST refinement?
Show answer
Answer: A.
Extraction needs an explicit schema and a rule for missing values, which makes it complete and checkable. Length (B) does not define the missing fields; extended thinking (C) does not add a schema; re-pasting (D) changes nothing about the fields requested.
A board pack is due in an hour. It reads fluently, but the segment table's rows do not sum to the headline total AND one of five citations cannot be found anywhere. What should the analyst do FIRST?
Show answer
Answer: B.
The cheapest, highest-value first move is the internal-consistency recompute, which needs no external source. Sending on fluency (A) is the fluency trap; a same-session self-check (C) reuses the biased reasoning; lowering temperature and regenerating (D) changes variance, not grounding, and wastes the hour.
Claude cites a study with a plausible title, real-sounding authors and a real journal name to support a key claim. Before relying on it, what is the MOST appropriate validation?
Show answer
Answer: B.
Fabricated citations blend real-sounding elements, so the only valid check is opening the source and confirming it supports the claim. A real journal name (A) proves nothing; a same-chat DOI (C) can also be fabricated (self-report reliance); regeneration (D) tests stability, not existence.
A fully accurate three-page analysis is rejected by a CFO who says she cannot find the decision. Which lens failed, and what is the fix?
Show answer
Answer: B.
The facts are correct but the output is mis-pitched for an executive — a relevance/audience failure, fixed by leading with the decision. Accuracy (A) already passed; there is no totals problem (C) or asymmetry (D) described, so those lenses are distractors.
A dashboard shows 92% overall accuracy for a Claude-assisted tagging task and a manager wants to expand it. Which TWO checks are MOST important FIRST?
Show answer
Answer: A and B.
An aggregate can hide a category that fails badly, so per-segment breakdown and checking whether failures concentrate in high-stakes categories are essential. Reporting only the headline (C) and assuming uniformity (D) are aggregate-metric traps; cost (E) is unrelated to validation.
A colleague says an internal-only memo's revenue figure 'doesn't need checking because it's internal' — but it will feed next year's hiring plan. What is the BEST response?
Show answer
Answer: B.
Stakes follow the downstream use, not the 'internal' label; a figure feeding hiring must be verified. 'Internal' does not exempt it (A); tone (C) is the wrong lens; a same-chat confidence check (D) is self-report reliance, not verification.
Two colleagues verify a high-stakes number by asking two different AI tools; both return the same value, so they treat it as confirmed. What is the FLAW?
Show answer
Answer: B.
Two models agreeing is not independent verification because they can be wrong the same way; an authoritative source or human is needed. Agreement is not confirmation (A); repeating one tool (C) or adding a third AI (D) still lacks authority.
In research mode, an analyst spot-checks 2 of 8 citations; one is fine and one does not exist. What is the MOST appropriate next step?
Show answer
Answer: C.
A fabricated citation invalidates sample-based trust; move to exhaustive verification. Partial trust (A) and fixing only one paragraph (B) leave other fabrications in place; inventing a citation (D) compounds the problem.
A vendor comparison lists six strengths for Vendor A and six weaknesses for Vendor B. Which is the MOST likely issue and the BEST remedy?
Show answer
Answer: B.
Asymmetric treatment of options is framing bias, fixed by a neutral, criteria-based symmetric structure. Sources (A) may help but do not fix the asymmetry; it is not truncation (C) or a totals problem (D).
Which TWO outputs should receive line-by-line human verification before use?
Show answer
Answer: A and C.
A regulatory filing and medical dosage content are high-stakes triggers requiring exhaustive review. Generation speed (B), bullet formatting (D) and length (E) are irrelevant to stakes.
A financial analyst asks Claude to compute a quarterly total from a pasted 500-row table, and it 'looks about right'. What is the MOST appropriate verification?
Show answer
Answer: B.
Language models can err on long-table arithmetic, so financial figures should be recomputed independently. Arithmetic is not reliably deterministic for an LLM (A); a same-chat doubt (C) is self-report reliance; rounding (D) hides rather than verifies.
A translated policy renders the defined term 'data controller' three different ways. Which lens caught this, and what is the cheapest fix?
Show answer
Answer: B.
Inconsistent rendering of a defined term is an internal-consistency failure, fixed with a glossary and a checkable terminology table. It is not an accuracy-sourcing (A), length (C) or bias (D) issue.
Claude answers three of four sub-questions in a brief thoroughly and silently omits the fourth, yet the brief reads as complete. Which lens catches this and how?
Show answer
Answer: B.
A silently dropped sub-question is a completeness failure, caught by checking the output against the brief as a checklist. Accuracy sourcing (A), bias balancing (C) and audience changes (D) do not detect a missing part.
Which output format is MOST appropriate for 300 customer records with fixed fields destined for a CRM import?
Show answer
Answer: C.
Machine-consumed, fixed-field data needs a specified structured format with the exact schema. Prose (A), a narrative artifact (B) and slides (D) are not importable structured data.
A strategy team needs BOTH the same brand tone and reference docs across many chats AND a fixed weekly-report procedure run identically each time. Which combination is BEST?
Show answer
Answer: B.
Shared context is a Project's job; a repeatable multi-step procedure is a Skill's job — the two are complementary. A long prompt (A) and Memory (C) cover neither well; separate accounts (D) fragment the setup and governance.
An associate must accurately sum a 4,000-row sales spreadsheet AND obtain a current, sourced view of a new regulation. Which TWO surfaces are correct, respectively?
Show answer
Answer: A and B.
Reliable arithmetic over many rows is executed by the analysis tool; current sourced facts come from research mode with citation checks. Built-in knowledge (C) may be stale; research mode does not do arithmetic (D); an artifact (E) is a container, not a compute engine.
A subject-line rewrite must run thousands of times a day at the lowest cost and latency. Which model AND thinking setting is MOST cost-effective?
Show answer
Answer: B.
A trivial, high-volume, latency-sensitive task wants the cheapest fast model with no extra thinking. Opus and Fable (A, C) over-buy cost and latency; extended thinking on Sonnet (D) adds cost for no quality gain here.
A manager claims that switching from the web app to Claude Desktop will let the team safely use confidential data that policy currently disallows. Which statement is CORRECT?
Show answer
Answer: B.
The plan/account determines governance and the no-training guarantee; changing client does not change what data is allowed. Desktop is not inherently more compliant (A, C), and deletion (D) does not alter policy permissions.
A workload includes a complex, high-stakes board scenario analysis AND a routine batch of 5,000 ticket tags that must be produced cheaply. What is the BEST model strategy?
Show answer
Answer: B.
Match each task to its stakes: top tier plus thinking for hard high-stakes reasoning, cheapest fast model for routine high-volume tagging. One model for both (A) mismatches at least one task; Haiku for the board analysis (C) under-buys; Opus for tagging (D) over-buys.
A user needs Claude to act on the web page they are currently viewing to help complete a long form. Which client is MOST appropriate?
Show answer
Answer: B.
In-page assistance on the current site is what the Chrome extension provides, with permission. The mobile app (A) is for on-the-go capture; Claude Code (C) is developer tooling; research mode (D) does web synthesis, not acting on the open page.
A user wants Claude to recall their personal writing preferences across sessions, while their team wants a shared, curated set of policy docs available to all. Which pairing is CORRECT?
Show answer
Answer: A.
Memory handles personal cross-session recall; curated shared documents belong in a Project's knowledge. Swapping them (B), putting shared docs in personal Memory (C), or hiding shared docs in a personal Project (D) mis-scopes the need.
A support lead wants Claude to auto-send simple refund confirmations 'because it's confident on those' and to connect the whole shared inbox for convenience. Which TWO corrections are MOST appropriate?
Show answer
Answer: A and B.
Irreversible financial sends need a human gate, and connectors must follow least privilege because the inbox holds PII. Self-rated confidence (C) is self-report reliance, not a valid gate; whole-inbox connection (D) over-exposes; model size (E) does not remove risk.
A pilot cut handling time 40% with steady overall CSAT, and the manager wants to roll it out to all 12 teams next week. What should happen FIRST?
Show answer
Answer: B.
Time and aggregate CSAT can hide a weak segment; verify per-type quality and put governance in place before a phased rollout. Immediate big-bang rollout (A) skips both; a pricier model (C) is unrelated; removing review (D) is unsafe.
Which need CLEARLY crosses the escalation boundary to Developers/Architects?
Show answer
Answer: B.
Automatic, system-to-system, unattended 24/7 processing needs engineered integration — Developer/Architect territory. The others keep a human in the loop and stay in the business-user zone.
Mapping a claims-processing workflow, which TWO steps should stay FULLY human?
Show answer
Answer: A and C.
A payout decision and sending an external, irreversible decision letter are consequential human steps. Summarising (B), reformatting (D) and drafting internal notes (E) are reversible, language-heavy steps Claude assists with under review.
A stakeholder asks the Associate to promise the new workflow will 'basically run itself'. What is the BEST response?
Show answer
Answer: B.
Honest framing pairs concrete value with limitations, a review/ownership gate, and a feedback loop. Promising autonomy (A, C) overpromises and loses trust on the first error; refusing to communicate value (D) discards real benefit.
A team measures a pilot ONLY by hours saved. What is the RISK, and the fix?
Show answer
Answer: B.
A time-only metric can hide quality regressions and rework; per-task-type quality metrics complete the picture. Time is not the only metric (A); sign-ups (C) measure adoption, not quality; stopping measurement (D) removes the evidence needed to scale.
A manager wants to automate sending customer refund emails end-to-end with no human, triggered by Claude's judgement. What is the BEST design?
Show answer
Answer: B.
Irreversible, financial actions require a human approval gate. Full automation (A) removes it; self-rated confidence (C) is self-report reliance; model tier (D) does not remove the risk.
Which step is the BEST FIRST candidate for a Claude pilot?
Show answer
Answer: B.
A reversible, internal, language-heavy, low-risk task is the ideal first pilot. The others are regulated, irreversible or financial and are poor first pilots.
A CRM connector is proposed so Claude can draft personalised outreach. What is the KEY governance consideration?
Show answer
Answer: B.
A CRM connector brings customer PII into scope, triggering data-handling policy and least-privilege scoping. It is not harmless (A); model choice (C) is secondary; automation (D) is not required and would raise risk.
An Associate is planning the rollout of a validated pilot. Which sequence reflects responsible scaling?
Show answer
Answer: B.
Responsible scaling measures, standardises, governs, then phases out with monitoring. Scaling before governance (A) is unsafe; the priciest model plus no review (C) over-buys and removes safeguards; never scaling or documenting (D) forgoes value and repeatability.
A prompt returns a generic, off-target answer. Before anything else, what should the associate do FIRST?
Show answer
Answer: B.
Diagnosis precedes action: match the symptom (generic/off-target) to its cause (ambiguity/missing context) and fix it. A bigger model (A) does not add the missing context (constraint-blind); regenerating (C) varies wording, not relevance (blind regeneration); thinking (D) does not fix a vague ask.
A prompt contains two contradictory instructions ('include every detail' and 'answer in one sentence') and the output is confused. What is the cause and fix?
Show answer
Answer: B.
Contradictory instructions confuse the output; resolve the conflict and set one priority. It is not missing context (A), truncation (C) or a model issue (D) — those misdiagnose the symptom.
Configuring a company-wide HR assistant, an associate has the current handbook, last year's handbook, and a spreadsheet of individual salaries. Which TWO setup choices are correct?
Show answer
Answer: A and B.
Keep one authoritative current source, and make restricted data unreachable rather than trusting an instruction. Both handbooks (C) create conflicts; instruction-only salary protection (D) is prompt-as-enforcement; whole-Drive connection (E) violates least privilege and re-exposes salary data.
An HR assistant Project starts citing an outdated leave policy; both the 2025 and 2026 handbooks are in the knowledge files. What is the BEST fix?
Show answer
Answer: B.
Conflicting versions are a curation problem; keep one authoritative, current source. A bigger model (A) cannot decide authority; a prompt (C) does not reliably override a stale source; new Projects (D) do not address the duplicated files.
A company-wide assistant must behave identically for every employee and be auditable. At which scope should it be configured?
Show answer
Answer: B.
A must-be-identical, auditable, company-wide assistant belongs at org scope with central governance. Personal or copied-around setups (A, C, D) cannot guarantee consistency or apply org-level controls.
A support-triage Project should answer FAQs but hand off billing disputes over $500 to a human. Which instruction element handles this?
Show answer
Answer: B.
Escalation rules define when to hand off rather than answer, the correct mechanism for out-of-scope cases. Role length (A), length limits (C) and tone (D) do not define hand-off behaviour.
An assistant's answers have quietly drifted from current policy over six months, with no record of what changed. Which TWO maintenance practices would have prevented this?
Show answer
Answer: A and B.
Ownership and a review cadence with a change log keep knowledge fresh and traceable. Never editing (C) lets it rot; deleting after use (D) and removing oversight (E) are not maintenance practices.
Before connecting a Slack workspace to a team Project, what should the associate check FIRST?
Show answer
Answer: B.
Connectors inherit the account's access, so the first check is the data scope and its appropriateness, scoped to least privilege. Model cost (A), formatting (C) and chat length (D) are not the governing concern.
An associate proposes adding every departmental document to a Project 'so it knows everything'. What is the MOST likely outcome and the better approach?
Show answer
Answer: B.
Over-stuffing knowledge lowers answer quality and raises conflict risk; curate for relevance and resolve conflicts. It does not sharpen answers (A), affect speed (C) or replace connectors (D).
Under deadline, a recruiter plans to paste 200 CVs (with names and dates of birth) into a personal free AI account because the work account is slow, then delete the chat afterwards. Which TWO statements are CORRECT?
Show answer
Answer: A and B.
Using an unapproved personal account for confidential PII is shadow AI, and deletion cannot reverse an exposure that has already occurred. The output being deleted (C) does not cure the breach, CVs contain personal data (D), and account speed (E) is irrelevant to compliance.
A recruiter wants Claude to rank and automatically reject the bottom half of candidates 'to remove human bias'. What is the BEST response?
Show answer
Answer: B.
Hiring is fairness-sensitive and regulated; auto-rejection launders bias and removes accountability, so humans must decide with bias review. Automation does not remove bias (A); models are not inherently neutral (C); cost (D) is beside the point.
An associate is unsure whether a dataset is 'restricted' and whether policy permits it in Claude. What is the BEST FIRST action?
Show answer
Answer: B.
Uncertainty about restricted data is resolved by defaulting to caution and escalating to the right owner. Proceeding on a guess (A, D) risks a violation; Claude is not an authoritative policy source (C).
Company policy requires disclosing AI assistance in client deliverables, but an associate omits it 'to look more rigorous'. Which statement is CORRECT?
Show answer
Answer: B.
Where policy requires disclosure, omitting it violates policy and the transparency expectation. It is not optional (A), not limited to copyright (C), and not limited to regulated industries (D).
A finance analyst wants to paste customer payment card numbers into a chat to reformat them. What is the BEST response?
Show answer
Answer: B.
Payment card data falls under PCI DSS and policy and must not be pasted into an unapproved tool; minimise and escalate. Reformatting is not harmless (A); model tier (C) and deletion (D) do not make handling PCI data compliant.
A team is deciding between a personal free plan and a Team/Enterprise plan for handling confidential business data. Which TWO reasons favour the commercial plan?
Show answer
Answer: A and B.
The no-training guarantee and admin controls (SSO, audit logs) are why confidential data belongs on commercial tooling. Paid plans do not change intelligence (C), personal plans lack org governance (D), and no plan removes human review (E).
A manager says the flawed customer email 'was the AI's fault, not ours'. Which principle applies and what should follow?
Show answer
Answer: B.
Human accountability is the anchor principle; the sender owns the output and should fix the gap that let the error through. 'The AI decided' is never a defence (A, D); tier is irrelevant (C).
A clinic wants Claude to help with patient records containing PHI. Which consideration is MOST important FIRST?
Show answer
Answer: B.
PHI triggers HIPAA and policy questions that must be settled before use; minimise and escalate if unclear. Model cost (A), format (C) and speed (D) are irrelevant to the governing legal/policy question.
Which scenario BEST illustrates an INAPPROPRIATE use case?
Show answer
Answer: B.
Deceptive impersonation to mislead is prohibited and clearly inappropriate, regardless of tool. The others are ordinary, appropriate business tasks on approved material.
A weekly report chat, excellent two hours ago, now contradicts an earlier decision, changed the table format, invented a figure, and dropped a section. What is the BEST FIRST action?
Show answer
Answer: B.
These are context-bloat/drift symptoms; the cure is to capture state and reset to a clean context. A bigger model with the pasted transcript (A) carries the bloat forward; in-place correction (C) fights the bloat; thinking (D) does not clear it.
To cut cost, a colleague proposes moving legal-contract review to Haiku AND dropping the lawyer sign-off, while leaving high-volume ticket-tagging on Opus. What is the BEST correction?
Show answer
Answer: B.
Protect quality and the review gate on high-stakes contract work; redirect optimisation to the routine, high-volume tagging. Cheaper-is-always-better (A, C) under-buys and removes a gate on high-stakes work; Opus for tagging (D) over-buys.
Claude's long report is repeatedly truncated before it finishes. Which TWO responses are BEST?
Show answer
Answer: A and B.
Truncation is handled by continuing or by sectioning the deliverable. Retrying the same oversized request (C) repeats the limit; model tier (D) and thinking (E) do not address output length.
An associate applied a prompt fix; one run looked great, so they declared success. What is the FLAW and the better practice?
Show answer
Answer: B.
Measuring improvement means before/after comparison, consistency across runs, and per-task-type checks — not a single lucky run. One run (A) can mislead; aggregate-only satisfaction (C) hides segment failures; switching models (D) does not validate the fix.
Last updated Sep 18, 2026