Question Anatomy
How scenario-based items are constructed, how distractors are written, and a repeatable technique for answering them.
The item template
Almost every item on these exams follows the same skeleton:
[Context] A team / company / role and what they are trying to achieve.[Constraint] One or two explicit constraints: cost, latency, compliance, scale, reliability, team size, existing stack.[Signal] A detail that points to the correct principle (e.g., "overnight", "must never", "10,000 documents", "confident but wrong").[Question] "Which approach BEST…", "What should the architect do FIRST…", "Which TWO actions…" (multiple-response items state the count).[Options] 3–4 plausible options. Usually: • the correct trade-off • a correct-sounding but over-engineered option • an option that ignores the constraint • an option that violates a stated principle (an anti-pattern)How distractors are written
Item writers produce wrong answers by applying predictable transformations to the right one. Learn to recognise them:
| Distractor type | Example | Why it fails |
|---|---|---|
| Constraint-blind | Recommends realtime API when the stem says “overnight, cost matters” | Ignores the signal |
| Over-engineered | Multi-agent system for a linear three-step task | Violates “simplest thing that works” |
| Prompt-as-enforcement | “Add to the system prompt: never issue refunds over $500” | Critical rules need programmatic hooks |
| Self-report reliance | “Escalate when Claude’s self-rated confidence is below 0.7” | Self-reported confidence is not calibrated |
| Silent failure | “Return an empty list if the tool errors” | Hides diagnostic context |
| Recall-only | Correct definition, wrong situation | Exams test application, not definition |
| Aggregate metric | “Overall accuracy is 96%, ship it” | Masks per-segment failure |
| Sentiment = complexity | “Escalate angry customers” | Sentiment ≠ complexity |
A repeatable technique
- Read the question sentence first (the last line before the options). Note the qualifier: BEST, FIRST, MOST cost-effective, TWO.
- Read the stem and underline the constraint and the signal. There is almost always exactly one decisive detail.
- Predict the principle before looking at options (“this is a batch-vs-realtime question”).
- Eliminate the constraint-blind and anti-pattern options – usually two of four.
- Between the remaining two, pick the simpler one that fully satisfies the constraint. Over-engineering is a distractor pattern, not a virtue.
- Multiple-response: the count is given. Each correct option must independently be right; do not pick options that are only right “together”.
- Flag and move on after ~2.5 minutes. Return with leftover time.
Worked example
A financial-services team is building a support agent with Claude. Refunds above $500 must never be issued without a human approval. The team proposes adding the rule to the system prompt. What should the architect recommend?
A. Keep the rule in the system prompt and add few-shot examples of correct refusals. B. Implement a pre-tool-use hook that blocks the
issue_refundtool when amount > 500 unless an approval token is present. C. Ask the model to output a confidence score and only issue refunds when confidence > 0.9. D. Lower the temperature to 0 so the model follows the instruction deterministically.
Show the answer and reasoning
B. The signal is “must never” plus a financial threshold: a critical business rule. Prompt-based enforcement (A, D) is best-effort, not a guarantee – temperature 0 does not make instruction-following deterministic. Self-reported confidence (C) is uncalibrated. Programmatic hooks enforce the rule deterministically outside the model. This maps to anti-pattern 3 (“prompt-based enforcement for critical business rules”) and anti-pattern 4 (“self-reported confidence”).
Time budget
| Exam | Items | Seconds per item | Suggested first pass | Reserve |
|---|---|---|---|---|
| CCAO-F | 60 | 120 | 95 min | 25 min |
| CCDV-F | 53 | 136 | 95 min | 25 min |
| CCAR-F | 60 | 120 | 95 min | 25 min |
| CCAR-P | 63 | 114 | 100 min | 20 min |
The reserve is for flagged items and a final check that no item is left blank (no guessing penalty).
Last updated Sep 18, 2026