# The 6 Exam Scenarios

The six reference scenarios the CCAR-F draws items from — business context, reference architecture, key decisions with exam-correct answers, distractor anti-patterns, and practice questions for each.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

The Architect – Foundations exam draws its items from **4 of these 6 scenarios**. You will not know which four in advance, so prepare all six. Each scenario below gives the business context, a reference architecture, the key architectural decisions with the **exam-correct** answer, the anti-patterns that appear as **distractors**, and three practice questions.

The scenarios are not independent of the domains — they are how the domains are *tested*. Every question in this course is tagged to a scenario as well as a domain.

---

## Scenario 1 · Customer Support Resolution Agent

**Business context.** A SaaS company wants an agent that resolves common support requests end-to-end — order status, refunds, subscription changes — over chat, handing off to humans only when necessary. Built on the **Agent SDK**, integrating backend systems through **MCP tools**, with an **escalation** path.

```text
Customer ─► [Support Agent · Agent SDK loop]
                 │  stop_reason-driven loop
                 ├─ MCP tools: get_order_status, issue_refund(idempotency_key),
                 │             update_subscription, search_kb
                 ├─ PreToolUse hook: refund > $X requires human approval
                 └─ escalate: explicit request → now; capability gap → after attempt
                                 │
                          [Human agent queue]  (with correlation ID + transcript)
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| Workflow or agent? | **Agent** — resolutions are open-ended, steps vary per request |
| How does the loop terminate? | On **`stop_reason`** (`end_turn`/`tool_use`), with a cap as backstop |
| When to escalate? | **Explicit request** (immediately) or **capability gap** (after attempting) — never sentiment/self-report |
| How to enforce "refunds over $X need a human"? | **PreToolUse hook** (exit code 2), not a prompt instruction |
| How does `issue_refund` avoid double-refunds on retry? | **Idempotency key** |
| Where do backend integrations live? | **MCP tools**, 4–5 focused ones |

### Distractor anti-patterns

- Escalating because the customer "sounds frustrated" (#5) or the model "reports low confidence" (#4).
- Enforcing the refund limit via a system-prompt line (#3).
- Giving the agent 18 tools "for completeness" (#8).
- `issue_refund` returning empty/generic on failure (#6, #7).

---

## Scenario 2 · Code Generation with Claude Code

**Business context.** An engineering team standardises on Claude Code. They want shared conventions, safe permissions, enforced quality gates, and repeatable review workflows across the team.

```text
Repo/
 ├─ CLAUDE.md              (checked in: build/test cmds, architecture, conventions)
 ├─ .claude/
 │   ├─ settings.json      (checked in: permissions allow/deny/ask, hooks, model)
 │   ├─ agents/security-reviewer.md   (isolated context, tools allowlist, model)
 │   ├─ commands/review.md            (/review $ARGUMENTS)
 │   └─ hooks/guard.sh                (PreToolUse: block rm -rf; tests before commit)
 └─ .mcp.json             (project-scope shared MCP servers)
Managed policy (org) ──► overrides everything below
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| Where do team conventions live? | **Project `./CLAUDE.md`**, checked in |
| Where do mandatory, non-overridable rules live? | **Managed policy** |
| Enforce "tests must pass before commit"? | **PreToolUse hook** exiting 2 |
| Large, unfamiliar refactor — first step? | **Plan mode** (read-only, review plan first) |
| A repeatable review prompt invoked by name? | **Slash command** with `$ARGUMENTS` |
| Keep CLAUDE.md cheap? | **Concise**, `@import` reference material |

### Distractor anti-patterns

- Team rule placed in git-ignored `CLAUDE.local.md`.
- Quality gate enforced via CLAUDE.md prose (#3).
- Direct execution on a large unfamiliar refactor instead of plan mode.
- Secrets stored in CLAUDE.md.

---

## Scenario 3 · Multi-Agent Research System

**Business context.** A research tool answers broad questions by decomposing them into sub-questions, dispatching subagents to investigate in parallel, and synthesising a sourced answer. Robust **error handling** and **coordinator/subagent** design are the focus.

```text
Question ─► [Coordinator]  plans sub-questions, aggregates, decides done
                 │ explicit context passing (never inheritance)
     ┌───────────┼───────────────┐
[Subagent A]  [Subagent B]   [Subagent C]   isolated contexts, own tools
     └── structured result {finding, sources, confidence} ──┐
                 aggregate + partial-failure handling ◄──────┘
                 (quorum? retry retryable? escalate?)  provenance of gaps
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| Pattern? | **Orchestrator-workers** — subtasks decided at runtime |
| How do subagents get context? | **Explicit passing** in the task prompt; no auto-inheritance |
| A subagent fails — coordinator behaviour? | **Structured error → retry retryable → proceed with quorum noting gap, or escalate** |
| How to protect the main context? | **Subagent isolation** — raw material stays in subagent windows |
| How to trace a failed task? | **Per-agent traces + correlation ID** |
| Cost concern? | Model per subagent (Haiku for simple), justify fan-out |

### Distractor anti-patterns

- Assuming subagents inherit coordinator findings.
- Silently dropping a failed subagent and presenting the rest as complete (#7).
- Returning a generic "research failed" message (#6).
- Building 8 subagents where a 2-step workflow meets the bar (over-engineering).

---

## Scenario 4 · Developer Productivity

**Business context.** Developers use Claude to explore large codebases, run analyses, and integrate internal systems. The focus is **built-in/server-side tools** and **MCP servers** for codebase exploration.

```text
Developer ─► [Claude · Opus 5]
                 ├─ server-side tools: web search, code execution
                 ├─ MCP servers: internal docs, ticketing, code search
                 │     (project .mcp.json; least-privilege scopes)
                 └─ subagent for large codebase exploration (isolated context)
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| Ground answers in current info with citations? | **Web search** server-side tool |
| Run computation/data analysis? | **Code execution** server-side tool |
| Integrate internal systems reusably across clients? | **MCP servers** (project scope) |
| Explore a huge codebase without flooding context? | **Subagent** isolation |
| Deterministic step your code can do itself? | **Direct API/CLI**, not a model tool |
| How many tools per agent? | **4–5 focused**; tool search + `defer_loading` beyond ~10 |

### Distractor anti-patterns

- Wrapping a deterministic call as a model tool.
- Overloading the agent with tools (#8).
- Choosing a Skill where a cross-client integration needs an MCP server.

---

## Scenario 5 · Claude Code for CI/CD

**Business context.** A pipeline runs Claude Code **headlessly** to review PRs, generate release notes, and gate merges. Emphasis on **structured output**, the **Batch API**, and **multi-pass review** of large PRs.

```text
CI trigger ─► claude -p "review diff" --output-format json
                 --allowedTools "Read,Grep,Bash(git diff:*)" --permission-mode acceptEdits
                 │  parse JSON findings; gate pipeline on exit code
Large PR ─► partition by module ─► review each pass/subagent ─► aggregate + rank
Bulk jobs (release notes for 500 PRs) ─► Message Batches API (50% off, ≤24h)
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| How to run in CI? | **`claude -p` headless**, `--output-format json`, minimal `--allowedTools`, gate on **exit code** |
| Machine-readable findings? | **Structured output** (JSON schema / structured outputs) |
| Large PR review? | **Multi-pass**: partition → review → aggregate |
| Bulk latency-tolerant jobs? | **Batch API** (50% discount, within 24h) |
| Least privilege in CI? | Narrow `--allowedTools` + deny dangerous ops; not `bypassPermissions` |

### Distractor anti-patterns

- `--permission-mode bypassPermissions` with all tools in CI.
- Reviewing a 4,000-line PR in one context pass.
- Grepping prose instead of parsing JSON output.
- Aggregate accuracy across PR types masking a weak type (#10).

---

## Scenario 6 · Structured Data Extraction

**Business context.** A pipeline extracts structured records from heterogeneous documents (invoices, contracts, receipts) and feeds a database. Emphasis on **JSON schemas**, **tool_use**, and **validation-retry**.

```text
Document ─► classify type ─┬─ invoice schema  ┐
                           ├─ contract schema ├─► extract (structured outputs /
                           └─ receipt schema  ┘   strict tool) ─► validate
                                                     │ fail: feed error back, retry
                                                     ▼ ok
                                              per-type metrics ─► DB (with provenance)
```

### Key decisions and the exam-correct answer

| Decision | Exam-correct answer |
| --- | --- |
| Guarantee schema-conformant output? | **Structured outputs** (`output_config.format`) or **strict tools** |
| On Fable 5.1, force a tool? | **No** — `auto` + instruction / `strict` / structured outputs |
| Validation fails? | **Feed the specific error back** and retry; parse defensively |
| Measure quality? | **Per document type**, gate on the worst type (not aggregate — #10) |
| `max_tokens` on a big doc? | Truncation — raise limit or chunk; not a completion |
| Optional field absent? | Model **nullable** in the schema, not a hallucinated value |

### Distractor anti-patterns

- Forcing `tool_choice` on Fable 5.1 (400).
- One aggregate accuracy number across types (#10).
- Treating truncated `max_tokens` output as complete.
- Grading extraction quality in the same session that produced it (#9).

---

## Cross-scenario failure-mode reference

Each scenario has a signature way it breaks in production. Memorising the failure → root cause → fix chain is how you eliminate distractors fast.

| Scenario | Common failure | Root cause | Exam-correct fix |
| --- | --- | --- | --- |
| S1 Support agent | Loop never stops / stops early | Prose parsing for termination (#1) | Drive from `stop_reason` |
| S1 Support agent | Double refund on retry | Non-idempotent write retried | Idempotency key |
| S1 Support agent | Hijacked by fetched content | Tool result treated as instructions | Content boundaries + least privilege + human gate |
| S2 Code gen | Rule not enforced | Prompt/CLAUDE.md enforcement (#3) | PreToolUse hook, exit 2 |
| S2 Code gen | Wrong file holds a setting | Shared/personal or advisory/mandatory confusion | Managed policy / project / local mapping |
| S3 Research | Subagent ignores known findings | Isolated context, no explicit passing | Pass context explicitly at every tier |
| S3 Research | Partial report shipped as complete | Silent suppression (#7) | Structured error + quorum/escalate |
| S3 Research | 12× cost, marginal gain | Over-engineered topology | Simplest design that meets the bar + caching |
| S4 Dev productivity | Wrong tool called | Too many tools (#8) | 4–5 focused / tool search + defer_loading |
| S4 Dev productivity | Deterministic step is flaky | Wrapped as a model tool | Call the API/CLI directly |
| S5 CI/CD | Build not gated | Grepping prose / ignoring exit code | JSON output + exit-code gating |
| S5 CI/CD | Arbitrary shell from a diff | Broad permissions in CI | Minimal `--allowedTools`, no bypass |
| S6 Extraction | 400 on Fable 5.1 | Forced `tool_choice` | auto+instruction / strict / structured outputs |
| S6 Extraction | Ships bad contracts at "94%" | Aggregate metric (#10) | Per-type metrics, gate on worst |
| S6 Extraction | Partial JSON persisted | `max_tokens` treated as complete | Raise limit / chunk, never persist partial |

---

## Practice questions

Five per scenario — 30 in total. Each is tagged with its scenario and domain.

<Accordions>
  <AccordionItem title="Q1 · [S1/D1] A support agent should hand off to a human. The customer writes an angry message but the request (order status) is fully resolvable. What is correct? (Select one)">
    A. Escalate immediately because of the negative sentiment.
    B. Resolve the order-status request; sentiment alone is not an escalation trigger.
    C. Escalate because the model's confidence is below 70%.
    D. Cap the conversation at 3 turns then escalate.

    **Answer: B.** Sentiment is not a trigger (#5); the request is within capability, so resolve it. Confidence-based escalation is #4, and an arbitrary cap is not an escalation rule.
  </AccordionItem>

  <AccordionItem title="Q2 · [S1/D1] The rule 'refunds over $500 need human approval' must always hold. How is it enforced? (Select one)">
    A. A sentence in the agent's system prompt.
    B. A PreToolUse hook on `issue_refund` that inspects the amount and exits 2 to block, routing to a human.
    C. Asking the model to double-check itself.
    D. A note in CLAUDE.md.

    **Answer: B.** Critical rules need deterministic hooks (#3). Prompt/CLAUDE.md text (A, D) and self-check (C) are probabilistic.
  </AccordionItem>

  <AccordionItem title="Q3 · [S1/D4] The refund tool is retried after a 429 and a customer is refunded twice. What prevents this? (Select one)">
    A. Longer backoff.
    B. An idempotency key on `issue_refund` so retries do not duplicate the effect.
    C. Never retrying anything.
    D. A bigger model.

    **Answer: B.** Idempotency keys make write retries safe. Backoff (A) does not prevent duplication, never retrying (C) is unnecessary, and model size (D) is irrelevant.
  </AccordionItem>

  <AccordionItem title="Q4 · [S2/D2] A quality rule 'tests must pass before commit' must be unbypassable for the whole team. What is correct? (Select one)">
    A. Add it to project CLAUDE.md.
    B. A checked-in PreToolUse hook that runs tests on commit and exits 2 on failure.
    C. Ask each developer to remember.
    D. A slash command that runs tests.

    **Answer: B.** A checked-in hook enforces it deterministically for everyone. CLAUDE.md (A) and memory (C) are probabilistic; a slash command (D) is opt-in.
  </AccordionItem>

  <AccordionItem title="Q5 · [S2/D2] Which belongs in a git-ignored `CLAUDE.local.md` rather than project CLAUDE.md? (Select one)">
    A. The team's build and test commands.
    B. A developer's personal scratch notes and local machine paths.
    C. The repo's architecture conventions.
    D. A mandatory security rule.

    **Answer: B.** Personal, non-shared notes go in the git-ignored local file. Team commands/conventions (A, C) go in project CLAUDE.md; mandatory rules (D) go in managed policy.
  </AccordionItem>

  <AccordionItem title="Q6 · [S2/D2] An engineer must refactor an unfamiliar 12-file module. What is the BEST first step? (Select one)">
    A. Direct execution to move fast.
    B. Plan mode: read-only exploration then a reviewable plan before edits.
    C. Delete failing tests.
    D. Increase max_tokens.

    **Answer: B.** Large, unfamiliar, multi-file work is the canonical plan-mode case. Direct execution (A) skips review, deleting tests (C) is destructive, and max_tokens (D) is irrelevant.
  </AccordionItem>

  <AccordionItem title="Q7 · [S3/D1] A coordinator delegates to a subagent, which produces an answer that ignores prior findings. Why? (Select one)">
    A. The subagent needs a bigger window.
    B. Subagent contexts are isolated and do not inherit findings; the coordinator must pass context explicitly.
    C. The subagent used the wrong model.
    D. Thinking was disabled.

    **Answer: B.** Isolated contexts never auto-inherit. Window size (A), model (C) and thinking (D) do not supply context that was never passed.
  </AccordionItem>

  <AccordionItem title="Q8 · [S3/D1] Two of five research subagents time out. What should the coordinator do? (Select one)">
    A. Present the three results as the complete answer.
    B. Record structured errors (timeout, retryable), retry the retryables, then proceed with a quorum noting the gap or escalate.
    C. Discard everything and restart.
    D. Return 'research failed'.

    **Answer: B.** Structured errors plus an explicit decision. Silent drop (A) is #7, restart (C) wastes good work, and a generic message (D) is #6.
  </AccordionItem>

  <AccordionItem title="Q9 · [S3/D5] The main context fills because subagents return their entire raw source material. What is the BEST fix? (Select one)">
    A. Move to Haiku 4.5 for a bigger window.
    B. Have subagents return only distilled, structured results; their isolated contexts hold the raw material.
    C. Turn off thinking.
    D. Remove error handling.

    **Answer: B.** Subagent isolation plus distilled returns keep the main window small. Haiku 4.5 (A) has a smaller (200k) window, and C/D are unrelated or harmful.
  </AccordionItem>

  <AccordionItem title="Q10 · [S4/D4] A developer needs answers grounded in current external information with citations. Which tool? (Select one)">
    A. Code execution.
    B. Web search server-side tool.
    C. Memory tool.
    D. Computer use.

    **Answer: B.** Web search grounds with citations. Code execution (A) runs code, memory (C) persists state, computer use (D) drives a desktop.
  </AccordionItem>

  <AccordionItem title="Q11 · [S4/D4] An internal system must be reachable from Claude Code, Desktop and the Messages API. What should be built? (Select one)">
    A. Three separate custom tools.
    B. An MCP server exposing the capability, connected by each host and via the Messages API MCP connector.
    C. A slash command.
    D. A CLAUDE.md note.

    **Answer: B.** MCP provides one reusable cross-client integration. Three tools (A) duplicate work; a slash command (C) and CLAUDE.md (D) are Claude Code artefacts, not integrations.
  </AccordionItem>

  <AccordionItem title="Q12 · [S4/D4] A step is fully deterministic and your code can call it directly. Should it be a model tool? (Select one)">
    A. Yes, for consistency.
    B. No — call the API/CLI directly; wrapping it as a model tool needlessly adds latency, cost and non-determinism.
    C. Yes, expose it as an MCP resource.
    D. Only if it is on Fable 5.1.

    **Answer: B.** Deterministic steps your code owns should be called directly. Wrapping (A, C) is over-engineering; the model version (D) is irrelevant.
  </AccordionItem>

  <AccordionItem title="Q13 · [S5/D2] A CI job must review PRs, emit machine-readable findings, and fail the build on issues. Which is correct? (Select one)">
    A. `claude -p "…" --output-format json --allowedTools "Read,Grep,Bash(git diff:*)"`, parse JSON, gate on exit code.
    B. Interactive `claude`, copy results by hand.
    C. `--permission-mode bypassPermissions` with all tools.
    D. Plain `claude -p` and grep the prose.

    **Answer: A.** Headless JSON output, minimal allowlist, exit-code gating. Interactive (B) does not automate, bypassing permissions (C) is unsafe, and grepping prose (D) is unreliable.
  </AccordionItem>

  <AccordionItem title="Q14 · [S5/D3] Release notes must be generated for 500 merged PRs overnight at minimum cost. Which choice? (Select one)">
    A. Real-time Messages API at high concurrency.
    B. The Message Batches API — 50% discount, results within 24h, latency-tolerant.
    C. One giant request with all 500 diffs.
    D. Force tool_choice on Fable 5.1.

    **Answer: B.** Batch API fits latency-tolerant bulk work at half price. High concurrency (A) risks limits and costs more, one request (C) will not fit, and forcing tool_choice on Fable 5.1 (D) 400s.
  </AccordionItem>

  <AccordionItem title="Q15 · [S5/D2] A 4,000-line PR does not fit one review pass. What is the BEST approach? (Select one)">
    A. Truncate to 500 lines.
    B. Multi-pass: partition by module, review each in its own pass/subagent, then aggregate and rank findings.
    C. One giant prompt with the whole diff.
    D. Skip review.

    **Answer: B.** Partition-review-aggregate preserves quality. Truncation (A) misses code, one giant prompt (C) degrades quality, and skipping (D) is unacceptable.
  </AccordionItem>

  <AccordionItem title="Q16 · [S6/D3] On Fable 5.1, an extraction sets a forced tool_choice of type tool and gets 400s. What is correct? (Select one)">
    A. Retry with backoff.
    B. Use `tool_choice: 'auto'` with an instruction to call the tool, `strict: true` schemas, or structured outputs.
    C. Use `tool_choice: 'any'`.
    D. Lower max_tokens.

    **Answer: B.** Fable 5.1 forbids forced tool choice; use auto+instruction, strict, or structured outputs. Backoff (A) does not fix a 400, `any` (C) is also blocked, and max_tokens (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q17 · [S6/D3] Overall extraction accuracy is 94% but contracts are frequently wrong. What is the correct evaluation change? (Select one)">
    A. Increase the sample size.
    B. Report per-document-type accuracy and gate on the worst type.
    C. Raise temperature for contracts.
    D. Average more runs.

    **Answer: B.** Aggregate metrics hide a failing type (#10); per-type metrics expose it. Larger samples (A) still aggregate, temperature (C) does not fix accuracy, and averaging (D) hides the problem more.
  </AccordionItem>

  <AccordionItem title="Q18 · [S6/D3] A validation-retry loop just re-sends the same prompt and keeps failing. What most helps? (Select one)">
    A. More blind retries.
    B. Feed the specific validation error back to the model and require valid JSON matching the schema; parse defensively.
    C. Use `eval()` to parse.
    D. Grade the output in the same session that produced it.

    **Answer: B.** Specific error feedback drives self-correction. Blind retries (A) rarely help, `eval()` (C) is unsafe, and same-session grading (D) is #9.
  </AccordionItem>

  <AccordionItem title="Q19 · [S1/D5] A support agent calls a billing tool that occasionally hangs, stalling the whole turn, and once a hung call was reported as 'done'. What is the BEST design? (Select one)">
    A. Remove the timeout so slow calls eventually return.
    B. Set a boundary timeout; on timeout return a structured `{category:'timeout', retryable:true}` error so the agent can retry, proceed noting the gap, or escalate.
    C. Catch the timeout and return an empty success.
    D. Use a larger model so the tool responds faster.

    **Answer: B.** Boundary timeouts plus a structured error keep the agent responsive and honest. No timeout (A) lets a hung tool stall the agent, empty success (C) is #7, and model size (D) does not change a downstream tool's latency.
  </AccordionItem>

  <AccordionItem title="Q20 · [S1/D1] A payment agent must always require human sign-off above $10,000, even if a fetched document claims pre-approval. Which control is correct? (Select one)">
    A. A confidence threshold on the model.
    B. A deterministic PreToolUse hook on the transfer tool that inspects the amount and exits 2 above the threshold, ignoring any injected 'pre-approval'.
    C. Escalate only if the customer seems anxious.
    D. A system-prompt instruction to ask before large transfers.

    **Answer: B.** Irreversible high-value actions need a deterministic gate keyed on the amount, immune to injection. Confidence (A) is #4, sentiment (C) is #5, and a prompt rule (D) is #3.
  </AccordionItem>

  <AccordionItem title="Q21 · [S2/D2] A team needs a linter to run after each edit AND commits blocked when tests fail. Which hook events are correct? (Select one)">
    A. PostToolUse to block the commit; PreToolUse to lint.
    B. PreToolUse on the commit (run tests, exit 2 on failure) and PostToolUse on Edit to lint.
    C. UserPromptSubmit for both.
    D. SessionStart to lint and Stop to run tests.

    **Answer: B.** Blocking happens before the action (PreToolUse, exit 2); linting reacts after the edit (PostToolUse). A swaps the events, C uses a prompt event, and D fires at the wrong times.
  </AccordionItem>

  <AccordionItem title="Q22 · [S2/D2] Managed policy sets Sonnet 5, project settings set Opus 5, and a user file sets Haiku 4.5. Which model wins? (Select one)">
    A. Haiku 4.5 (user is most personal).
    B. Sonnet 5 — managed policy cannot be overridden.
    C. Opus 5 (project is closest to code).
    D. Whichever file changed last.

    **Answer: B.** Managed policy overrides local, project and user. A and C invert precedence; D is not how resolution works.
  </AccordionItem>

  <AccordionItem title="Q23 · [S3/D5] A research agent degrades after many turns as the window fills with large, stale search results, but the dialogue must stay intact. Which mechanism, and which would be wrong? (Select one)">
    A. Compaction, because it summarises everything.
    B. Context editing to clear the stale tool results while preserving the narrative; compaction would be wrong here because it summarises the dialogue rather than targeting bulky tool outputs.
    C. Switch to Haiku 4.5 for its larger window.
    D. Increase max_tokens.

    **Answer: B.** Clearing bulky stale tool results is context editing, which keeps the dialogue. Compaction (A) targets the narrative, Haiku 4.5 (C) has a smaller 200k window, and max_tokens (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q24 · [S3/D1] A hierarchical system (root → sub-coordinators → workers) loses root-level findings three tiers down. Why? (Select one)">
    A. Workers need bigger windows.
    B. Context is isolated at every tier; each level must pass the relevant context explicitly to the next — inheritance never happens at any depth.
    C. The root should use Opus 5.
    D. Thinking is disabled at the worker tier.

    **Answer: B.** Isolation applies at every tier, so context must be threaded explicitly all the way down. Window size (A) and model (C) do not supply un-passed context; thinking (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q25 · [S4/D4] An MCP server must expose the refund-policy document (read-only reference the app supplies) and a `search_orders` action. Which primitives? (Select one)">
    A. Both as Tools.
    B. The policy as a Resource (application-controlled data) and `search_orders` as a Tool (model-controlled action).
    C. Both as Prompts.
    D. The policy as a Tool and `search_orders` as a Resource.

    **Answer: B.** App-supplied reference data is a Resource; a model-invoked action is a Tool. Making the policy a Tool (A, D) adds needless model decisions; Prompts (C) are user-invoked templates.
  </AccordionItem>

  <AccordionItem title="Q26 · [S4/D4] A shared remote MCP server is used by many teams; a support client only reads orders. How should its access be set? (Select one)">
    A. Full OAuth scopes so it never lacks a capability.
    B. Scope its OAuth 2.1 grant to read-only order access (least privilege applied to auth).
    C. Embed an admin API key in the prompt.
    D. Use stdio so no auth is needed.

    **Answer: B.** Least privilege applies to MCP auth. Full scopes (A) widen blast radius, prompt-embedded keys (C) are insecure, and stdio (D) is a local transport, not an option for a shared remote server.
  </AccordionItem>

  <AccordionItem title="Q27 · [S5/D2] A CI reviewer must emit machine-readable findings for a short task and fail the build on high-severity issues. Which is correct? (Select one)">
    A. `--output-format text`, then grep for 'HIGH'.
    B. `--output-format json`, parse findings, and fail on a non-zero exit code and any high-severity finding.
    C. `--output-format stream-json` and ignore the exit code.
    D. Interactive mode with a human reading output.

    **Answer: B.** A single JSON result plus exit-code-and-findings gating is correct for a short task. Grepping prose (A) is unreliable, ignoring the exit code (C) misses failures, and interactive mode (D) does not automate.
  </AccordionItem>

  <AccordionItem title="Q28 · [S5/D3] A million documents must be extracted overnight as cheaply as possible with a fixed schema. Which combination is MOST cost-effective? (Select two)">
    A. Cache the stable system + schema prefix so reads cost ~0.1x.
    B. Use the Message Batches API for the bulk run (50% off, ≤24h).
    C. Call the real-time API at maximum concurrency.
    D. Use Fable 5.1 at $10/$50 for every document.
    E. Randomise the prompt each call.

    **Answer: A and B.** Caching the fixed prefix plus the Batch API cut cost sharply for latency-tolerant bulk work. Real-time concurrency (C) is full price and limit-prone, the priciest model (D) raises cost, and randomising (E) destroys the cacheable prefix.
  </AccordionItem>

  <AccordionItem title="Q29 · [S6/D3] One pipeline handles invoices, contracts and receipts, each with different required fields. Which schema pattern is BEST? (Select one)">
    A. One loose object with everything optional.
    B. A discriminated union (`oneOf` with a `const` `doc_type` discriminator) so each branch enforces its own required fields.
    C. Three unrelated endpoints with no shared contract.
    D. A single string field holding raw JSON text.

    **Answer: B.** A discriminated union enforces per-type requirements in one schema. A loosens everything, C loses a shared contract, and D abandons schema guarantees.
  </AccordionItem>

  <AccordionItem title="Q30 · [S6/D3] A big contract's extracted JSON ends mid-array with `stop_reason` `max_tokens` and the pipeline writes the partial object. What is correct? (Select one)">
    A. Write the partial object; it is mostly complete.
    B. Treat `max_tokens` as truncation: raise the limit or chunk the document (or bound arrays with `maxItems`), then retry — never persist the partial output.
    C. Return an empty object so the pipeline continues.
    D. Ask the model in the same session whether it finished.

    **Answer: B.** `max_tokens` is truncation, not completion. Persisting the partial (A) corrupts data, empty (C) is #7, and same-session self-check (D) does not address truncation.
  </AccordionItem>
</Accordions>
