# D1 · Solution Design & Architecture

Translating business problems into Claude solutions, end-to-end reference architecture, choosing between workflow / agentic / augmented-LLM patterns, cloud placement, capacity planning, HA and fallback design, and ADRs.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This domain is worth **17% – roughly 11 of 63 items** – and it sets the frame for the whole exam. It tests whether you can take an ambiguous business problem, extract the constraints and success criteria, and produce a defensible end-to-end architecture. Items here rarely have a single "feature" answer; they reward the option that **starts from business value and reasons under constraint**. The recurring failure mode is over-engineering (reaching for multi-agent orchestration when a linear workflow meets the SLA at lower cost and higher reliability).

## Learning objectives

By the end of this page you should be able to:

1. Run **discovery**: extract the real problem, success criteria, and hard constraints from a stakeholder brief.
2. Draw an **end-to-end reference architecture** (input → processing → output → feedback loop) and name each component's responsibility.
3. Choose among **workflow, agentic and augmented-LLM** patterns using signal-based decision rules.
4. Decompose work into **multi-agent orchestration** only when the signals justify it.
5. Align a design to the **business value pillars** (efficiency, cost, performance SLAs) and make a **build-vs-buy** call.
6. Choose **cloud placement** (direct API vs Bedrock vs Vertex vs Foundry) on data-residency, IAM and commitment grounds.
7. Do **capacity planning** against rate-limit tiers and design **HA and fallback**.
8. Record decisions as **ADRs**.

---

## 1.1 Discovery: from business problem to solvable spec

Architects are handed *symptoms*, not specs. Discovery converts "we want to use AI for support" into a bounded problem with measurable success criteria and enumerated constraints. If you skip this, every downstream decision is unanchored.

### The discovery question set

| Category | Questions to ask | Why it changes the design |
| --- | --- | --- |
| Problem & value | What decision or task are we automating? What is the value pillar – efficiency, cost, or an SLA? | Determines whether latency, accuracy or unit cost dominates |
| Success criteria | What does "good enough" look like numerically (accuracy, p95 latency, deflection rate, cost per task)? | Becomes the eval target and the go/no-go gate |
| Volume & shape | Requests per second, peak vs average, sync vs async, payload sizes | Drives capacity planning and Batch vs real-time |
| Data | Sensitivity class, residency, retention obligations, sources, freshness | Drives cloud placement, RAG design, compliance |
| Constraints | Existing cloud commitments, IAM, budget ceiling, deadline, team skills | Narrows build-vs-buy and placement |
| Risk & reversibility | What is the blast radius of a wrong answer? Which actions are irreversible? | Determines human gates and guardrail depth |
| Integration | What systems must it read/write? What identity model? | Drives tool design, MCP vs API, authn/authz |

:::tip[Exam signal]
When a stem gives you a vague goal plus one hard number (a latency SLA, a monthly budget, a residency requirement), the correct answer is the one that **treats that number as the binding constraint** and rejects options that ignore it. "Constraint-blind" is the most common wrong-answer pattern in this domain.
:::

### Success criteria must be measurable and per-segment

"Improve support quality" is not a success criterion. "Resolve ≥ 60% of Tier-1 billing tickets without escalation, p95 latency under 6 s, cost under \$0.05 per resolved ticket, with no regression on refund-related tickets" is. Note the **per-segment** clause – aggregate targets hide the failures the exam wants you to catch.

---

## 1.2 The end-to-end reference architecture

Every Claude solution decomposes into four stages plus cross-cutting concerns. Learn to place any component into this skeleton.

```text
        INPUT                PROCESSING                 OUTPUT              FEEDBACK LOOP
  ┌───────────────┐   ┌───────────────────────┐   ┌───────────────┐   ┌───────────────────┐
  │ user / event  │   │  orchestration layer  │   │ validation /  │   │ evals (offline)   │
  │ API / webhook │──▶│  ┌─────────────────┐  │──▶│ structured    │──▶│ online metrics    │
  │ queue / batch │   │  │ Claude model(s) │  │   │ output check  │   │ traces + cost     │
  │ document      │   │  │ + tools (MCP)   │  │   │ human gate    │   │ user thumbs / QA  │
  └───────────────┘   │  │ + retrieval/RAG │  │   └───────┬───────┘   └─────────┬─────────┘
                      │  └─────────────────┘  │           │                     │
                      └───────────┬───────────┘           ▼                     │
                                  │               downstream systems            │
   cross-cutting:  auth · secrets · guardrails · observability · caching  ◀─────┘ (retrain prompts,
                                                                                   reroute models,
                                                                                   tune retrieval)
```

| Stage | Responsibility | Key decisions |
| --- | --- | --- |
| Input | Normalise and admit work | Sync API vs webhook vs queue vs Batch API; payload validation |
| Processing | Reason and act | Model choice/routing, tools, retrieval, orchestration pattern |
| Output | Guarantee shape and safety | Structured outputs, validation-retry, human gate on irreversible actions |
| Feedback | Improve the system | Offline evals, online metrics, traces, cost telemetry feeding back into prompts/models/retrieval |

The **feedback loop is not optional** at Professional level. A design with no path from production signals back into prompts, routing and retrieval is incomplete – expect distractors that omit it.

---

## 1.3 Choosing the pattern: workflow vs agentic vs augmented-LLM

This is the single most tested decision in D1. The default should be the **simplest pattern that meets the criteria**. Escalate to agentic only when the task genuinely requires dynamic, model-driven control flow.

| Pattern | What it is | Choose when the stem shows… | Avoid when… |
| --- | --- | --- | --- |
| **Augmented LLM** | A single call with retrieval + tools + structured output | Task is one bounded step; deterministic inputs; tight latency/cost SLA | The task needs multiple dependent steps with branching |
| **Workflow (orchestrated)** | Predefined graph of steps; code owns control flow; models fill steps | Steps are known and stable; you can enumerate the DAG; you need reproducibility and easy debugging | The path genuinely can't be known ahead of time |
| **Agentic** | Model decides next action in a loop until `stop_reason` | Open-ended goal; number/order of steps unknown; tool use is exploratory | A fixed workflow meets the SLA more cheaply and reliably (most cases) |

```text
Is the sequence of steps knowable in advance?
  ├─ Yes → Can it be done in one call?
  │        ├─ Yes → AUGMENTED LLM (retrieval + tools + structured output)
  │        └─ No  → WORKFLOW (orchestrated DAG; code owns control flow)
  └─ No  → Does the task truly need dynamic decisions AND is the cost/latency/reliability
           hit justified by the value?
           ├─ Yes → AGENTIC (loop on stop_reason; bounded tools; guardrails)
           └─ No  → Re-decompose into a WORKFLOW
```

:::caution[Control-flow anti-patterns (exam favourites)]
An agentic loop must terminate on **`stop_reason`** (`end_turn` / `tool_use`), never by **parsing natural language** for "done" and never with an **arbitrary iteration cap as the primary stop**. A cap is a safety backstop, not the control mechanism. Options that use string matching or a hard cap alone are wrong.
:::

---

## 1.4 Multi-agent orchestration and decomposition

Multi-agent (coordinator + subagents) is powerful and expensive. Each subagent multiplies token cost and latency and adds failure surface. Decompose into multiple agents only when the signals are present.

| Signal for multi-agent | Signal against (keep it single/workflow) |
| --- | --- |
| Genuinely parallel subtasks with independent context | Steps are sequential and share context |
| Distinct skill/tool sets that would bloat one agent's tool list | A single tool set under ~5–7 tools covers it |
| Need for context isolation (a subagent explores without polluting the coordinator) | Latency/cost budget is tight |
| Long-horizon research/synthesis where fan-out/fan-in helps | Reproducibility and easy debugging are priorities |

```text
        ┌──────────────┐
        │ Coordinator  │  owns plan, aggregates, applies guardrails
        └──────┬───────┘
     ┌─────────┼─────────┐
     ▼         ▼         ▼
 ┌───────┐ ┌───────┐ ┌───────┐   subagents: isolated context windows,
 │ sub A │ │ sub B │ │ sub C │   narrow tool allowlists, own model tier
 └───────┘ └───────┘ └───────┘   (e.g. Haiku for retrieval, Opus for synthesis)
```

Assign **cheaper models to narrow subagents** (Haiku 4.5 for extraction/routing) and reserve **Opus 5** for the coordinator or the hardest synthesis step. This is a portfolio decision, revisited in D2.

---

## 1.5 Aligning to business value pillars

Every design serves one dominant pillar; naming it resolves most trade-offs.

| Pillar | Dominant metric | Design levers | Typical model posture |
| --- | --- | --- | --- |
| Efficiency (throughput / deflection) | tasks automated, deflection rate | workflow simplification, batching, caching | Sonnet 5 default, Haiku for high-volume |
| Cost (unit economics) | cost per task | routing/cascades, prompt caching, Batch API, output trimming | Haiku 4.5 first, escalate only on failure |
| Performance SLA (latency) | p50/p95 latency | fast mode, smaller model, fewer tool round-trips, streaming | Haiku 4.5 / Sonnet 5, fast mode |

When two pillars conflict (cheap vs fast, or accurate vs cheap), the **stated success criterion breaks the tie**. If the stem never states one, the correct answer is usually to **go define it with the stakeholder**, not to guess.

---

## 1.6 Build vs buy

| Factor | Lean build | Lean buy |
| --- | --- | --- |
| Differentiation | Core to competitive advantage | Commodity capability |
| Team capability | You have ML/infra skills to operate it | You lack ops capacity |
| Time to value | You have runway | You need it now |
| Total cost of ownership | Volume amortises build cost | Low/uncertain volume |
| Compliance control | You need full control of data path | Vendor's certifications suffice |

For Claude specifically, "buy" often means using **managed agents** (Anthropic hosts the loop and sandbox) or a higher-level product, while "build" means the **Agent SDK / Tool Runner** where you host the loop. Choose managed when you want speed and less ops; choose self-hosted when you need control over the execution environment, data path or custom tooling.

---

## 1.7 Cloud placement: direct API vs Bedrock vs Vertex vs Foundry

Claude is available directly and via **Amazon Bedrock**, **Google Vertex AI**, and **Microsoft Foundry**. This is a compliance-and-commitment decision far more than a capability one.

| Dimension | Direct (Anthropic API) | Amazon Bedrock | Google Vertex AI | Microsoft Foundry |
| --- | --- | --- | --- | --- |
| Data residency | Anthropic regions | AWS regions incl. FedRAMP High | GCP regions incl. FedRAMP High | Azure regions |
| IAM | Anthropic API keys | AWS IAM / SigV4 / roles | GCP IAM / service accounts | Entra ID / Azure RBAC |
| Existing commitment | none | AWS spend commit / EDP | GCP commit | Azure commit / MACC |
| Compliance | Anthropic certs | Inherit AWS + FedRAMP High | Inherit GCP + FedRAMP High | Inherit Azure |
| Latest features first | Usually earliest | Slight lag | Slight lag | Slight lag |

:::tip[Exam signal]
"Data must stay in-region", "FedRAMP High", "we already have an AWS EDP", "identity must flow through corporate SSO" → choose the **cloud platform whose IAM and residency you already own** (Bedrock/Vertex/Foundry). "We want the newest model the day it ships" and no residency constraint → **direct API**. Don't pick direct API when a residency or existing-commitment signal is present.
:::

---

## 1.8 Capacity planning and rate-limit tiers

Rate limits are enforced per model as **RPM** (requests/min), **ITPM** (input tokens/min) and **OTPM** (output tokens/min), scaled by usage tier. Capacity planning means proving your peak load fits the tier – or designing around it.

<Steps>

1. Estimate peak: `peak_RPM = peak_requests_per_sec × 60`; `peak_ITPM = peak_RPM × avg_input_tokens`; likewise OTPM.

2. Compare against the model's tier limits (`client.models.retrieve(id)` and account tier). If you exceed any of the three, you are limited by that one.

3. Design around limits: shift latency-tolerant work to the **Message Batches API** (50% discount, results within 24 h), spread load, request a tier increase, or route overflow to a second model.

4. Handle 429s with **exponential backoff + jitter**, honouring the `retry-after` header. A 429 is expected under burst, not an error to swallow.

</Steps>

```python
import time, random, anthropic

client = anthropic.Anthropic()

def call_with_backoff(**kwargs):
    for attempt in range(6):
        try:
            return client.messages.create(**kwargs)
        except anthropic.RateLimitError as e:
            retry_after = float(getattr(e, "retry_after", 0) or 0)
            sleep = retry_after or min(2 ** attempt + random.random(), 30)
            time.sleep(sleep)
    raise RuntimeError("exhausted retries")
```

---

## 1.9 High availability and fallback design

Production Claude systems must degrade gracefully, not fail hard. Retry `429/500/529` with backoff; for sustained unavailability, **fall back to another model or provider**.

```text
             ┌─────────────────────────┐
 request ──▶ │ primary: claude-opus-5  │──✓──▶ response
             └───────────┬─────────────┘
                         │ 529 overloaded / timeout (after backoff)
                         ▼
             ┌─────────────────────────┐
             │ fallback: claude-sonnet-5│──✓──▶ response (log degraded mode)
             └───────────┬─────────────┘
                         │ still failing
                         ▼
             ┌─────────────────────────┐
             │ cross-provider: Bedrock │──✓──▶ response
             │ or queued for retry     │
             └─────────────────────────┘
```

| Failure | Mitigation |
| --- | --- |
| 429 rate limit | backoff + jitter, honour `retry-after`, spillover routing |
| 529 overloaded / 5xx | retry with backoff; fall back to secondary model/provider |
| Regional outage | multi-region via Bedrock/Vertex; queue and replay |
| Bad output shape | structured outputs + validation-retry (D4) |
| Irreversible action | human approval gate before execution |

:::caution[Fallback that silently drops capability]
Fable 5.1's **thinking blocks are readable only by the producing model or newer**; a silent fallback to an older model **drops those blocks** and can corrupt a multi-turn agentic session. Fallback design must account for feature parity, not just availability. Never mask a degraded path as a normal success.
:::

---

## 1.10 Architecture Decision Records (ADRs)

An ADR captures **one decision, its context, the options considered, the choice, and the consequences**. On the exam, the "best" answer often mirrors ADR discipline: it states the constraint, names the rejected alternative and why, and accepts an explicit trade-off.

```markdown
# ADR-014: Retrieval layer for the policy-QA assistant
## Status: Accepted
## Context
Regulated (GDPR); 2M docs; answers must cite source clause; p95 < 8 s; budget $0.04/query.
## Decision
Hybrid retrieval (BM25 + dense) with reranking; Claude Sonnet 5 for synthesis on Vertex AI (EU residency).
## Alternatives considered
- Long-context stuffing (1M): rejected — cost per query > budget, no citation granularity.
- Fine-tuning: rejected — corpus changes weekly; retraining cadence infeasible.
- Direct API: rejected — EU data residency requires Vertex EU region.
## Consequences
+ Citations at clause level; cost within budget.
- Added reranker latency (~300 ms) and an index-refresh pipeline to operate.
```

---

## 1.11 Capacity and cost modelling with arithmetic

Capacity planning at Professional level is a numbers exercise, not a vibe. Prove the load fits the tier and the budget *before* the design review.

### Worked capacity model

A support workload: peak **20 requests/sec**, average **8 req/s**; each call ~**6k input** + **1.5k output** tokens on Sonnet 5.

```text
peak_RPM  = 20 req/s × 60            = 1,200 RPM
peak_ITPM = 1,200 × 6,000            = 7,200,000 input tokens/min
peak_OTPM = 1,200 × 1,500            = 1,800,000 output tokens/min
```

You are limited by whichever of RPM / ITPM / OTPM you breach first. If the tier caps ITPM at 4,000,000, ITPM is the binding limit at ~55% of peak — so ~45% of peak requests must shift to Batch, spill to a second model, or you request a tier increase.

### Worked cost model (per month)

Average 8 req/s → `8 × 86,400 = 691,200 req/day ≈ 20.7M req/month`.

| Design | Per-request cost | Monthly (20.7M req) |
| --- | --- | --- |
| All Sonnet 5 (6k in @ \$2, 1.5k out @ \$10) | 6k×\$2/1e6 + 1.5k×\$10/1e6 = \$0.012 + \$0.015 = **\$0.027** | **\$559k** |
| + prompt cache (80% hit on a 4k stable prefix, read 0.1×) | saves ~0.8×(4k×\$2×0.9)/1e6 = ~\$0.0058 → **\$0.021** | **\$435k** |
| Cascade: 70% Haiku (\$0.0135), 30% Sonnet (\$0.027) | 0.7×\$0.0135 + 0.3×\$0.027 = **\$0.0176** | **\$364k** |
| Cascade + cache + 30% batchable at 50% off | ≈ **\$0.013** | **\$269k** |

:::tip[Exam signal]
When a stem gives request rate, token sizes and a budget, the correct answer is the design whose **arithmetic clears the budget** — usually cache + cascade + batch, not "buy a bigger model" or "hope the tier is enough". Show the maths in your head: cost = (input_tokens × in_price + output_tokens × out_price) ÷ 1e6.
:::

| Lever | Typical saving | Precondition |
| --- | --- | --- |
| Prompt caching | 40–90% of input cost on the cached prefix | Stable prefix ≥ ~1024 tokens |
| Cascade / routing | 30–70% overall | A reliable validation check to escalate on |
| Batch API | 50% on eligible traffic | Latency tolerance up to 24 h |
| Output trimming / structured output | 10–40% output cost | Don't drop needed content |

---

## 1.12 Latency budget decomposition

A p95 SLA is a *budget* to be allocated across the request path. If the parts sum above the SLA, the design fails before it ships.

```text
p95 target: 6,000 ms
  ├─ network + auth ingress ......  150 ms
  ├─ retrieval (hybrid + rerank) ..  400 ms
  ├─ model TTFT (Sonnet 5) ........  600 ms
  ├─ model generation (1.5k tok) .. 3,200 ms
  ├─ tool round-trip (1 call) .....  700 ms
  ├─ output validation ............  120 ms
  └─ headroom ..................... ~830 ms  ✓ fits
```

| If the budget is blown by… | Lever |
| --- | --- |
| Generation time | Smaller/faster model, fast mode, shorter output, streaming (improves perceived latency) |
| Too many tool round-trips | Fewer tools, parallelise independent calls, cache tool results |
| Retrieval | Lower k with reranking, warm the index, cache embeddings |
| Model queueing under load | Higher tier, spillover routing, Batch for non-interactive work |

:::tip[Exam signal]
"p95 is over budget and most of it is generation" → shrink the model / output / effort, or stream. Adding a multi-agent layer *increases* latency; it is the wrong answer whenever a latency budget is the binding constraint.
:::

---

## 1.13 Scenario walkthrough: designing an insurance-claims triage assistant

**Scenario.** A mid-size insurer wants to "use AI to speed up claims". Discovery surfaces: 12,000 claims/day (peak 3×), each claim has a PDF plus structured metadata; the assistant must classify claim type, extract key fields, flag likely fraud for human review, and draft a customer acknowledgement. Regulated (GDPR, EU residents), on Azure already, budget \$40k/month, p95 under 10 s for the interactive draft, and fraud flags **must not auto-deny** — a human adjuster decides. Historical fraud rate is ~4%.

**Expert reasoning trace.**

<Steps>

1. **Anchor on constraints.** GDPR + EU residents + existing Azure → placement is **Microsoft Foundry (Azure) in an EU region**; direct API is rejected on residency. Fraud auto-deny is irreversible and regulated → a **human gate** is mandatory, not optional.

2. **Pattern choice.** The steps are enumerable (classify → extract → fraud-score → draft), so this is a **workflow**, not an agentic loop. Reject multi-agent: the steps are sequential and share context; fan-out buys nothing and adds cost/latency.

3. **Model portfolio.** Extraction and classification are narrow, high-volume → **Haiku 4.5**. Fraud reasoning and the customer draft are higher-stakes → **Sonnet 5**. Reserve escalation to **Opus 5** only for low-confidence fraud cases (a cascade on a *validation* signal, never on the model's self-reported confidence).

4. **Capacity + cost.** 12k/day base, peak 3× → ~0.4 req/s average, ~1.25 req/s peak — well within tier; no Batch needed for the interactive path, but the nightly bulk re-scoring can use Batch. Cost is dominated by the PDF input tokens → cache the stable system/policy prefix; the arithmetic lands under \$40k.

5. **Feedback loop.** Adjuster accept/override on fraud flags is the gold-label stream → feeds a per-segment eval (by claim type) and recalibrates the fraud threshold. Without this loop the design is incomplete.

6. **Record the ADR.** Placement (Foundry EU), pattern (workflow), portfolio (Haiku/Sonnet/Opus cascade), human gate on fraud, and the rejected alternatives (direct API, multi-agent, auto-deny) with reasons.

</Steps>

**Why each tempting alternative is wrong:** direct API ignores EU residency; multi-agent over-engineers a sequential workflow; auto-deny removes the mandatory human gate on an irreversible, regulated action; escalating on self-reported confidence is anti-pattern #4; a single aggregate accuracy number would hide per-claim-type failure.

---

## 1.14 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Agentic is more advanced, so it's the better design." | Agentic is a tool for *unknowable* control flow; a workflow is better when steps are known. | The over-engineering distractor is the most common wrong answer in D1. |
| "A bigger context window removes the need for RAG." | A 1M window still costs per token and can degrade on needle-in-haystack; retrieval is cheaper and fresher. | Distractors offer "stuff the corpus into context" — reject it on cost/freshness. |
| "The most capable model is the safe default." | The cheapest model that clears the quality bar is the right default; capability is routed to where it's needed. | Blanket-Opus answers blow cost/latency budgets. |
| "A 429 means something is broken." | 429 is expected back-pressure; handle with backoff + `retry-after`. | Answers that treat 429 as fatal or swallow it are wrong. |
| "Fallback just means retry a cheaper model." | Fallback must preserve feature parity (e.g. Fable 5.1 thinking blocks) and log degraded mode. | Silent-downgrade distractor corrupts agentic sessions. |
| "Cloud placement is a performance choice." | It is primarily a residency/IAM/commitment (compliance) choice. | Residency signals in the stem override novelty and cost. |
| "Success criteria are a single accuracy number." | Criteria must be per-segment and tied to the value pillar. | Aggregate-metric trap hides segment failures. |

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Reaching for multi-agent orchestration for a task a linear workflow handles | Over-engineered; higher cost, latency and failure surface for no benefit |
| Terminating an agent loop by parsing text for "done" | Should check `stop_reason`; string matching is brittle and unsafe |
| Using an iteration cap as the primary stopping mechanism | A cap is a backstop; control flow should be driven by `stop_reason` |
| Choosing direct API when the stem states EU residency / FedRAMP | Ignores a binding constraint; use Bedrock/Vertex/Foundry |
| Designing with no feedback loop | Incomplete architecture; no path from production signals to improvement |
| Picking the most powerful model everywhere to "be safe" | Blows cost/latency budgets; routing/cascades exist for this |
| Treating a 429 as a fatal error | Expected under burst; retry with backoff + `retry-after` |
| Silent fallback to an older model in a Fable 5.1 session | Drops thinking blocks; corrupts the agentic loop; hides degradation |
| Setting success criteria as a single aggregate number | Masks per-segment failure; criteria must be per-segment |
| Skipping discovery and designing from the vague ask | Unanchored design; the binding constraint is never surfaced |
| Answering a latency-budget stem by adding a multi-agent layer | Multi-agent increases latency; wrong when p95 is binding |
| "Buy a bigger model" as the cost fix when the arithmetic favours cache + cascade + batch | Ignores the cost model; a bigger model raises unit cost |
| Choosing placement on performance/novelty when a residency signal is present | Residency/IAM/commitment govern placement, not speed |
| Auto-executing an irreversible/regulated action without a human gate | Removes the mandatory approval gate; unsafe and non-compliant |
| Treating capacity as "the tier is probably fine" without computing RPM/ITPM/OTPM | The binding limit is whichever of the three you breach first |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A retailer asks for 'an AI agent to handle customer emails'. Volume is 400 emails/hour, mostly order-status lookups against one API, with a p95 latency target of 5 s and a tight cost ceiling. What should the architect propose FIRST? (Select one)">
    A. A multi-agent system with a coordinator and specialised subagents for each email type.
    B. An augmented-LLM or simple workflow: classify intent, call the order-status tool, return a structured reply — because the steps are knowable and the SLA/cost budget favour the simplest pattern.
    C. An agentic loop with a large tool catalogue so it can handle anything.
    D. Fine-tuning a model on historical emails before any pipeline exists.

    **Answer: B.** The steps are enumerable (classify → lookup → reply), latency and cost are binding, and the volume is modest. The simplest pattern that meets the criteria wins. Multi-agent (A) and a broad agentic loop (C) are over-engineered, raising cost, latency and failure surface. Fine-tuning (D) is premature with no pipeline or eval baseline.
  </AccordionItem>

  <AccordionItem title="Q2 · A healthcare provider (EU-based, GDPR, data must not leave the EU) wants a Claude assistant. They already run everything on Google Cloud. Which placement is BEST? (Select one)">
    A. Direct Anthropic API for earliest access to new models.
    B. Google Vertex AI in an EU region, inheriting GCP IAM and residency.
    C. Amazon Bedrock, because it has the most compliance certifications.
    D. Whichever is cheapest per token.

    **Answer: B.** Two binding constraints — EU residency and an existing GCP footprint (IAM, commitment). Vertex AI in an EU region satisfies both. Direct API (A) risks residency. Bedrock (C) is compliant but ignores the existing GCP investment and identity model. Cost (D) cannot override a legal residency requirement.
  </AccordionItem>

  <AccordionItem title="Q3 · An agent occasionally never finishes a task; a junior engineer proposes stopping the loop after 10 iterations and also scanning the model's text for the word 'complete'. What is the architect's correct guidance? (Select two)">
    A. Drive termination from `stop_reason` (`end_turn` / no further `tool_use`).
    B. Keep the iteration cap, but only as a safety backstop, not the primary control.
    C. Keep parsing the text for 'complete' as the main signal.
    D. Remove all limits and trust the model to stop.
    E. Lower the temperature so it stops sooner.

    **Answer: A and B.** Control flow must key off `stop_reason`; an iteration cap is a legitimate *backstop* against runaway loops but not the primary mechanism. Parsing natural language (C) is brittle and unsafe. Removing all limits (D) risks runaway cost. Temperature (E) does not govern termination.
  </AccordionItem>

  <AccordionItem title="Q4 · A stakeholder says 'make support better with AI' and offers no numbers. What is the BEST first action? (Select one)">
    A. Start building an agentic system immediately.
    B. Run discovery to define measurable, per-segment success criteria and enumerate constraints before designing.
    C. Pick Opus 5 because it is the most capable.
    D. Assume a 90% deflection target and proceed.

    **Answer: B.** With no success criteria or constraints, any design is unanchored. Discovery surfaces the value pillar, numeric targets (per segment), volume, data sensitivity and constraints. Building (A), defaulting to the biggest model (C), or inventing a target (D) all skip the anchoring step the exam rewards.
  </AccordionItem>

  <AccordionItem title="Q5 · Peak load is 30 requests/sec with ~8k input tokens each. The team hits frequent 429s on the target model's tier. Which combination is the SOUNDEST response? (Select two)">
    A. Move latency-tolerant jobs to the Message Batches API and add exponential backoff with jitter honouring `retry-after`.
    B. Retry immediately in a tight loop until it succeeds.
    C. Request a higher usage tier and/or route overflow to a second model.
    D. Swallow the 429 and return an empty result as success.
    E. Switch every request to Opus 5.

    **Answer: A and C.** ITPM/RPM ceilings are exceeded at peak; shifting tolerant work to Batch (50% discount, 24 h) plus backoff, and raising the tier or spilling over to a second model, are the correct capacity levers. Tight-loop retry (B) worsens the storm. Swallowing errors (D) is a silent-failure anti-pattern. Upgrading every request to Opus (E) raises cost and does not fix rate limits.
  </AccordionItem>

  <AccordionItem title="Q6 · Which scenario genuinely justifies a multi-agent (coordinator + subagents) design? (Select one)">
    A. A three-step, sequential document-cleanup pipeline with shared context.
    B. A research task that fans out into several independent investigations with isolated context, then synthesises results.
    C. A single order-status lookup with a tight latency budget.
    D. Any task, to be safe.

    **Answer: B.** Independent parallel subtasks with context isolation and fan-in synthesis are the textbook multi-agent signal. Sequential shared-context work (A) is a workflow. A single lookup (C) is augmented-LLM. "Any task" (D) is the over-engineering trap.
  </AccordionItem>

  <AccordionItem title="Q7 · A design routes all traffic to Opus 5. Cost is 4× budget but accuracy is fine. What is the BEST optimisation that preserves quality? (Select one)">
    A. Switch everything to Haiku 4.5 and accept lower accuracy.
    B. Introduce a cascade: attempt Haiku 4.5 / Sonnet 5 first, escalate to Opus 5 only when confidence/validation checks fail, and add prompt caching for the stable prefix.
    C. Reduce the number of users.
    D. Remove the eval suite to save compute.

    **Answer: B.** Cascades route cheap-first and escalate on failure, cutting cost while preserving quality on hard cases; caching amortises the stable prefix. Blanket Haiku (A) sacrifices accuracy. Cutting users (C) or evals (D) does not address unit economics responsibly.
  </AccordionItem>

  <AccordionItem title="Q8 · The primary model returns 529 (overloaded) under a traffic spike. Which fallback design is BEST? (Select one)">
    A. Return an error to every user until it recovers.
    B. Retry with backoff; if still failing, fall back to a secondary model/provider and log the request as served in degraded mode.
    C. Silently return cached-but-stale answers as if fresh.
    D. Immediately fail over to a much older model in an active Fable 5.1 thinking session.

    **Answer: B.** Retry-then-fallback with explicit degraded-mode logging keeps the system available and observable. Hard failure (A) is poor HA. Passing stale answers off as fresh (C) is silent failure. Failing an active Fable 5.1 session to an older model (D) drops thinking blocks and corrupts the session.
  </AccordionItem>

  <AccordionItem title="Q9 · An architect must decide build vs buy for the agent runtime. The company lacks ops capacity, needs it live in six weeks, and the capability is not a differentiator. What is the BEST call? (Select one)">
    A. Build a custom Agent SDK runtime and sandbox from scratch.
    B. Use managed agents (Anthropic hosts the loop and sandbox) to reduce ops burden and hit the deadline.
    C. Delay the project until an ML platform team is hired.
    D. Buy a competitor's product with no Claude support.

    **Answer: B.** No ops capacity, tight deadline, commodity capability → buy/managed. Managed agents remove loop and sandbox operations. Building (A) contradicts the constraints; delaying (C) fails the deadline; (D) abandons the requirement.
  </AccordionItem>

  <AccordionItem title="Q10 · What belongs in the 'feedback loop' stage of a reference architecture? (Select two)">
    A. Offline eval runs against a golden set on each release.
    B. Online metrics, traces and cost telemetry that inform prompt/model/retrieval changes.
    C. The initial request normalisation and payload validation.
    D. The structured-output schema definition.
    E. The load balancer in front of the API.

    **Answer: A and B.** The feedback stage closes the loop from production signals back into the system's prompts, routing and retrieval — offline evals and online telemetry/traces do exactly that. Input normalisation (C) and schema definition (D) belong to input/output stages; the load balancer (E) is infrastructure, not feedback.
  </AccordionItem>

  <AccordionItem title="Q11 · A stem states a hard p95 latency of 3 s and a modest accuracy bar for a high-volume classification task. Which posture is BEST? (Select one)">
    A. Opus 5 with xhigh effort for maximum quality.
    B. Haiku 4.5 (or Sonnet 5) possibly with fast mode, minimal tool round-trips, and prompt caching — because latency and volume are the binding pillars and the accuracy bar is modest.
    C. A multi-agent pipeline for robustness.
    D. Long-context stuffing of the entire knowledge base per request.

    **Answer: B.** The binding pillar is latency at volume with only a modest accuracy need — the cheapest/fastest model that clears the bar, with fewer round-trips and caching. Opus + xhigh (A) maximises latency and cost against the SLA. Multi-agent (C) adds latency. Long-context stuffing (D) inflates tokens, cost and latency.
  </AccordionItem>

  <AccordionItem title="Q12 · Why record an ADR for the workflow-vs-agentic choice? (Select one)">
    A. To satisfy a documentation quota.
    B. To capture the context, the rejected alternatives and their reasons, and the accepted trade-off, so the decision can be reviewed and revisited as constraints change.
    C. Because ADRs replace evals.
    D. Because agentic systems cannot be built without one.

    **Answer: B.** An ADR makes the reasoning and trade-offs explicit and reviewable — exactly the altitude the exam rewards. It is not a formality (A), does not replace evaluation (C), and is not a technical prerequisite (D).
  </AccordionItem>

  <AccordionItem title="Q13 · A workload runs 8 req/s average at 6k input + 1.5k output tokens on Sonnet 5 (\$2/\$10 per MTok). Roughly what is the per-request cost, and which lever cuts it MOST for a stable-prefix workload? (Select one)">
    A. About \$0.027; the biggest single lever is prompt caching on the stable prefix.
    B. About \$0.27; switch everyone to Opus 5.
    C. About \$0.003; do nothing.
    D. Cost is unknowable without a load test.

    **Answer: A.** `6k×\$2/1e6 + 1.5k×\$10/1e6 = \$0.012 + \$0.015 = \$0.027`. With a stable prefix, prompt caching (read ≈ 0.1× input) removes most of the input cost — the largest lever here. Opus (B) raises unit cost; \$0.003 (C) is off by an order of magnitude; the cost is directly computable (D).
  </AccordionItem>

  <AccordionItem title="Q14 · A p95 budget of 6 s is being blown, and traces show generation of a long output dominates the time. Which change BEST fits? (Select one)">
    A. Add a coordinator and three subagents for robustness.
    B. Use a smaller/faster model or fast mode, shorten/stream the output, and lower effort where adequate — because generation is the dominant term.
    C. Increase k in retrieval to improve quality.
    D. Move the interactive path to the Batch API.

    **Answer: B.** The latency budget is dominated by generation, so shrink the model/output/effort or stream. Multi-agent (A) adds latency; larger k (C) adds retrieval time; Batch (D) has up-to-24 h latency and cannot serve an interactive p95 SLA.
  </AccordionItem>

  <AccordionItem title="Q15 · Peak load is 20 req/s at 6k input tokens each; the tier caps ITPM at 4,000,000. What is the binding constraint and the sound response? (Select two)">
    A. ITPM is the binding limit (peak ITPM ≈ 7.2M > 4M).
    B. RPM is the binding limit; nothing else matters.
    C. Shift latency-tolerant traffic to Batch and/or request a higher tier or spill overflow to a second model.
    D. Ignore it; bursts are rare.
    E. Send everything to Opus 5 to be safe.

    **Answer: A and C.** `peak_ITPM = 20×60×6,000 = 7.2M`, which exceeds the 4M cap, so ITPM binds first. The fix is to reduce interactive ITPM via Batch, a higher tier, or spillover routing. RPM-only (B) misreads the maths; ignoring bursts (D) causes 429 storms; Opus for all (E) raises cost and ITPM.
  </AccordionItem>

  <AccordionItem title="Q16 · An EU-regulated insurer on Azure wants claims triaged with a fraud flag; fraud must never auto-deny. Which combination is BEST? (Select two)">
    A. Deploy on Microsoft Foundry in an EU region for residency and existing IAM.
    B. Route fraud auto-deny straight to the model to save adjuster time.
    C. Insert a mandatory human approval gate before any denial (irreversible, regulated action).
    D. Use the direct Anthropic API for the newest model.
    E. Set a single aggregate accuracy target across all claim types.

    **Answer: A and C.** EU residency + existing Azure → Foundry EU; an irreversible regulated action requires a human gate. Auto-deny (B) removes the mandatory gate; direct API (D) risks residency; a single aggregate target (E) hides per-claim-type failure.
  </AccordionItem>

  <AccordionItem title="Q17 · A design falls back from Opus 5 to a much older model automatically during a 529 spike, inside an active Fable 5.1 thinking session. Users see corrupted multi-turn behaviour. What is the correct fix? (Select one)">
    A. Increase the iteration cap.
    B. Fall back only to a model with feature parity (equal-or-newer thinking support), log the degraded mode, and never silently downgrade a thinking session.
    C. Disable retries so it fails fast.
    D. Parse the output text to detect corruption.

    **Answer: B.** Fallback must preserve feature parity — thinking blocks are readable only by the producing model or newer — and degraded mode must be logged, not silent. Iteration caps (A) and text parsing (D) don't address feature parity; disabling retries (C) hurts availability.
  </AccordionItem>

  <AccordionItem title="Q18 · A stem gives request rate, token sizes, and a hard monthly budget, and asks the MOST cost-effective design that preserves quality. What is the BEST approach? (Select one)">
    A. Pick the most capable model so quality is never in question.
    B. Compute per-request cost, then combine prompt caching on the stable prefix, a cheap-first cascade escalating on validation failure, and Batch for latency-tolerant traffic until the arithmetic clears the budget.
    C. Guess a design and adjust after launch.
    D. Remove evaluation to reduce compute cost.

    **Answer: B.** Cost-effectiveness under a stated budget is an arithmetic exercise: cache + cascade + batch layered until cost < budget, with quality preserved on hard cases via the cascade. Blanket-capable (A) overspends; guessing (C) is unanchored; cutting evals (D) removes the quality guard.
  </AccordionItem>
</Accordions>

## Key takeaways

- Start every design from discovery: the value pillar, measurable **per-segment** success criteria, and the binding constraints.
- Decompose into input → processing → output → **feedback loop**; a design without the loop is incomplete.
- Choose the **simplest pattern** that meets the criteria: augmented-LLM → workflow → agentic. Over-engineering is the top wrong answer.
- Agentic loops terminate on **`stop_reason`**; iteration caps are backstops, never the primary control.
- Multi-agent only when subtasks are genuinely parallel/isolated; assign cheaper models to narrow subagents.
- Cloud placement follows **residency, IAM and existing commitments** — Bedrock/Vertex/Foundry for those constraints, direct API otherwise.
- Plan capacity against RPM/ITPM/OTPM tiers; use Batch API, backoff and spillover routing.
- Design for failure with retries and **model/provider fallback**, minding feature parity (Fable 5.1 thinking blocks).
- Record decisions as **ADRs** that name rejected alternatives and accepted trade-offs.
- **Model the arithmetic**: per-request cost = (in_tok×in_price + out_tok×out_price)/1e6; layer caching + cascade + Batch until it clears the budget.
- Treat a **p95 SLA as a budget** decomposed across the request path; if generation dominates, shrink model/output/effort or stream — never add latency with multi-agent.
- Compute **RPM / ITPM / OTPM** and design around whichever binds first; "the tier is probably fine" is not capacity planning.
- Put a **human gate** on every irreversible/regulated action, and make fallbacks preserve **feature parity** with degraded-mode logging.
