# Decision Frameworks

Every A-versus-B trade-off the four exams test, consolidated into one reference with the scenario signal that points to each answer.

The exams are trade-off exams. This page consolidates the decisions across all four blueprints. For each: the options, the signal words in a stem, and the exam-correct choice.

## Architecture

| Decision | Option A | Option B | Signal → choice |
| --- | --- | --- | --- |
| Workflow vs agent | **Workflow** – predefined code paths | **Agent** – model directs its own steps | Steps known and repeatable → workflow. Open-ended, path depends on findings → agent. Default to the simplest. |
| Single call vs chain | One prompt | Prompt chaining with gates | Distinct stages with checkable intermediate output → chain |
| Routing | Single generalist prompt | Classifier → specialist prompts/models | Heterogeneous inputs with different handling → route (cheap router) |
| Parallelization | Sequential | Sectioning (independent parts) / voting (same task ×N) | Independent subtasks or need for consensus → parallel |
| Orchestrator-workers | Fixed decomposition | Orchestrator decomposes dynamically | Subtasks unknown until inspected → orchestrator |
| Evaluator-optimizer | Single pass | Generate → evaluate → refine | Clear evaluation criteria and iterative value → loop; evaluator in separate session |
| Hosting | **Managed Agents** – Anthropic hosts loop + sandbox | **Agent SDK / Tool Runner** – you host | Need control over runtime, network, data locality → self-host; want least ops → managed |
| Subagents vs dynamic workflows | Subagents/forks | Scripted dynamic workflows | A few delegated tasks → subagents; dozens–hundreds → dynamic workflows |
| Build vs escalate (Associate) | Configure in claude.ai/Projects | Escalate to Developer/Architect | Needs API, automation, integration, or system-level guarantees → escalate |

## Loop control and enforcement

| Decision | Correct | Anti-pattern |
| --- | --- | --- |
| Loop termination | Branch on `stop_reason` | Parse prose for "done"; arbitrary iteration cap as primary stop |
| Critical business rule | Programmatic **hook** (PreToolUse) / tool permission | System-prompt sentence |
| Escalation trigger | Explicit request → immediate; beyond capability → resolve-then-escalate | Sentiment; self-reported confidence |
| Error reporting | Structured: category, retryable, partial results | Generic message; silent empty success |
| Self-review | Fresh session or different model | Same-session "are you sure?" |
| Quality metric | Per-segment / per-document-type | Aggregate only |
| Tool count | 4–5 focused tools; tool search + `defer_loading` beyond ~10 | 18 tools loaded upfront |

## API mechanics

| Decision | Option A | Option B | Signal → choice |
| --- | --- | --- | --- |
| Realtime vs Batch | Messages API | Message Batches (50% off, ≤24 h) | "Overnight", "cost matters, latency doesn't", bulk → Batch |
| Streaming | Non-streaming | SSE streaming | Interactive UI, long outputs, perceived latency → stream |
| Structured output | Prompt "return JSON" | `output_config.format` / `strict` tools | Machine-consumed output → schema-enforced; always validate + retry |
| Forced tool call | `tool_choice: any/tool` | `auto` + instruction / `strict` / structured output | Fable 5.1 → B (A is 400) |
| Thinking | `budget_tokens` | `{"type":"adaptive"}` | Current models → adaptive (Haiku 4.5 is the exception) |
| Effort | `high` (default) | `low`/`medium` for mechanical; `xhigh` for hardest | Subagent doing routine work → low/medium |
| Caching | None | `cache_control` on stable prefix | Repeated long prefix ≥1024 tokens → cache; stable first |
| Context too long | Client-side truncation | Context editing / compaction (server-side) | Verbose tool results → edit; long narrative → compact; Fable 5.1 → server-side only |
| Persistent state | Conversation history | Memory tool / files / DB | Must survive compaction or sessions → external |
| Retry | Retry everything | Backoff on 429/5xx/529 only; honour `retry-after` | 4xx client errors → fix, don't retry |
| Fallback model | Any | Newer-or-equal model | From Fable 5.1 to older → thinking blocks dropped; plan for re-planning |

## Model selection

| Signal | Choice |
| --- | --- |
| Classification, routing, extraction at scale, sub-second | Haiku 4.5 |
| Balanced quality/cost general tasks, customer chat | Sonnet 5 |
| Complex agentic coding, enterprise default | Opus 5 |
| Frontier reasoning, budget secondary, accept constraints | Fable 5.1 |
| Mixed difficulty | Cascade cheap → expensive with external validator |
| Latency critical | Smaller model, streaming, shorter output, caching, fast mode |
| Cost critical | Smaller model, caching, Batch, trim output, lower effort |

## Prompting and context

| Decision | Choice | Signal |
| --- | --- | --- |
| Zero-shot vs few-shot | Few-shot | Format-sensitive, ambiguous, edge cases |
| CoT vs extended thinking | Extended/adaptive thinking | Hard reasoning; keep visible output clean |
| Document placement | Documents first, question last | Long context recall + cacheable prefix |
| Data vs instruction boundary | XML tags | Untrusted content; injection risk |
| Prompt as asset | Version, test, cache-stable | Production system prompt |
| Context isolation | Subagent with explicit context | Protect main context; parallel exploration |
| Plan vs direct (Claude Code) | Plan mode | Multi-file, architectural, unfamiliar, irreversible |
| Skill vs CLAUDE.md vs command vs subagent vs MCP | See table | Procedure → Skill; conventions → CLAUDE.md; parameterised one-liner → command; isolated delegated task → subagent; external system → MCP |

## Tools and MCP

| Decision | Choice | Signal |
| --- | --- | --- |
| Custom tool vs MCP server | MCP | Reused across hosts/agents; external system |
| stdio vs Streamable HTTP | HTTP | Remote, multi-user, needs OAuth |
| Tool description quality | Rich: what/when/when-not/returns | Model picks wrong tool → fix description first |
| Destructive tool exposed but unneeded | **Remove it** | Not "log it" or "confirm it" |
| Identity | Per-user OAuth propagation | Multi-tenant app; authz gap |
| Parallel tool calls | Allow | Independent lookups |

## RAG (Professional)

| Decision | Choice | Signal |
| --- | --- | --- |
| RAG vs long context vs fine-tuning | RAG | Large/changing corpus, citations needed. Long context: small stable corpus fits 1M. Fine-tuning: style/format, not facts |
| Chunking | Structural/document-aware; parent-child | Structured docs; need precise hits with broad context |
| Retrieval | Hybrid (dense + BM25) + reranker | Mixed paraphrase and exact-ID queries |
| Query handling | Rewriting / multi-query / HyDE | Vague or multi-intent questions |
| Confident-but-wrong after doc refresh | **Suspect retrieval/indexing first** | Stale index, chunk drift |
| Retrieval eval | recall@k, MRR, faithfulness, answer relevance | Separate retrieval quality from generation quality |

## Evaluation

| Decision | Choice | Signal |
| --- | --- | --- |
| Grader | Exact match → rubric → LLM-as-judge (separate model, calibrated) | Increasing subjectivity |
| A/B | Pairwise with significance | Comparing prompts/models |
| Where | Offline golden set + online canary | Before and after release |
| CI | Regression suite gating deploy | Prompt or model change |
| Metric view | Per-segment + aggregate | Always both |
| Root cause | Integration-layer vs model-output vs retrieval | Check traces before changing prompts |

## Governance and safety

| Decision | Choice | Signal |
| --- | --- | --- |
| Irreversible / regulated / external / PII action | Human-in-the-loop gate | "must never", "publish", "pay", "delete", "diagnosis" |
| Data class → tool | Follow classification and policy | Restricted data → approved enterprise surface only |
| Retention | ZDR where required | Regulated; note Fable 5.1 exclusion |
| Compliance framework | GDPR (EU personal data), HIPAA (PHI, BAA), FedRAMP High (US federal via Bedrock/Vertex), SOC 2 | Sector words in the stem |
| Guardrails | Layered | Never single-layer |
| Disclosure | Tell users when AI is involved | External comms, decisions about people |

## Stakeholders and lifecycle (Professional)

| Decision | Choice | Signal |
| --- | --- | --- |
| Discovery | Structured interviews → success metrics → constraints | Vague requirements |
| Communicating trade-offs | ADR + decision matrix + cost model + risk register | Non-technical sponsor |
| Expectations | SLA/SLO alignment with measured baselines | "Executives expect 100% accuracy" |
| Lifecycle | Discovery → design → build → handoff → monitoring → iteration with exit criteria | Handoff without runbooks is the distractor |
| Model deprecation | Treat as planned lifecycle event with eval-gated migration | Retirement notice |

## Cost and performance optimization

| Decision | Choice | Signal |
| --- | --- | --- |
| Latency too high | Smaller model → streaming → shorter `max_tokens` → caching → fast mode | "p95 too slow", "users wait" |
| Cost too high | Smaller model → caching → Batch → trim output → lower effort | "spend is too high", "reduce cost" |
| Repeated long prefix | Prompt caching, stable-first layout | "same system prompt every call" |
| Bulk, latency-tolerant | Batch API (50%) | "overnight", "millions of items" |
| Mixed difficulty stream | Cascade cheap→expensive with external validator | "most tickets are simple, some hard" |
| Cutting round trips in multi-tool logic | Programmatic tool calling | "many sequential tool calls" |
| Reducing tokens in long agent runs | Context editing (clear tool results) | "context keeps growing", "verbose tool output" |
| Effort tuning | `low`/`medium` mechanical, `high` default, `xhigh` hardest | "subagent renames files" → low |

## Reliability and failure handling

| Decision | Choice | Anti-pattern rejected |
| --- | --- | --- |
| Loop exit | Branch on `stop_reason` | Prose parse; iteration cap as primary stop |
| Truncated output (`max_tokens`) | Continue / raise cap | Treat as complete |
| `refusal` stop | Explicit fallback path | Generic error; blind retry |
| `pause_turn` | Re-send to continue | Treat as failure |
| Transient 429/5xx/529 | Backoff + jitter + `retry-after` | Retry storm; retry 4xx |
| Tool failure | Structured `is_error` (category, retryable, partial) | Generic message; silent empty success |
| Side-effecting retry | Idempotency key | Double-charge on retry |
| Model fallback | Newer-or-equal only | Down-migration drops Fable 5.1 thinking blocks |
| Escalation trigger | Explicit request / beyond capability | Sentiment; self-reported confidence |

## Data and knowledge management

| Decision | Choice | Signal |
| --- | --- | --- |
| Small stable corpus fits window | Long context (no retrieval) | "50-page handbook, rarely changes" |
| Large / changing corpus, citations | RAG | "thousands of docs", "must cite" |
| Style/format not facts | Fine-tuning / examples | "match our tone", "always this format" |
| State across sessions | Memory tool / external store | "remember across chats", "survive compaction" |
| Structured docs, precise + broad | Parent-child chunking | "tables and sections", "need context around hits" |
| Vague / multi-intent query | Query rewriting / multi-query / HyDE | "users ask fuzzy questions" |
| Mixed exact-ID + paraphrase queries | Hybrid (dense + BM25) + reranker | "order numbers and natural language" |
| Confident-but-wrong after refresh | Suspect retrieval/indexing first | "worked before the doc update" |

## Prompting depth

| Decision | Choice | Signal |
| --- | --- | --- |
| Format-sensitive / edge cases | Few-shot with representative examples | "output varies", "specific format" |
| Hard reasoning, clean output | Adaptive/extended thinking | "show your work but keep answer clean" |
| Untrusted content in prompt | XML boundaries, treat as data | "user-supplied", "web content" |
| Steer output start | Prefill (or structured output for guarantees) | "always begins with", "JSON only" |
| Machine-consumed output | `output_config.format` + validate/retry | "downstream system parses it" |
| Protect main context | Subagent with explicit context | "parallel exploration", "don't pollute context" |

## Signal word → principle quick index

Fast lookup: the exact phrase item writers use, and the principle it points to. When two phrases appear in one stem, the **constraint** (cost, latency, compliance, reliability) usually decides.

| Stem phrase | Points to |
| --- | --- |
| "MOST cost-effective" | Cheapest model / Batch / caching that still meets constraints |
| "FIRST step" | Cheapest reversible diagnostic before big changes |
| "BEST approach" | Simplest design meeting every stated constraint |
| "Select TWO" | Two independently-correct actions; watch for one-right-one-plausible |
| "at scale" / "millions" | Haiku 4.5 + Batch + caching |
| "sub-second" / "real time" | Smaller model, streaming, low effort; not Batch |
| "overnight" / "latency does not matter" | Batch API (50%) |
| "same long prompt every request" | Prompt caching, stable-prefix first |
| "must never" / "under no circumstances" | Deterministic hook / tool permission, not a prompt sentence |
| "how does the loop know it is done" | `stop_reason: end_turn`, not prose parsing |
| "when should it escalate" | Explicit request or beyond-capability; not sentiment/self-report |
| "the agent kept going forever" | Terminate on `stop_reason`, not iteration cap |
| "returns nothing / empty" | Structured `is_error` / not_found; silent-success anti-pattern |
| "confidence" / "how sure are you" | Do not route on self-reported confidence |
| "angry" / "frustrated customer" | Sentiment ≠ complexity; do not escalate on tone |
| "overall accuracy is 95%" | Demand per-segment metrics |
| "are you sure?" in the same chat | Fresh session / different model for review |
| "18 tools" / "too many tools" | 4–8 tools; tool search + `defer_loading` |
| "forced to call a tool" + Fable 5.1 | `auto`+instruction / `strict` / structured output (forced = 400) |
| "guaranteed JSON shape" | `output_config.format` / `strict: true` + validation |
| "thinking blocks disappeared after fallback" | Fable 5.1 binding; only migrate up |
| "context window is full" | Context editing (tool results) or compaction (narrative) |
| "edited an earlier turn" + Fable 5.1 | Append-only history; use mid-conversation system messages |
| "remember across sessions" | Memory tool / external store |
| "must cite sources" | RAG + citations; provenance test |
| "worked before the document update" | Suspect retrieval/indexing first |
| "exact IDs and paraphrases" | Hybrid retrieval + reranker |
| "vague user questions" | Query rewriting / HyDE / multi-query |
| "50-page doc that rarely changes" | Long context, not RAG |
| "match our writing style" | Few-shot / fine-tuning, not RAG |
| "reused across several apps/agents" | MCP server, not a custom tool |
| "remote, multi-user tool server" | Streamable HTTP + OAuth 2.1, per-user identity |
| "acts as a shared admin account" | Authz gap; propagate end-user identity |
| "delete/refund tool it doesn't need" | Remove it (least privilege), not log/confirm |
| "web content / tool output told it to…" | Indirect injection; treat as data, boundaries |
| "hard rule in the system prompt" | Prompt-as-enforcement anti-pattern; use a hook |
| "multi-file / architectural change" (Claude Code) | Plan mode |
| "single obvious edit" | Direct mode |
| "reusable multi-step procedure" | Skill |
| "repo conventions / build commands" | CLAUDE.md |
| "parameterised one-liner" | Slash command |
| "isolated delegated task" | Subagent |
| "dozens–hundreds of agents" | Dynamic workflows |
| "a few delegated tasks" | Subagents / forks |
| "named roles collaborating" | Agent teams |
| "scheduled recurring job" | Routines |
| "distribute commands/agents/hooks/MCP" | Plugins |
| "run in CI / headless" | `claude -p --output-format json`, restricted tools |
| "grade subjective quality" | Rubric → LLM-as-judge (separate session), calibrated |
| "compare two prompts/models" | Pairwise A/B with statistical significance |
| "gate the deploy" | CI regression suite on the golden set |
| "before and after release" | Offline golden set + online canary |
| "regulated / irreversible / external action" | Human-in-the-loop gate |
| "EU personal data" | GDPR; DPIA for high-risk |
| "health data / PHI" | HIPAA + BAA |
| "US federal" | FedRAMP High via Bedrock/Vertex |
| "prompts must not be retained" | ZDR (note Fable 5.1 exclusion) |
| "non-technical sponsor / executives" | ADR + decision matrix + cost model + risk register |
| "expect 100% accuracy" | Set SLO/SLA against measured baselines |
| "handoff to the team" | Runbooks, monitoring, exit criteria |
| "which model should we use" | Match capability matrix to the binding constraint |
| "reduce hallucinations" | Grounding (RAG/citations), validation, not just "tell it not to" |

:::tip[Exam signal]
When a stem stacks a capability need against a compliance or cost constraint, the constraint wins. "We want the MCP connector but we are ZDR-only on Fable 5.1" is impossible as stated — recognise the conflict rather than picking the shiny feature.
:::
