AI Cert Prep
Type to search documentation.

Domains

D1 · Agentic Architecture and Orchestration

Choosing between workflows and agents, the six orchestration patterns, the agentic loop and stop_reason-driven termination, coordinator/subagent hierarchies, escalation design, error propagation, state and memory, and the cost/latency of multi-agent systems.

This is the heaviest domain on the Architect exam – roughly 16 of 60 items – and it is where the ten anti-patterns bite hardest. It tests whether you can look at a described system and choose the simplest correct architecture: a single prompt, a fixed workflow, or a genuine agent; and if an agent, which orchestration pattern, how it terminates, how it escalates, and how it fails safely. Nearly every item is a trap between an over-engineered answer and a constraint-blind one.

Learning objectives

By the end of this page you should be able to:

  1. Apply the workflow vs agent decision framework and justify it against constraints.
  2. Select among the six orchestration patterns (prompt chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, autonomous agent) and describe each one’s failure modes.
  3. Implement the agentic loop and drive termination from stop_reason – not natural-language parsing (anti-pattern 1) or arbitrary iteration caps (anti-pattern 2).
  4. Design coordinator/subagent hierarchies with explicit context passing, result aggregation and partial-failure handling.
  5. Design escalation on explicit request (immediate) or capability (resolve first) – never sentiment or self-reported confidence (anti-patterns 4–5).
  6. Choose hooks vs prompts for enforcement (anti-pattern 3) and Managed Agents vs Agent SDK/Tool Runner for hosting.
  7. Propagate structured errors (category, retryable, partial results) with idempotency and retries (anti-patterns 6–7).
  8. Decide between conversation, files, database and the memory tool for state, and model the cost and latency of multi-agent systems with per-agent observability.

1.1 Workflow vs agent: the first decision

An agent is a system where the model dynamically directs its own process and tool use, deciding what to do next in a loop until it judges the task complete. A workflow is a system where the code orchestrates the model through predefined paths – the control flow is written by you, not decided by the model.

The exam’s default posture, straight from Anthropic’s guidance, is: find the simplest solution and only increase complexity when it demonstrably improves outcomes. A single well-engineered prompt beats a workflow; a workflow beats an agent; an agent beats a multi-agent system. Add autonomy only when the task genuinely requires open-ended decision-making at runtime.

Signal in the stemPoints toWhy
Steps are known in advance, fixed orderWorkflowDeterministic control flow is cheaper, faster, testable
Predictable branching on a classifiable inputWorkflow (routing)You can enumerate the branches
Task requires runtime decisions you cannot enumerateAgentOnly the model can decide the path from context
Open-ended research, debugging, explorationAgentNumber and order of steps unknown up front
Latency, cost or auditability are paramountWorkflowFewer model calls, deterministic, easy to trace
One model call with good prompting sufficesSingle promptDo not build a system at all

Exam signal

Words like “fixed sequence”, “known steps”, “every request follows the same path” point to a workflow. Words like “open-ended”, “the number of steps varies”, “must decide at runtime”, “explore until” point to an agent. When the stem gives you a simple task and an elaborate architecture as the “recommended” answer, that answer is almost always the over-engineered distractor.


1.2 The six orchestration patterns

The exam expects fluency in all six. For each: what it is, an ASCII diagram, when to use it, and its failure modes.

Pattern 1 – Prompt chaining

Decompose a task into a fixed sequence of steps, each model call consuming the previous output. Add programmatic gates (“checks”) between steps.

text
Input ──► [Call 1] ──► gate ──► [Call 2] ──► gate ──► [Call 3] ──► Output
│ │
fail→fix fail→stop

Use when: the task cleanly splits into fixed subtasks and you trade a little latency for much higher accuracy per step (e.g., outline → draft → polish).

Failure modes: error compounding down the chain; a wrong early step poisons everything after; latency is the sum of all calls.

Pattern 2 – Routing

Classify the input, then dispatch to a specialised prompt/model/flow.

text
┌─► [Billing prompt · Haiku]
Input ─► [Router]─┼─► [Technical prompt · Sonnet]
└─► [Escalation flow]

Use when: inputs fall into distinct classes best handled separately, and misclassification is cheaper than one-size-fits-all. Lets you send easy classes to cheaper models.

Failure modes: router misclassification cascades; too many classes make the router brittle; an unhandled class silently mishandled.

Pattern 3 – Parallelization (sectioning and voting)

Run independent subtasks concurrently (sectioning) or run the same task multiple times and aggregate (voting).

text
Sectioning Voting
┌─►[Worker A]─┐ ┌─►[Run 1]─┐
Input ──┼─►[Worker B]─┼─►[Merge] Input ──┼─►[Run 2]─┼─►[Vote]─► Output
└─►[Worker C]─┘ └─►[Run 3]─┘

Use when: subtasks are independent (sectioning) and latency matters, or when multiple attempts raise confidence (voting) – e.g., guardrail plus main answer, or majority-vote code review.

Failure modes: sectioning that assumes independence when tasks actually interact; voting that hides a systematic bias shared by all runs; cost multiplied by fan-out.

Pattern 4 – Orchestrator-workers

A central model (orchestrator) dynamically decomposes a task, spawns worker calls, and synthesises. Unlike parallelization, the subtasks are not predefined – the orchestrator decides them at runtime.

text
Input ─► [Orchestrator] ─┬─► [Worker: search docs]
▲ ├─► [Worker: read files]
│ └─► [Worker: summarise]
└────────────── synthesise ◄────────┘

Use when: you cannot predict the subtasks in advance (e.g., multi-file code changes, research where the sub-questions depend on findings).

Failure modes: the coordinator/subagent anti-patterns – relying on context auto-inheritance, no partial-failure handling, unbounded fan-out cost. See §1.5.

Pattern 5 – Evaluator-optimizer

One call generates, a second evaluates against explicit criteria and returns feedback, and the loop repeats until the criteria pass.

text
┌───────────────────────────┐
Input ─►[Generator]─► draft ─►[Evaluator]─► pass? ─► Output
▲ │ fail
└────────── feedback ◄─────────┘

Use when: you have clear, checkable success criteria and iteration measurably improves the result (translation quality, code that must pass tests, writing against a rubric).

Failure modes: the evaluator and generator share a session and therefore a bias (anti-pattern 9 – same-session self-review); no objective stopping criterion so it loops forever or stops arbitrarily; vague criteria produce useless feedback.

Evaluator independence

The evaluator should be an independent context – a fresh session, ideally a different model – so it does not inherit the generator’s reasoning. Reusing the same conversation to “check its own work” retains the reasoning-context bias that produced the flaw.

Pattern 6 – Autonomous agent

A single model runs an open-ended loop: it decides an action, calls a tool, observes the result, and repeats until it judges the task done or a stopping condition fires.

text
┌─────────────────────────────┐
User goal ─► [Claude] ─► tool_use ─► [Env] ──┘
▲ result
└── loop while stop_reason == "tool_use"
stop when end_turn / gate / human

Use when: the task is genuinely open-ended, the environment gives reliable feedback, and errors are recoverable or gated by human approval for irreversible actions.

Failure modes: looping forever (bad termination); acting on stale or hallucinated state; taking irreversible actions without a gate; the two loop anti-patterns below.

PatternControl flow decided byBest forSignature risk
Prompt chainingYou (fixed)Decomposable fixed tasksError compounding
RoutingYou (branch on class)Distinct input classesMisclassification
ParallelizationYou (fan-out)Independent subtasks / votingFalse independence, cost
Orchestrator-workersModel (runtime subtasks)Unpredictable decompositionContext/partial-failure
Evaluator-optimizerYou (loop on criteria)Checkable quality barShared-bias self-review
Autonomous agentModel (open loop)Open-ended tasksTermination, irreversibility

1.3 The agentic loop and stop_reason-driven termination

Every agent is a loop. The only correct way to drive it is the API’s stop_reason field. The two most-tested anti-patterns live here.

stop_reason values: end_turn, tool_use, max_tokens, stop_sequence, pause_turn, refusal.

python
from anthropic import Anthropic
client = Anthropic()
messages = [{"role": "user", "content": "Investigate the failing test and fix it."}]
while True:
resp = client.messages.create(
model="claude-opus-5",
max_tokens=4096,
tools=TOOLS,
messages=messages,
)
messages.append({"role": "assistant", "content": resp.content})
if resp.stop_reason == "tool_use":
results = run_tools(resp.content) # execute requested tools
messages.append({"role": "user", "content": results})
continue # loop: model saw the results
elif resp.stop_reason == "end_turn":
break # model is done
elif resp.stop_reason == "max_tokens":
raise OutputTruncated() # do NOT treat as completion
elif resp.stop_reason == "pause_turn":
continue # long-running server tool; resume
elif resp.stop_reason == "refusal":
escalate_to_human(resp) # model declined; stop the loop
break

Anti-pattern 1 · Parsing prose for termination

Do not terminate by checking whether the model’s text says “done”, “task complete” or “I have finished”. The model’s prose is not a control signal – it varies, it lies, it can be prompt-injected. The loop must key off stop_reason == "end_turn".

Anti-pattern 2 · Iteration caps as the primary stop

A hard cap like for i in range(10) is a safety backstop, never the primary stopping mechanism. If your loop only stops because it hit the cap, you have no idea whether the task finished. The primary stop is stop_reason; the cap exists so a runaway loop cannot burn unbounded cost. Both together: while stop_reason == "tool_use" and iterations < CAP.

A correct termination design combines: stop_reason as the primary signal, an iteration/token/cost cap as a backstop, and explicit gates (human approval for irreversible actions).


1.4 Managed Agents vs Agent SDK / Tool Runner: the hosting decision

OptionWho hosts the loop and sandboxChoose whenTrade-off
Managed AgentsAnthropic hosts the loop and execution sandboxYou want speed to production, standard tools, less infra to runLess control over the loop, environment and custom tooling
Claude Agent SDKYou host the loop (pip install claude-agent-sdk / npm i @anthropic-ai/claude-agent-sdk)You need custom tools, your own environment, deep control, self-hostingYou own retries, observability, sandboxing, scaling
Tool RunnerYou host tool execution, Anthropic drives orchestrationYou want managed orchestration but your own tool executionSplit responsibility

Exam signal

“We need to ship fast with standard tools and minimal infrastructure” → Managed Agents. “We have proprietary tools / a bespoke sandbox / must control the loop” → Agent SDK. The SDK was renamed from the Claude Code SDK; both package names appear on the exam.


1.5 Coordinator/subagent hierarchies

When one agent cannot hold the whole task, a coordinator decomposes it and delegates to subagents, each with its own isolated context window. This is the orchestrator-workers pattern at system scale (the Multi-Agent Research System scenario).

text
┌──────────────┐
User goal ───►│ Coordinator │ holds plan, aggregates, decides done
└──────┬───────┘
explicit context│ (never rely on inheritance)
┌───────────────┼────────────────┐
┌────▼────┐ ┌────▼────┐ ┌────▼────┐
│Subagent1│ │Subagent2│ │Subagent3│ isolated context each
└────┬────┘ └────┬────┘ └────┬────┘
└── structured result ──────────┘
aggregate + handle partial failure

Explicit context passing

Subagents have isolated context windows. They do not automatically inherit the coordinator’s conversation, files or findings. The coordinator must explicitly pass everything the subagent needs in its task prompt.

python
# CORRECT: the subagent gets exactly what it needs, explicitly.
subagent_task = {
"role": "user",
"content": (
f"Objective: {objective}\n"
f"Context you must use:\n{relevant_findings}\n"
f"Constraints: {constraints}\n"
f"Return: a JSON object {{'finding': str, 'sources': [str], 'confidence': str}}."
),
}

Relying on context auto-inheritance

Assuming a subagent “already knows” what the coordinator discovered is a top scenario distractor. Subagent contexts are isolated by design (this is a feature – it protects the main context). Always pass context explicitly; specify the exact return shape so results aggregate cleanly.

Result aggregation and partial-failure handling

Subagents fail independently. The coordinator must aggregate structured results and decide what to do when some fail.

python
results, failures = [], []
for sub in subagents:
try:
r = run_subagent(sub, timeout=SUB_TIMEOUT)
if r.status == "ok":
results.append(r)
else:
failures.append((sub, r.error)) # structured error, not swallowed
except TimeoutError as e:
failures.append((sub, {"category": "timeout", "retryable": True}))
# Decide: proceed with partial results, retry retryables, or escalate.
if len(results) >= MIN_QUORUM:
answer = coordinator_synthesise(results, note_missing=failures) # provenance of gaps
else:
escalate("insufficient subagent results", partial=results, failures=failures)

The correct design surfaces partial failure to the coordinator with structured detail, then makes an explicit decision (proceed with a quorum and note the gap, retry retryables, or escalate). Silently dropping a failed subagent and presenting the rest as complete is anti-pattern 7.


1.6 Escalation design

Escalation is one of the most reliably tested topics because three of the ten anti-patterns are escalation traps. The rule set:

TriggerCorrect behaviourWhy
Explicit request (“I want a human”, “speak to an agent”)Escalate immediately, no further attemptsThe user has stated intent; honour it
Capability boundary (task needs a tool/permission/authority the agent lacks)Attempt resolution first, escalate only when genuinely blockedEscalating prematurely wastes a capable path
Sentiment (user is angry/frustrated)Not an escalation trigger by itselfSentiment ≠ complexity (anti-pattern 5)
Self-reported confidence (“I’m 60% sure”)Not a triggerModels are poorly calibrated; self-report is unreliable (anti-pattern 4)
python
def should_escalate(turn) -> bool:
if turn.user_explicitly_requested_human:
return True # immediate
if turn.requires_capability_agent_lacks and turn.resolution_attempts_exhausted:
return True # capability-based, after trying
# NOT: turn.sentiment == "angry"
# NOT: turn.model_confidence < 0.7
return False

Exam signal

An option that escalates because the customer “sounds frustrated” or because the model “reported low confidence” is a distractor. The correct answer escalates on explicit request (now) or an objective capability gap (after attempting). Anger is handled with good responses, not routing.


1.7 Hooks vs prompts for enforcement

Some rules are critical – they must always hold (never delete production data, never commit secrets, always run tests before commit). Enforce these with deterministic code: Claude Code hooks or programmatic guards in your loop. Never enforce a critical rule with a prompt instruction.

Enforcement mechanismGuaranteesUse for
Prompt instructionBest-effort; probabilistic; bypassable by injectionPreferences, style, soft guidance
Programmatic hook / guardDeterministic; exit code 2 blocks the actionCritical business rules, safety, irreversible-action gates

Anti-pattern 3 · Prompt-based enforcement of critical rules

“Add a line to the system prompt telling Claude never to run destructive commands” is the wrong answer whenever the rule is critical. Prompts are probabilistic and injection-vulnerable. The correct control is a PreToolUse hook that inspects the command and returns exit code 2 to block it. (Hook mechanics are Domain 2.)


1.8 Error propagation with structured context

Errors must carry enough structure for a caller (or the model) to decide what to do: category, retryable flag, and any partial results.

python
class AgentError(Exception):
def __init__(self, category, message, retryable, partial=None):
self.category = category # e.g. "rate_limit", "validation", "timeout", "auth"
self.message = message # specific, diagnostic
self.retryable = retryable # drives backoff vs. escalate
self.partial = partial # data gathered before failure
# On the boundary, return structured error content the model can reason about:
tool_result = {
"type": "tool_result",
"tool_use_id": tu_id,
"is_error": True,
"content": json.dumps({
"category": "rate_limit", "retryable": True, "retry_after": 12,
"partial": partial_rows,
}),
}

Anti-patterns 6 & 7 · Generic errors and silent suppression

Returning "Something went wrong" (anti-pattern 6) strips the diagnostic context the model or operator needs to recover. Returning an empty result as if it were success (anti-pattern 7) is worse – it converts a failure into silent wrong data downstream. Always return the specific category, whether it is retryable, and any partial results.

Idempotency and retries

Retry 429, 5xx and 529 with exponential backoff + jitter, and respect retry-after. But only safely retry operations that are idempotent – give write operations an idempotency key so a retried “create order” does not create two orders.

python
def with_retry(fn, *, max_attempts=5):
for attempt in range(max_attempts):
try:
return fn()
except RetryableError as e:
if attempt == max_attempts - 1:
raise
sleep(min(2 ** attempt + random.random(), 30)) # backoff + jitter

1.9 State and memory vs conversation

Conversation history is not durable state. It is trimmed by context editing, summarised by compaction, and lost across sessions. Anything that must survive belongs in explicit state.

StoreSurvivesUse for
ConversationOnly within the (untrimmed) windowImmediate working context
Memory toolAcross sessions, model-managed filesFacts/learnings the agent should recall later
FilesDurable, you manageArtefacts, intermediate outputs, checkpoints
Database / external storeDurable, queryable, transactionalOrders, tickets, canonical business state

Exam signal

“Must persist across sessions” or “must survive compaction” → memory tool or external state, never conversation history. Canonical business data (an order, a ticket status) always lives in a database, with the conversation referencing it, not owning it.


1.10 Cost and latency modelling of multi-agent systems

Every agent and subagent is model calls, and multi-agent fan-out multiplies both cost and token consumption. Model this before building.

  • Cost ≈ Σ over calls of (input tokens × in-price + output tokens × out-price). Subagents each carry their own system prompt and tools as input – fan-out of 5 means ~5× the fixed overhead. Prompt caching (0.1× on cache reads) and cheaper models for simple subtasks (Haiku for classification) cut this sharply.
  • Latency: chaining adds latencies serially; parallelization bounds latency by the slowest branch; deep coordinator/subagent trees add round-trips.
  • The trade: a multi-agent research system can be 10–15× the token cost of a single call. Justify it only when the quality or breadth gain is real. If a workflow meets the bar, use the workflow.
text
Single agent: 1 system + N tool round-trips (cheapest, serial)
Parallel workers: 1 orchestrator + K×(system+tools) (fast, K× fixed cost)
Deep hierarchy: coordinator + Σ subagents + aggregation (most capable, most $$)

1.11 Observability

You cannot operate what you cannot see. A multi-agent system needs:

  • Traces per agent – each agent/subagent emits a span with inputs, tool calls, stop_reason, tokens and cost.
  • Correlation IDs – one request ID threaded through the coordinator and every subagent so you can reconstruct a single task end-to-end.
  • Metrics – per-agent latency (p50/p95), error rate by category, retry counts, tokens and cost per task, cache-hit rate.
  • Structured logs without secrets.

Exam signal

“We cannot tell which subagent caused the failure” → the missing control is per-agent traces plus a correlation ID, not more retries or a bigger model.


1.12 Orchestration variants and their failure modes

Beyond the six base patterns, item writers combine them into named variants. Recognising the variant tells you the failure mode to guard against.

Hierarchical (coordinator → sub-coordinators → workers)

text
[Root coordinator]
/ \
[Sub-coord A] [Sub-coord B]
/ \ / \
[W1] [W2] [W3] [W4] isolated contexts throughout

Use when a task decomposes into large sub-domains that themselves need decomposition. Failure modes: context loss deepens with each tier (every level must pass context explicitly); cost and latency compound (each tier adds round-trips); a mid-tier failure can orphan a whole subtree.

Pipeline with feedback (chaining + evaluator-optimizer)

text
[Extract] ─► [Transform] ─► [Load] ─► [Evaluate]
▲ │ fail
└──────────── targeted retry ◄───────┘

Use when a fixed pipeline needs a quality gate at the end that can send work back to a specific stage. Failure modes: feedback to the wrong stage; unbounded feedback loops without an objective stop; the evaluator sharing the generator’s session (#9).

Blackboard / shared-state coordination

text
[Agent A] ─┐ ┌─► reads/writes
[Agent B] ─┼─► [Shared state store] ◄─┼─ [Agent C]
[Agent D] ─┘ └─► reads/writes

Use when several agents contribute to a common evolving artefact (a research doc, a plan). Failure modes: race conditions and lost updates on the shared store; no single owner of “done”; stale reads. Fix with transactional writes, versioning, and one coordinator that decides completion.

VariantBest forSignature failureGuard
HierarchicalDeep, decomposable domainsCompounding context loss & costExplicit context each tier; justify depth
Pipeline + feedbackFixed flow with a quality gateWrong-stage / unbounded feedbackObjective stop; route to the right stage
BlackboardShared evolving artefactRaces, no owner of doneTransactions, versioning, one coordinator

Exam signal

A stem that adds tiers or agents without a stated benefit is testing whether you will over-engineer. The correct answer keeps the shallowest topology that meets the bar and names the failure mode the extra structure would introduce.


1.13 Cost and latency arithmetic you can be asked to do

The exam can give you token counts and prices and ask for the relative cost of two topologies. Use the September-2026 prices (input/output per million tokens): Opus 5 $5/$25, Sonnet 5 $2/$10, Haiku 4.5 $1/$5, Fable 5.1 $10/$50. Cache reads ≈ 0.1× input; Batch ≈ 0.5×.

Worked example — single agent vs 5-worker fan-out. Each call carries a 10k-token stable prefix (system + tools) and produces 1k output. On Opus 5:

text
Per call input = 10,000 tokens × $5 / 1,000,000 = $0.050
Per call output = 1,000 tokens × $25 / 1,000,000 = $0.025
Per call total = $0.075
Single agent, 4 tool round-trips ≈ 4 × $0.075 = $0.300
5-worker fan-out (each 1 call) + 1 synthesis = 6 × $0.075 = $0.450
...but each worker re-sends the 10k prefix, so the fixed overhead
is paid 6× instead of being amortised — fan-out is ~1.5× here,
and grows with prefix size and worker count.

Two levers that change the answer:

  • Prompt caching. If the 10k prefix is cached, each read costs 10,000 × $0.50/1,000,000 = $0.005 instead of $0.050 — the fan-out’s fixed overhead collapses and the multi-agent penalty shrinks dramatically.
  • Model mix. Route simple workers to Haiku ($1/$5) and reserve Opus for synthesis; a 5-worker Haiku + Opus-synthesis design can undercut an all-Opus single agent while adding breadth.
TopologyFixed overhead paidLatencyWhen it wins
Single agentOnce (amortised)Serial round-tripsCheapest; task fits one context
Parallel workersK× (per worker)≈ slowest branchLatency-critical, independent work
Deep hierarchyΣ across tiersSum of tier round-tripsOnly when breadth/quality gain is real

Exam signal

“MOST cost-effective” with fan-out in the stem usually rewards caching the stable prefix and routing simple subtasks to cheaper models — not “use a bigger model” and not “add more agents”. If a workflow meets the bar, the workflow is the cost-effective answer.


1.14 Human-in-the-loop gates as first-class architecture

Autonomy is a spectrum. The architect chooses where a human sits in the loop, and makes that gate deterministic, not a prompt request.

Autonomy levelHuman roleCorrect mechanism
Suggest-onlyHuman approves every actionPlan mode / ask permission on all effects
Gated-writeHuman approves irreversible actions onlyPreToolUse hook (exit 2) on payments, deletes, sends
Autonomous-with-auditHuman reviews after the factStructured logs, traces, provenance
Fully autonomousNoneOnly for reversible, low-stakes actions

The gate is placed on irreversibility and blast radius, not on the model’s confidence or the user’s tone. A refund over a threshold, a production deploy, an outbound email to a customer list — these get a hard gate regardless of how sure the model claims to be.

Exam signal

“Must always require human approval before X” → a deterministic gate (hook / permission), and the gate is chosen by X’s irreversibility. Any option that gates on confidence (#4) or sentiment (#5) is wrong even when the action is genuinely risky.


Common misconceptions

MisconceptionRealityWhy it matters on the exam
“More agents = better answers.”More agents multiply cost/latency and add failure modes; they help only when breadth or quality demonstrably improves.The over-engineering distractor is the most common wrong answer in D1.
“The model saying ‘done’ means it is done.”Only stop_reason == end_turn means done; prose is not a control signal.Anti-pattern 1 appears in many stems disguised as “flexible”.
“An iteration cap keeps my loop correct.”A cap only bounds cost; correctness comes from stop_reason.Anti-pattern 2; the cap-as-completion trap.
“Subagents can see what the coordinator knows.”Contexts are isolated; nothing is inherited — you must pass context explicitly.The signature Multi-Agent Research distractor.
“Escalate frustrated or low-confidence cases.”Escalate on explicit request or a capability gap; sentiment and self-report are unreliable.Anti-patterns 4 and 5, tested together.
“A strong system-prompt rule is enough for critical policy.”Prompts are probabilistic and injectable; critical rules need deterministic hooks.Anti-pattern 3; hooks-vs-prompts is a recurring decision.
“Retrying any failed call is safe.”Only idempotent operations are safe to retry; writes need an idempotency key.Double-charge/double-refund stems.
“Managed Agents and the Agent SDK are interchangeable.”Managed Agents host the loop/sandbox (fast, less control); the SDK is self-hosted (control, custom tools, your VPC).Hosting-choice items hinge on the stated constraint.

Scenario walkthrough — a research system that keeps “losing the plot”

Situation. A team ships a market-research assistant. A coordinator breaks a question into sub-questions and dispatches subagents that each search and summarise. In production, two problems appear: (1) subagents sometimes contradict findings the coordinator already established, and (2) when one subagent’s search API times out, the final report is delivered as if complete, occasionally missing a whole competitor. Leadership also complains the system costs 12× a single Opus call and asks whether it should be simplified. The accuracy bar was met in testing by a two-step “plan → parallel-summarise → synthesise” design.

Expert reasoning trace.

  1. Name the pattern. Sub-questions are decided at runtime from findings → this is orchestrator-workers, so expect the coordinator/subagent failure modes, not a routing or chaining fix.

  2. Diagnose problem (1). Subagents “contradicting known findings” is the classic isolated-context symptom: the subagents never received the coordinator’s established facts. The fix is explicit context passing in each subagent’s task prompt with a fixed return shape — not a bigger window (a tempting but constraint-blind option) and not a better model.

  3. Diagnose problem (2). Presenting a partial report as complete after a timeout is silent suppression (#7). The fix is a structured error (category:"timeout", retryable:true) surfaced to the coordinator, which then retries the retryable, proceeds with a quorum noting the gap, or escalates — never a silent drop and never a generic “research failed” (#6).

  4. Address the cost/simplification question. The two-step design already met the accuracy bar. The deep, expensive topology is over-engineering. The correct call is to use the simpler design and add complexity only where it demonstrably improves outcomes — while caching the stable prefix and routing simple summaries to Haiku to cut the remaining cost. Rejecting “use a bigger model” and “add more subagents” is essential.

  5. Make failures observable. Add per-agent trace spans + a correlation ID so the next timeout is attributable, rejecting “add more retries” as an observability fix.

Exam-correct decision: explicit context passing + structured partial-failure handling with a quorum + collapse to the two-step design + caching/model-mix for cost + per-agent traces. Every rejected option maps to a named anti-pattern (isolation-ignorance, #7, #6, over-engineering, confidence/sentiment escalation) — which is exactly how the distractors are built.


Exam traps in this domain

TrapWhy it is wrong
Build a multi-agent system for a fixed, known sequenceOver-engineered; a workflow (or single prompt) is correct
Terminate the loop when the model’s text says “done”Prose is not a control signal (anti-pattern 1)
Stop the agent purely by an iteration capCap is a backstop, not the primary stop; use stop_reason (anti-pattern 2)
Treat max_tokens as task completionIt means truncated output, not done
Escalate because the customer sounds angrySentiment ≠ complexity (anti-pattern 5)
Escalate on the model’s self-reported confidenceModels are poorly calibrated (anti-pattern 4)
Enforce a critical rule via a system-prompt linePrompts are probabilistic; use a hook (anti-pattern 3)
Assume subagents inherit the coordinator’s contextContexts are isolated; pass context explicitly
Return a generic error messageHides diagnostics needed to recover (anti-pattern 6)
Drop a failed subagent and present the rest as completeSilent failure becomes wrong data (anti-pattern 7)
Keep must-persist state only in conversation historyTrimmed/compacted/lost; use memory tool or a database
Retry a non-idempotent write on 429Duplicates the side effect; use an idempotency key
Add tiers/agents without a stated benefitCompounds cost, latency and context loss; keep the shallowest topology that meets the bar
“Use a bigger model” to cut fan-out costCache the stable prefix and route simple workers to cheaper models instead
Gate an irreversible action on model confidenceGate on irreversibility with a deterministic hook, never on self-report (#4)
Treat pause_turn as completionIt means resume a long-running server tool; continue the loop
Rely on a shared blackboard with no owner of “done”Use transactional writes, versioning and one coordinator that decides completion

Practice questions

Each item states how many responses to select. Attempt before revealing.

Q1 · A team processes support tickets that always follow the same three steps: classify, draft a reply, and check the reply against policy. They ask whether to build an autonomous agent. What is the BEST architecture? (Select one)

A. An autonomous agent that decides each step at runtime for flexibility. B. A prompt-chaining workflow with a programmatic policy gate between the draft and send steps. C. A multi-agent system with a coordinator and three subagents. D. A single mega-prompt that does all three at once.

Answer: B. The steps are fixed and known, so control flow should be code, not the model – prompt chaining with a gate. An autonomous agent (A) and multi-agent system (C) are over-engineered for a fixed sequence. A single mega-prompt (D) loses the per-step accuracy and the policy gate.

Q2 · An engineer's agent loop terminates when Claude's response text contains the phrase 'task complete'. Occasionally it stops early or never stops. What is the correct fix? (Select one)

A. Add more phrases to match, like ‘finished’ and ‘done’. B. Cap the loop at 10 iterations and stop there. C. Drive termination from stop_reason: continue while it is tool_use, stop on end_turn, and treat max_tokens and refusal explicitly. D. Lower the temperature so the phrasing is consistent.

Answer: C. This is anti-pattern 1 – parsing prose for termination. The control signal is stop_reason. More phrases (A) still parses prose. An iteration cap (B) is only a backstop (anti-pattern 2). Temperature (D) does not make prose a reliable control signal.

Q3 · A coordinator agent delegates a sub-question to a subagent but the subagent's answer ignores findings the coordinator already gathered. What is the root cause and fix? (Select one)

A. The subagent needs a larger context window. B. The subagent context is isolated and did not inherit the findings; the coordinator must pass the relevant context explicitly in the subagent’s task prompt. C. The coordinator should use a more capable model. D. Increase max_tokens on the subagent.

Answer: B. Subagent contexts are isolated by design. Relying on auto-inheritance is a core distractor. Context size (A, D) and model choice (C) do not address missing context that was never passed.

Q4 · A customer support agent should escalate to a human. Which TWO triggers are correct? (Select two)

A. The customer explicitly asks to speak to a human. B. The customer’s message has negative sentiment. C. The task requires a refund authority the agent lacks, after the agent has attempted the resolvable parts. D. The model self-reports 55% confidence. E. The reply is longer than 200 words.

Answer: A and C. Explicit request escalates immediately; a genuine capability gap escalates after attempting resolution. Sentiment (B) is anti-pattern 5, self-reported confidence (D) is anti-pattern 4, and length (E) is irrelevant.

Q5 · A system must never commit code that fails the test suite. A proposal adds 'Always run tests before committing' to the project's system prompt. Why is this insufficient and what is correct? (Select one)

A. It is sufficient if the instruction is emphatic enough. B. Prompt instructions are probabilistic and injection-vulnerable; enforce the rule with a deterministic hook (PreToolUse on the commit) that blocks with exit code 2 when tests fail. C. Use a larger model so it follows instructions reliably. D. Put the instruction in CLAUDE.md instead of the system prompt.

Answer: B. This is anti-pattern 3 – prompt-based enforcement of a critical rule. Critical rules need deterministic hooks. Emphasis (A), model size (C) and file location (D) do not make a probabilistic control deterministic.

Q6 · One subagent in a five-subagent research task times out. What is the BEST coordinator behaviour? (Select one)

A. Silently drop the failed subagent and present the four results as the complete answer. B. Record a structured error (category timeout, retryable true), retry the retryable task, and if still failing proceed with a quorum while noting the gap or escalate. C. Discard all results and restart the whole task. D. Return a generic ‘research failed’ message.

Answer: B. Structured error plus explicit decision. Dropping it silently (A) is anti-pattern 7. Restarting everything (C) wastes the good results. A generic message (D) is anti-pattern 6.

Q7 · A team must ship an agent quickly, using standard tools, with minimal infrastructure to operate. Which hosting choice fits BEST? (Select one)

A. Claude Agent SDK, self-hosting the loop and sandbox. B. Managed Agents, where Anthropic hosts the loop and sandbox. C. Build a custom orchestration engine from scratch. D. Tool Runner with a fully custom sandbox.

Answer: B. Managed Agents minimise infrastructure and are fastest to production for standard tools. The Agent SDK (A) and custom engine (C) are for teams needing bespoke tools and control. Tool Runner (D) implies you host tool execution.

Q8 · An orchestrator agent must decide subtasks at runtime because they depend on what it finds. Which pattern is this and what is its signature risk? (Select one)

A. Prompt chaining; risk is misclassification. B. Orchestrator-workers; risk includes relying on context inheritance and unhandled partial failures. C. Routing; risk is error compounding. D. Voting; risk is shared bias.

Answer: B. Runtime-decided subtasks define orchestrator-workers. Its signature risks are the coordinator/subagent failure modes. The other options misname the pattern or its risk.

Q9 · A billing agent retries a 'create refund' API call after a 429 and the customer receives two refunds. What went wrong and how should retries be designed? (Select one)

A. The backoff was too short; increase it. B. A non-idempotent write was retried; use an idempotency key so retries do not duplicate the side effect, and only retry with backoff and jitter respecting retry-after. C. The agent should never retry any request. D. Use a bigger model to avoid the error.

Answer: B. Retrying a non-idempotent write duplicates the effect. Idempotency keys make retries safe. Longer backoff (A) does not prevent duplication; never retrying (C) is unnecessary for idempotent calls; model size (D) is irrelevant.

Q10 · An agent must remember a customer's stated preference across separate sessions days apart. Where should this live? (Select one)

A. In the conversation history, which persists automatically. B. In the memory tool or an external store, because conversation history is trimmed, compacted and not retained across sessions. C. In the system prompt, hard-coded per customer. D. In a longer max_tokens setting.

Answer: B. Cross-session persistence requires the memory tool or external state. Conversation history (A) does not survive across sessions or compaction. Hard-coding (C) does not scale; max_tokens (D) is unrelated.

Q11 · An evaluator-optimizer loop reuses the same conversation for both generating and evaluating a translation, and quality plateaus. What is the flaw and fix? (Select one)

A. The generator model is too small; upgrade it. B. Same-session self-review retains the generator’s reasoning bias; run the evaluator as an independent context (fresh session, ideally a different model) against explicit criteria. C. The loop needs more iterations. D. Lower the temperature to reduce variance.

Answer: B. This is anti-pattern 9 – same-session self-review. An independent evaluator removes the shared bias. Model size (A), more iterations (C) and temperature (D) do not fix the bias source.

Q12 · A multi-agent system's operators cannot tell which subagent caused a failed task. Which change addresses this MOST directly? (Select one)

A. Add more retries to every subagent. B. Emit a trace span per agent and thread a correlation ID through the coordinator and all subagents. C. Switch every subagent to Opus 5. D. Increase the iteration cap.

Answer: B. Per-agent traces plus a correlation ID let you reconstruct a single task and locate the failing agent. Retries (A), model choice (C) and caps (D) do not improve observability.

Q13 · Three independent document-summarisation subtasks must complete as fast as possible. Which pattern and why? (Select one)

A. Prompt chaining, because it is simplest. B. Parallelization by sectioning, because the subtasks are independent and latency is bounded by the slowest branch rather than the sum. C. Evaluator-optimizer, because it improves quality. D. Autonomous agent, for flexibility.

Answer: B. Independent subtasks plus a latency goal is textbook sectioning. Chaining (A) runs them serially. Evaluator-optimizer (C) and autonomous agent (D) do not match independent parallel work.

Q14 · Which statements about loop termination are correct? (Select two)

A. stop_reason == "end_turn" is the primary signal that the model is done. B. stop_reason == "max_tokens" means the task completed successfully. C. An iteration cap should be the sole stopping mechanism. D. An iteration or cost cap is a valid backstop alongside stop_reason. E. Parsing the response text for ‘complete’ is the recommended approach.

Answer: A and D. end_turn signals completion; a cap is a legitimate backstop. max_tokens (B) means truncation, a sole cap (C) is anti-pattern 2, and prose parsing (E) is anti-pattern 1.

Q15 · A proposed research system uses a coordinator with eight subagents where a two-step workflow would meet the accuracy bar. Latency and cost are constraints. What is the BEST call? (Select one)

A. Build the eight-subagent system for maximum capability. B. Use the two-step workflow, since it meets the bar at far lower cost and latency; add complexity only if it demonstrably improves outcomes. C. Use a single autonomous agent with eighteen tools. D. Use voting across eight runs of the same prompt.

Answer: B. Simplest solution that meets the requirement. The eight-subagent system (A) multiplies cost/latency without a proven benefit. An 18-tool agent (C) is anti-pattern 8. Eight-way voting (D) is expensive and unjustified.

Q16 · A tool call fails and the tool returns an empty list, which the agent treats as 'no results found' and reports success. What anti-pattern is this and what is correct? (Select one)

A. Anti-pattern 2; add an iteration cap. B. Anti-pattern 7 (silent error suppression); return a structured error with category and retryable flag so the agent distinguishes a real empty result from a failure. C. Anti-pattern 5; stop escalating on sentiment. D. There is no problem; empty results are fine.

Answer: B. Converting a failure into an empty success is silent suppression (anti-pattern 7). The fix is structured error content so ‘failed’ is never mistaken for ‘no results’. The other options misidentify the issue.

Q17 · A hierarchical system uses a root coordinator, two sub-coordinators, and four workers. Findings established at the root are missing three tiers down. What is the root cause and fix? (Select one)

A. The workers need bigger context windows. B. Context is isolated at every tier; each level must pass the relevant context explicitly down to the next — inheritance never happens automatically at any depth. C. The root coordinator should use Opus 5. D. Thinking should be disabled at the worker tier.

Answer: B. Isolation applies at every tier, so context must be threaded explicitly all the way down. Window size (A) and model (C) do not supply context that was never passed; thinking (D) is unrelated.

Q18 · Each call carries a 10k-token stable prefix at $5/M input. Fan-out to five workers plus one synthesis pays the prefix six times. Which change MOST reduces the fan-out's cost penalty? (Select one)

A. Increase max_tokens on each worker. B. Cache the stable prefix so each read costs ~0.1x, collapsing the repeated fixed overhead, and route simple workers to Haiku 4.5. C. Add more workers to parallelise further. D. Switch every worker to Fable 5.1 at $10/$50.

Answer: B. Caching cuts the repeated 10k prefix from ~$0.05 to ~$0.005 per read and cheaper models cut it further. max_tokens (A) raises output cost, more workers (C) pay the overhead more times, and Fable 5.1 (D) is the most expensive model.

Q19 · A payment agent must always require human approval before any transfer over $10,000, even if a tool result claims pre-approval. Which gate is correct? (Select one)

A. Escalate only if the model’s confidence is below 80%. B. A deterministic PreToolUse hook on the transfer tool that inspects the amount and exits 2 above the threshold, routing to a human — regardless of any claimed pre-approval. C. Escalate if the customer sounds anxious. D. A system-prompt rule to always ask before large transfers.

Answer: B. Irreversible high-value actions get a deterministic gate keyed on the amount, immune to injected ‘pre-approval’. Confidence (A) is #4, sentiment (C) is #5, and a prompt rule (D) is #3.

Q20 · A fixed ETL pipeline needs an end-of-run quality gate that can send failures back to the specific stage that produced them. Which variant fits, and what must be bounded? (Select one)

A. Blackboard coordination; bound the number of agents. B. Pipeline with feedback (chaining + evaluator-optimizer); bound the feedback with an objective stop criterion and route failures to the correct stage (using an independent evaluator). C. Autonomous agent; bound the tools. D. Routing; bound the classes.

Answer: B. A fixed flow with a returning quality gate is pipeline-with-feedback; it needs an objective stop and correct-stage routing, with an independent evaluator to avoid #9. The others misname the variant.

Q21 · Three agents write to a shared research document concurrently and updates are being lost; no component decides when the doc is 'done'. What is the BEST fix? (Select two)

A. Serialise all writes through transactional, versioned updates to the shared store. B. Let every agent decide independently when the document is complete. C. Designate one coordinator that owns the ‘done’ decision and aggregates. D. Remove the shared store and give each agent its own copy with no merge. E. Increase each agent’s context window.

Answer: A and C. Transactional/versioned writes prevent lost updates and a single owner of ‘done’ resolves the coordination gap. Independent completion (B) has no owner, un-merged copies (D) lose the shared artefact, and window size (E) is irrelevant.

Q22 · A team debates Managed Agents vs the Agent SDK. The only hard constraint is that tool execution must run inside their own VPC on their hardware; otherwise they want speed. Which is correct? (Select one)

A. Managed Agents, because speed matters most. B. The Agent SDK, self-hosting the loop and sandbox so execution stays in their VPC — the hard constraint overrides the speed preference. C. Either is fine; the constraint is a preference. D. A bespoke engine built from scratch.

Answer: B. A hard compliance constraint (own VPC/hardware) forces self-hosting via the Agent SDK. Managed Agents (A) host the sandbox at Anthropic; the constraint is not a preference (C); a from-scratch engine (D) is needless reinvention.

Key takeaways

  • Prefer the simplest architecture that meets the requirement: single prompt < workflow < agent < multi-agent.
  • Know all six patterns and their failure modes; match the pattern to whether you or the model decides control flow.
  • Drive loop termination from stop_reason; use caps only as a backstop; treat max_tokens as truncation, pause_turn as resume, and refusal as stop-and-escalate.
  • Pass context to subagents explicitly at every tier – isolated contexts never auto-inherit; specify the return shape for clean aggregation.
  • Handle partial failure with structured errors and an explicit decision (quorum, retry, escalate) – never silent drops.
  • Escalate on explicit request (now) or capability gap (after attempting); never on sentiment or self-reported confidence.
  • Enforce critical rules and irreversible-action gates with deterministic hooks/permissions, keyed on blast radius, not prompt instructions.
  • Persist anything durable in the memory tool or external state; retry only idempotent writes with backoff and jitter; and instrument every agent with traces and correlation IDs.
  • Recognise orchestration variants (hierarchical, pipeline-with-feedback, blackboard) and the failure mode each introduces; add tiers only when breadth or quality demonstrably improves.
  • Model cost/latency before building: caching the stable prefix and routing simple subtasks to cheaper models usually beats “bigger model” or “more agents”.
  • Fan-out pays the fixed prefix overhead per worker; caching collapses it, which is why it is the exam-correct cost lever.

Last updated Sep 18, 2026