# D3 · Agents and Workflows

Workflow vs agent decision criteria, the agentic loop, subagents and hierarchies, the Claude Agent SDK, managed vs self-hosted agents, hooks, memory, context management, frameworks, and safe termination.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is roughly **8 of 53 items**. It tests whether you can choose between a fixed workflow and an autonomous agent, build the agentic loop correctly (driven by `stop_reason`, terminated safely), structure subagents with isolated context, and use the Claude Agent SDK and hooks. The recurring theme: **structure and determinism where it matters; autonomy only where it pays.**

## Learning objectives

By the end of this page you should be able to:

1. Decide between a **workflow** and an **agent**, and pick the right Anthropic pattern.
2. Implement the **agentic loop** terminated by `stop_reason`, not iteration caps.
3. Design **manager/supervisor hierarchies** and **subagents** with isolated context.
4. Use the **Claude Agent SDK** (Python + TS) and its key options.
5. Distinguish **Managed Agents** (hosted) from **self-hosted** loops/harnesses.
6. Use **hooks** for deterministic actions, the **memory tool**, and manage the **context window** in long agents.
7. Position frameworks (**LangGraph, PydanticAI, CrewAI**) and apply **safe termination**.

---

## 3.1 Workflow vs agent

- A **workflow** orchestrates Claude and tools through **predefined code paths** – predictable, testable, cheap.
- An **agent** lets Claude **dynamically direct its own process and tool use** – flexible, but less predictable and harder to test.

Prefer the simplest thing that works: **use workflows for well-defined, repeatable tasks; use agents only when the path cannot be predicted in advance.**

| Signal | Choose |
| --- | --- |
| Fixed steps, known inputs/outputs | Workflow |
| Predictable branching | Workflow (routing) |
| Open-ended task, unknown number of steps | Agent |
| Need auditability and low cost | Workflow |
| Task requires the model to decide the plan | Agent |

---

## 3.2 The Anthropic 'Building effective agents' patterns

| Pattern | What it is | Use when |
| --- | --- | --- |
| **Prompt chaining** | Output of step N feeds step N+1 | Task decomposes into fixed sequential subtasks |
| **Routing** | Classify input, dispatch to a specialised path | Distinct categories need different handling |
| **Parallelization** | Run subtasks concurrently, aggregate | Independent subtasks or voting |
| **Orchestrator-workers** | A lead splits work and delegates to workers | Subtasks unknown until runtime |
| **Evaluator-optimizer** | One generates, another critiques and refines | Quality improves with iteration and clear criteria |
| **Autonomous agent** | Model plans and acts in a loop with tools | Open-ended, path not predictable |

:::tip[Exam signal]
"Steps are fixed" → chaining. "Different categories" → routing. "Independent subtasks / vote" → parallelization. "Lead delegates dynamically" → orchestrator-workers. "Generate then critique" → evaluator-optimizer. "Open-ended, decide its own steps" → autonomous agent.
:::

---

## 3.3 The agentic loop

The loop is: call the model → if `stop_reason == 'tool_use'`, run the tools, append `tool_result`, call again → repeat until `end_turn`.

```python
def run_agent(client, messages, tools):
    while True:
        resp = client.messages.create(
            model="claude-opus-5", max_tokens=2048, tools=tools, messages=messages)
        messages.append({"role": "assistant", "content": resp.content})
        if resp.stop_reason == "tool_use":
            tool_results = []
            for block in resp.content:
                if block.type == "tool_use":
                    out = dispatch(block.name, block.input)   # run the tool
                    tool_results.append({"type": "tool_result",
                                         "tool_use_id": block.id, "content": out})
            messages.append({"role": "user", "content": tool_results})
            continue
        return resp    # end_turn / refusal / max_tokens handled by caller
```

```text
┌─ model call ─┐
│              │  stop_reason == tool_use
│   Claude ────┼──────────────► run tool(s) ──► append tool_result ──┐
│              │                                                     │
└──────────────┘◄────────────────────────────────────────────────── ┘
        │ stop_reason == end_turn
        ▼
     return
```

:::danger[Safe termination]
Terminate on **`stop_reason`**, not on an arbitrary iteration count. An iteration cap is a *safety backstop* against runaway loops, never the primary stopping mechanism (anti-patterns #1 and #2). Also handle `max_tokens`, `pause_turn` and `refusal`.
:::

---

## 3.4 Subagents and manager/supervisor hierarchies

A **manager** (orchestrator) agent decomposes a task and delegates to **subagents**, each with its **own isolated context window**, tools and system prompt. Benefits:

- **Context isolation** – a subagent's noisy intermediate work does not pollute the manager's context.
- **Specialisation** – each subagent has a focused tool allowlist and prompt.
- **Parallelism** – independent subagents run concurrently.

```text
        ┌──────────── Manager / Coordinator ────────────┐
        │  plans, delegates, aggregates results         │
        └───┬───────────────┬───────────────┬───────────┘
            ▼               ▼               ▼
       Subagent A      Subagent B      Subagent C
       (own context)   (own context)   (own context)
```

:::tip[Exam signal]
"Intermediate research is bloating the context", "specialised subtasks", "run parts in parallel" → subagents with isolated context under a coordinator.
:::

---

## 3.5 The Claude Agent SDK

`pip install claude-agent-sdk` / `npm install @anthropic-ai/claude-agent-sdk` (renamed from the Claude Code SDK). It provides the agent loop, tool handling and Claude Code's harness.

<Tabs>
  <TabItem label="Python">
```python
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions

async def main():
    options = ClaudeAgentOptions(
        system_prompt="You are a careful coding assistant.",
        allowed_tools=["Read", "Grep", "Edit"],
        permission_mode="acceptEdits",
        mcp_servers={"docs": {"command": "python", "args": ["docs_server.py"]}},
        max_turns=8,
    )
    async for message in query(prompt="Fix the failing test in test_utils.py", options=options):
        print(message)

anyio.run(main)
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
import { query } from '@anthropic-ai/claude-agent-sdk';

for await (const message of query({
  prompt: 'Fix the failing test in test_utils.py',
  options: {
    systemPrompt: 'You are a careful coding assistant.',
    allowedTools: ['Read', 'Grep', 'Edit'],
    permissionMode: 'acceptEdits',
    mcpServers: { docs: { command: 'node', args: ['docs-server.js'] } },
    maxTurns: 8,
  },
})) {
  console.log(message);
}
```
  </TabItem>
</Tabs>

Key options: **`allowed_tools`** (allowlist – least privilege), **`permission_mode`** (`default`/`acceptEdits`/`plan`/`bypassPermissions`), **`system_prompt`**, **`mcp_servers`**, **`hooks`** (deterministic handlers), **`max_turns`** (backstop, not primary termination).

---

## 3.6 Managed vs self-hosted agents

| | Managed Agents | Agent SDK / Tool Runner (self-hosted) |
| --- | --- | --- |
| Who runs the loop | Anthropic hosts loop + sandbox | You host |
| Control over environment | Lower | Full |
| Ops burden | Minimal | You manage infra, sandboxing, scaling |
| Best for | Fast start, standard agentic tasks | Custom tools, custom environments, data control |

:::tip[Exam signal]
"We want Anthropic to run the sandbox and loop" → Managed Agents. "We need our own tools/environment/data controls" → self-hosted via the Agent SDK.
:::

---

## 3.7 Hooks, memory and context management

- **Hooks** are deterministic shell/HTTP handlers fired at lifecycle points (`PreToolUse`, `PostToolUse`, `UserPromptSubmit`, `Stop`, `SessionStart`, `Notification`, `SubagentStop`, `PreCompact`). Exit code **2** blocks the action. Use them to enforce rules that must not depend on the model's cooperation.
- **Memory tool** persists information across sessions (durable notes, learned facts) rather than re-deriving it each run.
- **Context-window management** in long agents: prune old tool results (**context editing**), summarise while preserving narrative (**compaction**), and offload durable state to memory. Without this, long agents blow the window and cost.

```python
# A PreToolUse hook that blocks destructive shell commands (exit 2 = block)
# hook script: reads JSON on stdin, exits 2 to deny
import json, sys
event = json.load(sys.stdin)
cmd = event.get("tool_input", {}).get("command", "")
if "rm -rf" in cmd:
    print("Blocked: destructive command", file=sys.stderr)
    sys.exit(2)
sys.exit(0)
```

:::danger[Deterministic enforcement]
Critical rules (never delete, never spend, always approve) belong in **hooks**, not the prompt (anti-pattern #3). A hook cannot be talked out of blocking; a prompt instruction can.
:::

---

## 3.8 Frameworks (positioning)

| Framework | Position |
| --- | --- |
| **Claude Agent SDK** | Anthropic's first-party harness; tightest Claude Code / tool integration |
| **LangGraph** | Graph-based orchestration of nodes/edges; explicit state machines |
| **PydanticAI** | Type-safe, Pydantic-validated agent outputs in Python |
| **CrewAI** | Multi-agent "crews" with roles and tasks |

Frameworks add orchestration and ergonomics; they do not change the fundamentals – you still drive the loop from `stop_reason`, enforce critical rules deterministically, and manage context.

---

## 3.9 A complete agentic loop dispatching on every `stop_reason`

The sketch in 3.3 only handles `tool_use` and `end_turn`. A production loop must dispatch on **every** `stop_reason`: `tool_use`, `end_turn`, `max_tokens`, `pause_turn`, `refusal` (and `stop_sequence`). The iteration cap is a *backstop*, never the primary stop.

<Tabs>
  <TabItem label="Python">
```python
from anthropic import Anthropic

client = Anthropic()

def run_agent(messages, tools, max_iters=25):
    for _ in range(max_iters):                      # backstop only, not the primary stop
        resp = client.messages.create(
            model="claude-opus-5", max_tokens=4096, tools=tools, messages=messages)
        messages.append({"role": "assistant", "content": resp.content})

        if resp.stop_reason == "tool_use":
            results = []
            for block in resp.content:
                if block.type == "tool_use":
                    out = dispatch(block.name, block.input)   # run the tool
                    results.append({"type": "tool_result",
                                    "tool_use_id": block.id, "content": out})
            messages.append({"role": "user", "content": results})
            continue

        if resp.stop_reason == "pause_turn":
            # long-running server tool paused; resend to resume
            continue

        if resp.stop_reason == "max_tokens":
            # truncated output — ask to continue or raise max_tokens; do not treat as done
            messages.append({"role": "user", "content": "Continue the previous response."})
            continue

        if resp.stop_reason == "refusal":
            # model declined for safety; surface to a human, do not retry blindly
            return {"status": "refused", "response": resp}

        if resp.stop_reason in ("end_turn", "stop_sequence"):
            return {"status": "done", "response": resp}

    raise RuntimeError("iteration backstop hit — investigate; loop did not reach end_turn")
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

async function runAgent(messages: any[], tools: any[], maxIters = 25) {
  for (let i = 0; i < maxIters; i++) {          // backstop only, not the primary stop
    const resp = await client.messages.create({
      model: 'claude-opus-5', max_tokens: 4096, tools, messages,
    });
    messages.push({ role: 'assistant', content: resp.content });

    if (resp.stop_reason === 'tool_use') {
      const results = [];
      for (const block of resp.content) {
        if (block.type === 'tool_use') {
          const out = await dispatch(block.name, block.input);   // run the tool
          results.push({ type: 'tool_result', tool_use_id: block.id, content: out });
        }
      }
      messages.push({ role: 'user', content: results });
      continue;
    }
    if (resp.stop_reason === 'pause_turn') continue;             // resume long-running server tool
    if (resp.stop_reason === 'max_tokens') {
      messages.push({ role: 'user', content: 'Continue the previous response.' });
      continue;                                                  // truncated, not done
    }
    if (resp.stop_reason === 'refusal') return { status: 'refused', response: resp };
    if (resp.stop_reason === 'end_turn' || resp.stop_reason === 'stop_sequence') {
      return { status: 'done', response: resp };
    }
  }
  throw new Error('iteration backstop hit — loop did not reach end_turn');
}
```
  </TabItem>
</Tabs>

:::danger[Dispatch on the signal, not the prose]
Treating `max_tokens` as completion silently truncates work; treating `refusal` as an error to retry blindly ignores a safety signal; ignoring `pause_turn` abandons a long-running server tool. Parsing the assistant text for 'I am done' is anti-pattern #1. Always branch on `stop_reason`.
:::

---

## 3.10 Orchestrator-workers with subagent context passing

In an orchestrator-workers pattern the lead agent decomposes the task, spawns workers each with an **isolated context window**, passes each worker only the slice it needs, and aggregates the distilled results back — the workers' verbose intermediate work never enters the orchestrator's context.

```python
from anthropic import Anthropic

client = Anthropic()

def worker(task_brief: str, documents: str) -> str:
    """A subagent with its OWN fresh context — only the brief + its slice, nothing else."""
    resp = client.messages.create(
        model="claude-sonnet-5", max_tokens=1024,
        system="You are a focused research worker. Return only a distilled 5-line summary.",
        messages=[{"role": "user", "content": f"<task>{task_brief}</task>\n<docs>{documents}</docs>"}])
    return "".join(b.text for b in resp.content if b.type == "text")

def orchestrator(question: str, corpus: dict[str, str]) -> str:
    # 1. Lead decomposes and delegates; each worker gets only its slice (context passing).
    summaries = []
    for topic, docs in corpus.items():
        brief = f"Summarise what the docs say about: {question} (focus: {topic})"
        summaries.append(f"[{topic}] {worker(brief, docs)}")     # only distilled result returns

    # 2. Lead aggregates distilled summaries — worker noise never entered this context.
    synthesis = client.messages.create(
        model="claude-opus-5", max_tokens=2048,
        system="You are the coordinator. Synthesise the worker summaries into one answer.",
        messages=[{"role": "user", "content": f"Question: {question}\n\n" + "\n".join(summaries)}])
    return "".join(b.text for b in synthesis.content if b.type == "text")
```

```text
        ┌──────── Orchestrator (Opus 5) ────────┐
        │ decompose → delegate slices → synthesise│
        └───┬───────────┬───────────┬────────────┘
   brief+slice     brief+slice   brief+slice     (context passing: only the needed slice)
            ▼           ▼           ▼
        Worker A     Worker B     Worker C        (each fresh, isolated context)
            │           │           │
        5-line      5-line       5-line           (only distilled results return)
         summary     summary      summary
```

:::tip[Exam signal]
'A lead splits work it cannot enumerate up front and delegates' → orchestrator-workers. 'Pass each subagent only the slice it needs and return distilled summaries' → context passing + isolation, which keeps the coordinator's window clean and cheap.
:::

---

## 3.11 Hooks in the Agent SDK: a PreToolUse example

Hooks are deterministic handlers fired at lifecycle points; **exit code 2 blocks the action**. Critical rules belong here, not in the prompt (anti-pattern #3). Below, a `PreToolUse` hook denies any `Bash` command containing a destructive pattern before it ever runs.

<Tabs>
  <TabItem label="Python (Agent SDK)">
```python
import anyio
from claude_agent_sdk import query, ClaudeAgentOptions

async def block_destructive(input_data, tool_use_id, context):
    """PreToolUse hook: return a deny decision to block the tool call."""
    cmd = input_data.get("tool_input", {}).get("command", "")
    if any(p in cmd for p in ("rm -rf", "DROP TABLE", "git push --force")):
        return {
            "hookSpecificOutput": {
                "hookEventName": "PreToolUse",
                "permissionDecision": "deny",
                "permissionDecisionReason": "Blocked destructive command by policy.",
            }
        }
    return {}

async def main():
    options = ClaudeAgentOptions(
        allowed_tools=["Read", "Grep", "Bash"],
        permission_mode="default",
        hooks={"PreToolUse": [block_destructive]},
        max_turns=8,
    )
    async for message in query(prompt="Clean up the build artifacts", options=options):
        print(message)

anyio.run(main)
```
  </TabItem>
  <TabItem label="Shell hook (exit 2 blocks)">
```python
# .claude/hooks/pretooluse.py — reads the event JSON on stdin; exit 2 denies.
import json, sys

event = json.load(sys.stdin)
cmd = event.get("tool_input", {}).get("command", "")
if any(p in cmd for p in ("rm -rf", "DROP TABLE", "git push --force")):
    print("Blocked destructive command by policy.", file=sys.stderr)
    sys.exit(2)          # exit code 2 → the tool call is blocked
sys.exit(0)
```
  </TabItem>
</Tabs>

:::danger[Deterministic enforcement, not prompt pleading]
A hook cannot be argued out of blocking; a system-prompt rule can be overridden by a clever tool result or injected instruction. Any rule that is irreversible or high-stakes (delete, spend, publish, push) must be a hook or programmatic check (anti-pattern #3).
:::

---

## 3.12 Managed Agents vs Agent SDK vs custom loop

| Dimension | Managed Agents | Claude Agent SDK (self-hosted) | Custom Messages API loop |
| --- | --- | --- | --- |
| Who runs the loop + sandbox | Anthropic hosts both | You host; SDK provides the harness | You build and host everything |
| Setup effort | Lowest | Moderate | Highest |
| Control over environment / tools | Lowest | High (own tools, MCP, hooks, permission modes) | Total |
| Built-in hooks, permissions, subagents | Yes (managed) | Yes (SDK primitives) | You implement it all |
| Data / infra control | Anthropic-managed | Your infra, your data controls | Your infra, your data controls |
| Best for | Fast start on standard agentic tasks | Custom tools/environments needing Claude Code harness ergonomics | Full control, unusual integrations, or minimal dependencies |
| Termination / safety burden | Managed | You wire hooks + `stop_reason` + `max_turns` | You implement `stop_reason` dispatch and backstops yourself |

:::tip[Exam signal]
'Anthropic should host the loop and sandbox, start fast' → Managed Agents. 'We need our own tools/environment/data controls with a first-party harness, hooks and permission modes' → Agent SDK. 'Minimal dependencies / unusual integration / full control' → custom Messages API loop (and you must implement `stop_reason` dispatch and a backstop yourself).
:::

---

## 3.13 The evaluator-optimizer pattern in code

Evaluator-optimizer improves quality by iterating: a **generator** produces a draft, a **separate** evaluator critiques it against explicit criteria, and the generator revises — repeating until the criteria are met or a backstop trips. The evaluator must not share the generator's session (anti-pattern #9).

```python
from anthropic import Anthropic

client = Anthropic()

def generate(brief, feedback=None):
    prompt = brief if not feedback else f"{brief}\n\nRevise to fix: {feedback}"
    r = client.messages.create(model="claude-sonnet-5", max_tokens=1024,
                               messages=[{"role": "user", "content": prompt}])
    return "".join(b.text for b in r.content if b.type == "text")

def evaluate(brief, draft):
    # Separate session AND ideally a different model — no shared reasoning bias.
    r = client.messages.create(
        model="claude-opus-5", max_tokens=512,
        system="Score the draft against the brief. Return JSON {'pass': bool, 'feedback': str}.",
        messages=[{"role": "user", "content": f"BRIEF:\n{brief}\n\nDRAFT:\n{draft}"}])
    import json
    return json.loads("".join(b.text for b in r.content if b.type == "text"))

def optimize(brief, max_rounds=3):          # backstop, not the primary stop
    draft = generate(brief)
    for _ in range(max_rounds):
        verdict = evaluate(brief, draft)
        if verdict["pass"]:
            return draft                     # primary stop: criteria met
        draft = generate(brief, verdict["feedback"])
    return draft                             # backstop reached — flag for review
```

| Element | Requirement |
| --- | --- |
| Generator and evaluator | **Different sessions**; ideally different models |
| Stopping | Primary: criteria met; backstop: max rounds |
| Feedback | Concrete and actionable, fed back into the next draft |
| Use when | Quality improves with iteration and criteria are explicit |

:::tip[Exam signal]
'Generate, then critique and refine against clear criteria' → evaluator-optimizer. If the *same* session grades its own draft, that is anti-pattern #9 — the judge inherits the generator's bias.
:::

---

## 3.14 Positioning frameworks against the fundamentals

Frameworks add orchestration ergonomics; they never remove the need for `stop_reason`-driven termination, deterministic enforcement, and context management.

| Framework | Sweet spot | What it does *not* change |
| --- | --- | --- |
| **Claude Agent SDK** | First-party harness; tightest Claude Code / tool / hook integration | You still wire allowlists, hooks, `stop_reason` dispatch, backstop |
| **LangGraph** | Explicit graph/state-machine orchestration of nodes and edges | Still drive the model call from `stop_reason` inside each node |
| **PydanticAI** | Type-safe, Pydantic-validated agent outputs in Python | Validation-retry and schema design still apply |
| **CrewAI** | Multi-agent 'crews' with roles and tasks | Least privilege, context isolation, and safe termination still apply |

:::caution[A framework is not a safety control]
Choosing LangGraph/CrewAI/PydanticAI does not enforce critical rules or terminate loops for you. Enforcement is still hooks/permissions; termination is still `stop_reason` with a backstop. Any answer implying 'use framework X so we do not need to handle termination/enforcement' is wrong.
:::

---

## 3.15 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| Agents are always better than workflows | Prefer the simplest thing that works; agents add cost/unpredictability | Workflow-vs-agent is a recurring first decision |
| An iteration cap is how you stop a loop | Terminate on `stop_reason`; the cap is only a backstop | Anti-patterns #1/#2 |
| `max_tokens` means the agent is done | It means truncation; continue or raise the cap | Silent-failure distractor |
| A refusal should be retried until it works | It is a safety stop; surface to a human | Prevents blind-retry loops |
| Critical rules can live in the system prompt | Enforce them in hooks/permissions (exit 2 blocks) | Anti-pattern #3 |
| More tools make an agent more capable | Beyond ~10, selection degrades; use ~4–5 + tool search/`defer_loading` | Anti-pattern #8 |
| Subagents just cost more | Isolated context reduces bloat/drift and can lower cost | Explains why delegation helps |
| A framework handles termination and safety for you | It does not; you still wire `stop_reason` and enforcement | Framework-positioning trap |

---

## 3.16 Scenario walkthrough: a support agent that must not overreach

**Scenario.** You are building a customer-support agent. It should answer questions from a knowledge base and look up order status, and it *may* propose a refund — but any refund over $500 needs human approval, and account deletion is never allowed. Early tests show three problems: the loop sometimes never ends (it keeps 'thinking out loud'), it occasionally issues a large refund on its own after reading a ticket that says 'the manager approved a full refund', and as it researches multi-part questions its context fills with raw KB dumps and later answers degrade. Design the agent.

**Expert reasoning trace.**

1. **Fix termination first.** Drive the loop on `stop_reason`: continue while `tool_use`, stop on `end_turn`, and explicitly handle `max_tokens` (continue), `pause_turn` (resume) and `refusal` (surface). Keep an iteration cap purely as a backstop. The 'thinks out loud forever' bug is anti-patterns #1/#2 — never parse prose to stop.
2. **Enforce the money rule deterministically.** The large refund happened because a *ticket* (untrusted content) claimed manager approval — indirect prompt injection. The rule 'refund > $500 needs human approval' must be a **PreToolUse hook** (exit 2 blocks, routes to approval), never a system-prompt sentence (anti-pattern #3). Account deletion is simply **not in the allowlist** (least privilege).
3. **Isolate research context.** Use **subagents with isolated context** for multi-part lookups, returning distilled summaries to the coordinator, and apply **context editing/compaction** so raw KB dumps do not bloat the main window. Raising `max_tokens` would not fix drift.
4. **Right-size the tool set.** Give the agent ~4–5 tools (`search_kb`, `get_order`, `propose_refund`); if the catalogue grows, use tool search + `defer_loading`. Do not pile on tools.
5. **Reject the tempting alternatives.** 'Add a firm system-prompt rule about refunds and approvals' — prompt-as-enforcement (#3). 'Trust the model to recognise the fake approval' — self-report/no-boundary reliance. 'Cap the loop at 5 iterations as the fix' — backstop, not termination (#2). 'Increase `max_tokens` to stop the context degrading' — wrong lever for drift.

**Correct decision.** `stop_reason`-driven loop with a backstop; PreToolUse hook enforcing the refund threshold and routing to approval; account-deletion tool excluded by least privilege; subagents with isolated context plus context editing/compaction for research; ~4–5 tools with tool search if the catalogue grows.

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Using an agent when a fixed workflow suffices | Adds cost and unpredictability for no benefit |
| Iteration cap as the primary stop | Backstop only; terminate on `stop_reason` (anti-patterns #1, #2) |
| Enforcing critical rules in the prompt | Use hooks / programmatic checks (anti-pattern #3) |
| One agent with 18 tools | Overloads the model; split into subagents / use tool search + `defer_loading` (anti-pattern #8) |
| Letting subagent work bloat the manager's context | Use isolated-context subagents |
| Same-session self-review of an agent's own work | Retains reasoning bias (anti-pattern #9); use a fresh evaluator |
| Parsing prose to decide delegation/termination | Drive from structured signals |
| Assuming a framework removes the need for safe termination | It does not |
| Treating `max_tokens` as task completion | Output was truncated; ask to continue or raise the cap, do not mark done |
| Retrying a `refusal` blindly | It is a safety signal; surface to a human, do not loop on it |
| Ignoring `pause_turn` | Abandons a long-running server tool; resend to resume |
| Passing full worker context back to the orchestrator | Bloats the coordinator window; return only distilled results |
| Having the same session grade the agent's own draft in evaluator-optimizer | Judge inherits the generator's bias (anti-pattern #9); use a separate session/model |
| Assuming a framework (LangGraph/CrewAI/PydanticAI) handles termination or enforcement | It does not; you still wire `stop_reason` dispatch and hooks/permissions |
| Acting on an 'approval' claimed inside untrusted ticket/tool content | Indirect injection; enforce approval in a hook, not by trusting the content |
| Choosing an agent when fixed ordered steps exist | Prompt chaining is cheaper, testable and predictable |
| Using an iteration cap as the fix for a runaway loop | The cap is a backstop; the fix is `stop_reason`-driven termination |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A task has three fixed, ordered steps with known inputs and outputs. What is the best design? (Select one)">
    A. An autonomous agent that decides the steps.
    B. A prompt-chaining workflow where each step feeds the next.
    C. A single mega-prompt.
    D. An orchestrator with dynamic subagents.

    **Answer: B.** Fixed sequential subtasks are the definition of prompt chaining – predictable, testable, cheap. Autonomy (A, D) adds unpredictability with no benefit; a mega-prompt (C) is harder to control.
  </AccordionItem>

  <AccordionItem title="Q2 · An agent loop sometimes never terminates. What is the correct primary termination mechanism? (Select one)">
    A. Stop after 5 iterations regardless.
    B. Stop when the assistant text says it is finished.
    C. Continue while `stop_reason == 'tool_use'` and stop on `end_turn`, with an iteration cap only as a safety backstop.
    D. Stop when output length exceeds a threshold.

    **Answer: C.** Terminate on `stop_reason`; the cap is a backstop, not the primary mechanism (anti-patterns #1, #2). Prose (B) and length (D) are not reliable signals.
  </AccordionItem>

  <AccordionItem title="Q3 · A research agent's context is filling with verbose intermediate tool output, degrading later steps. What is the best fix? (Select one)">
    A. Increase `max_tokens`.
    B. Delegate research to subagents with isolated context that return only distilled results to the coordinator.
    C. Lower temperature.
    D. Remove all tools.

    **Answer: B.** Isolated-context subagents keep noisy intermediate work out of the coordinator's window and return summaries. `max_tokens` (A) caps output; temperature (C) and removing tools (D) do not address context bloat.
  </AccordionItem>

  <AccordionItem title="Q4 · A business rule states the agent must never execute a refund over $500 without human approval. Where should this be enforced? (Select one)">
    A. In the system prompt as an instruction.
    B. In a PreToolUse hook that blocks (exit code 2) refunds over the threshold and routes to human approval.
    C. By asking the model to double-check.
    D. By lowering the model's temperature.

    **Answer: B.** Critical, irreversible rules must be enforced deterministically via hooks, not prompt instructions (anti-pattern #3). Prompt-based enforcement and self-checks can be bypassed.
  </AccordionItem>

  <AccordionItem title="Q5 · A team wants Anthropic to host the agent loop and sandbox so they can start fast with standard tools. Which option fits? (Select one)">
    A. Self-hosted Agent SDK on their own infrastructure.
    B. Managed Agents.
    C. A raw Messages API loop with no tools.
    D. A cron job.

    **Answer: B.** Managed Agents have Anthropic host the loop and sandbox. Self-hosting (A) is the opposite; a bare loop (C) or cron (D) does not provide a managed sandbox.
  </AccordionItem>

  <AccordionItem title="Q6 · Which TWO Agent SDK options directly support least-privilege and safety? (Select two)">
    A. `allowed_tools` restricting the tool allowlist.
    B. `max_turns` set to 1000.
    C. `hooks` that block dangerous actions.
    D. `system_prompt` length.
    E. `permission_mode: 'bypassPermissions'`.

    **Answer: A and C.** An explicit tool allowlist and blocking hooks enforce least privilege and safety. A huge `max_turns` (B) weakens the backstop; prompt length (D) is irrelevant; `bypassPermissions` (E) removes safeguards.
  </AccordionItem>

  <AccordionItem title="Q7 · One agent is configured with 18 tools and frequently picks the wrong one. What is the recommended remedy? (Select one)">
    A. Add more tools to cover edge cases.
    B. Reduce to a focused set (≈4–5), split responsibilities into subagents, and use tool search with `defer_loading` for large catalogues.
    C. Raise `temperature`.
    D. Increase `max_turns`.

    **Answer: B.** Too many tools per agent is anti-pattern #8; the fix is fewer tools, subagent specialisation, and tool search + `defer_loading` beyond ~10 tools. More tools (A) worsens it.
  </AccordionItem>

  <AccordionItem title="Q8 · An agent must remember user preferences across separate sessions. Which mechanism is designed for this? (Select one)">
    A. Increasing the context window.
    B. The memory tool for cross-session persistence.
    C. Higher effort thinking.
    D. Re-sending the full transcript every time forever.

    **Answer: B.** The memory tool persists durable information across sessions. A bigger window (A) does not persist between sessions; effort (C) is unrelated; re-sending everything (D) is costly and unbounded.
  </AccordionItem>

  <AccordionItem title="Q9 · An agent loop returns `stop_reason == 'max_tokens'` mid-way through a long answer, and the harness treats that as completion. What is the correct handling? (Select one)">
    A. Treat `max_tokens` as done; the answer is complete.
    B. Recognise the output was truncated: ask the model to continue (or raise `max_tokens`) rather than marking the turn complete.
    C. Retry the whole request at `temperature: 0`.
    D. Switch to Haiku 4.5.

    **Answer: B.** `max_tokens` means the output hit the output cap and was cut off — it is not `end_turn`. The loop should continue the response or raise the cap. Treating it as done (A) silently truncates work; regenerating (C) or switching model (D) does not recover the truncated content.
  </AccordionItem>

  <AccordionItem title="Q10 · While running, an agent receives `stop_reason == 'pause_turn'`. What does this indicate and what should the loop do? (Select one)">
    A. The model refused; stop and alert a human.
    B. A long-running server-side tool paused; resend the conversation to resume it.
    C. The iteration cap was hit; abort.
    D. The output was truncated; raise `max_tokens`.

    **Answer: B.** `pause_turn` signals a long-running (server) turn paused; the loop resends to resume. Refusal (A) is a different stop reason; the cap (C) is a harness backstop, not a `stop_reason`; truncation (D) is `max_tokens`.
  </AccordionItem>

  <AccordionItem title="Q11 · A research coordinator delegates topics to worker subagents. Which TWO design choices keep the coordinator's context clean and costs down? (Select two)">
    A. Give each worker its own isolated context window with only its task slice.
    B. Return each worker's full transcript, including intermediate tool output, to the coordinator.
    C. Have workers return only a short distilled summary to the coordinator.
    D. Run all work in the coordinator's single context.
    E. Give every worker all 18 tools.

    **Answer: A and C.** Isolated worker contexts plus distilled-summary returns keep the coordinator's window free of noisy intermediate work and reduce token cost. Returning full transcripts (B) or using one shared context (D) causes bloat; giving every worker 18 tools (E) is anti-pattern #8.
  </AccordionItem>

  <AccordionItem title="Q12 · A team must guarantee the agent never runs `git push --force`, regardless of what the model decides. Where and how should this be enforced in the Agent SDK? (Select one)">
    A. A strongly worded instruction in the system prompt.
    B. A `PreToolUse` hook that returns a deny decision (or a shell hook exiting with code 2) when the command matches the pattern.
    C. Ask the model to confirm before force-pushing.
    D. Lower the temperature so it behaves.

    **Answer: B.** Irreversible, critical rules must be enforced deterministically with a `PreToolUse` hook (deny / exit code 2), which the model cannot argue around (anti-pattern #3). Prompt instructions (A), model self-confirmation (C), and temperature (D) can all be bypassed.
  </AccordionItem>

  <AccordionItem title="Q13 · An organisation wants the tightest control over its own tools, sandboxing and data, using a first-party harness with hooks and permission modes. Which option fits, and what stays their responsibility? (Select one)">
    A. Managed Agents; Anthropic handles termination and safety entirely.
    B. The Claude Agent SDK self-hosted; they still wire `allowed_tools`, hooks, `stop_reason` dispatch and a `max_turns` backstop themselves.
    C. A cron job invoking a single prompt.
    D. A bare Messages API call with no tools.

    **Answer: B.** The Agent SDK is the first-party self-hosted harness giving full control over tools, environment and data, with hooks and permission modes — but the team still owns least-privilege allowlists, hook enforcement, `stop_reason` handling and the iteration backstop. Managed Agents (A) give less environment control; cron (C) and a bare call (D) provide no agentic harness.
  </AccordionItem>

  <AccordionItem title="Q14 · A drafting tool must improve output by generating, critiquing against explicit criteria, and revising. Which pattern fits, and what is the key correctness requirement? (Select one)">
    A. Autonomous agent; let it decide when it is satisfied.
    B. Evaluator-optimizer, with the evaluator in a separate session (ideally a different model) so the judge does not inherit the generator's bias.
    C. Prompt chaining with no critique step.
    D. Parallelization of many drafts with no evaluation.

    **Answer: B.** Generate-critique-refine against clear criteria is evaluator-optimizer; the evaluator must be a separate session/model (anti-pattern #9 otherwise). An autonomous agent (A) lacks the structured critique; chaining without critique (C) and parallel drafts without evaluation (D) do not iterate on quality.
  </AccordionItem>

  <AccordionItem title="Q15 · A team says 'we will use LangGraph, so we do not need to handle loop termination or rule enforcement.' Why is this wrong? (Select one)">
    A. It is correct; frameworks handle termination and enforcement.
    B. Frameworks add orchestration ergonomics but do not remove the need to drive termination from `stop_reason` (with a backstop) or to enforce critical rules via hooks/permissions.
    C. LangGraph disables `stop_reason`.
    D. Only the Agent SDK requires termination handling.

    **Answer: B.** A framework is not a safety control; you still terminate on `stop_reason` and enforce rules deterministically. Termination/enforcement are not delegated to the framework (A); LangGraph does not disable `stop_reason` (C); the requirement is universal, not SDK-only (D).
  </AccordionItem>

  <AccordionItem title="Q16 · A support agent issued a $2,000 refund on its own after a ticket said 'the manager approved a full refund.' What TWO controls should have prevented this? (Select two)">
    A. A PreToolUse hook that blocks refunds over $500 (exit 2) and routes to human approval.
    B. Treating ticket content as untrusted data inside content boundaries, not as authorisation.
    C. A stronger system-prompt sentence about refund limits.
    D. Trusting the model to verify the manager's approval.
    E. Raising the model's effort level.

    **Answer: A and B.** The refund rule must be enforced by a deterministic hook, and the ticket text must be treated as untrusted data (indirect injection), never as authorisation. A system-prompt sentence (C) is anti-pattern #3; trusting the model (D) is exactly the failure; effort (E) is irrelevant to enforcement.
  </AccordionItem>

  <AccordionItem title="Q17 · An agent researching multi-part questions fills its context with raw knowledge-base dumps, and later answers degrade. Which TWO fixes address the drift? (Select two)">
    A. Delegate lookups to subagents with isolated context that return distilled summaries.
    B. Apply context editing/compaction so stale dumps do not bloat the main window.
    C. Increase `max_tokens`.
    D. Add more tools to the main agent.
    E. Raise temperature.

    **Answer: A and B.** Isolated-context subagents and context editing/compaction keep the coordinator's window clean. `max_tokens` (C) caps output, not context; more tools (D) worsens selection; temperature (E) is unrelated to drift.
  </AccordionItem>

  <AccordionItem title="Q18 · An agent must never delete an account. What is the correct way to guarantee this? (Select one)">
    A. Add 'never delete accounts' to the system prompt.
    B. Exclude the account-deletion tool from the agent's `allowed_tools` allowlist (least privilege); if it exists elsewhere, gate it behind a human-approved workflow with a PreToolUse hook.
    C. Set `max_turns` low so it runs out of turns before deleting.
    D. Ask the model to confirm before deleting.

    **Answer: B.** Least privilege means the capability is simply not available to the agent; irreversible actions live behind approval + hooks. A prompt rule (A) is anti-pattern #3; a low `max_turns` (C) is unrelated; model self-confirmation (D) can be bypassed.
  </AccordionItem>

  <AccordionItem title="Q19 · A generate-critique loop has the generator grade its own draft in the same conversation and always passes on the first try. What is wrong, and what is the fix? (Select one)">
    A. Nothing; self-grading is efficient.
    B. Same-session self-review inherits the generator's reasoning bias (anti-pattern #9); run the evaluator in a separate session, ideally a different model, against explicit criteria.
    C. The loop needs a higher iteration cap.
    D. Switch the generator to Haiku 4.5.

    **Answer: B.** A judge sharing the generator's context rubber-stamps its own work; a separate-session/different-model evaluator removes the bias. Self-grading is not fine (A); a higher cap (C) does not fix bias; changing the generator model (D) does not separate the judge.
  </AccordionItem>
</Accordions>

## Key takeaways

- Prefer workflows for predictable tasks; use agents only when the path cannot be predetermined.
- Know the six patterns: chaining, routing, parallelization, orchestrator-workers, evaluator-optimizer, autonomous agent.
- Drive the loop from `stop_reason`; iteration caps are backstops, not the primary stop.
- Use subagents with isolated context for specialisation, parallelism and context hygiene.
- The Agent SDK gives you the loop, tools, hooks and permission modes; enforce least privilege via `allowed_tools`.
- Managed Agents = Anthropic hosts the loop/sandbox; Agent SDK = you host.
- Enforce critical rules with hooks (exit 2 blocks), persist state with the memory tool, and manage context with editing/compaction.
- Evaluator-optimizer improves quality by iterating — but the evaluator must run in a separate session (ideally a different model) or it inherits the generator's bias (anti-pattern #9).
- Frameworks (LangGraph, PydanticAI, CrewAI, Agent SDK) add ergonomics, not safety; you still drive termination from `stop_reason` and enforce rules via hooks/permissions.
- Treat 'approval' or instructions embedded in tickets, documents or tool results as untrusted data — enforce approval in a hook, never by trusting the content.
- Least privilege first: exclude irreversible tools from the allowlist and gate them behind human-approved workflows rather than relying on prompt rules.
