# Claude API Cheat Sheet

Messages API request and response anatomy, streaming events, tool use, structured outputs, thinking, caching, batches, errors and retries – with Python and TypeScript.

import { Tabs, TabItem, Steps } from '@prosefly/astro-components';

## Request anatomy

```json
{
  "model": "claude-opus-5",
  "max_tokens": 2048,
  "system": "You are a precise assistant. Answer only from <document>.",
  "messages": [
    { "role": "user", "content": [
      { "type": "text", "text": "<document>…</document>", "cache_control": { "type": "ephemeral" } },
      { "type": "text", "text": "Summarise the termination clause." }
    ]}
  ],
  "temperature": 0.2,
  "stop_sequences": ["</answer>"],
  "thinking": { "type": "adaptive" },
  "effort": "high",
  "tools": [],
  "tool_choice": { "type": "auto" },
  "metadata": { "user_id": "hashed-id" }
}
```

| Field | Notes |
| --- | --- |
| `model` | Pinned ID: `claude-fable-5-1`, `claude-opus-5`, `claude-sonnet-5`, `claude-haiku-4-5` |
| `max_tokens` | Hard output cap; `stop_reason: max_tokens` = truncated |
| `system` | Top-level string or content blocks; stable → cacheable |
| `messages` | Alternating `user`/`assistant`; content is string or block array (`text`, `image`, `document`, `tool_use`, `tool_result`, `thinking`) |
| `temperature` / `top_p` | Set one, not both; 0 reduces variance, not error |
| `thinking` | `{"type":"adaptive"}` (current models); `{"type":"enabled","budget_tokens":N}` only Haiku 4.5 |
| `effort` | `low` / `medium` / `high` / `xhigh` |
| `tools` / `tool_choice` | See tool use below; forced choice is 400 on Fable 5.1 |
| `output_config.format` | JSON Schema for structured output |

## Response anatomy

```json
{
  "id": "msg_01…",
  "type": "message",
  "role": "assistant",
  "model": "claude-opus-5",
  "content": [
    { "type": "thinking", "thinking": "…", "signature": "…" },
    { "type": "text", "text": "The notice period is 60 days." }
  ],
  "stop_reason": "end_turn",
  "stop_sequence": null,
  "usage": {
    "input_tokens": 1200,
    "cache_creation_input_tokens": 0,
    "cache_read_input_tokens": 18000,
    "output_tokens": 42
  }
}
```

### stop_reason – branch on it every time

| Value | Meaning | Action |
| --- | --- | --- |
| `end_turn` | Finished naturally | Done |
| `tool_use` | Wants tools run | Execute, append `tool_result`, call again |
| `max_tokens` | Truncated | Continue or raise cap; never treat as complete |
| `stop_sequence` | Hit a stop string | Done (check `stop_sequence`) |
| `pause_turn` | Long server-side tool turn paused | Re-send to continue |
| `refusal` | Safety decline | Explicit fallback path; do not retry blindly |

## Minimal calls

<Tabs>
  <TabItem label="Python">
```python
from anthropic import Anthropic

client = Anthropic()  # reads ANTHROPIC_API_KEY

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=1024,
    system="You are a concise analyst.",
    messages=[{"role": "user", "content": "Three risks of vendor lock-in?"}],
)
print(msg.content[0].text, msg.stop_reason, msg.usage)
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
import Anthropic from '@anthropic-ai/sdk';

const client = new Anthropic();

const msg = await client.messages.create({
  model: 'claude-sonnet-5',
  max_tokens: 1024,
  system: 'You are a concise analyst.',
  messages: [{ role: 'user', content: 'Three risks of vendor lock-in?' }],
});
console.log(msg.content[0].type === 'text' ? msg.content[0].text : '', msg.stop_reason);
```
  </TabItem>
</Tabs>

## Streaming (SSE)

Event order: `message_start` → (`content_block_start` → `content_block_delta`* → `content_block_stop`)* → `message_delta` (carries `stop_reason`, output usage) → `message_stop`.

<Tabs>
  <TabItem label="Python">
```python
with client.messages.stream(
    model="claude-sonnet-5", max_tokens=1024,
    messages=[{"role": "user", "content": "Write a haiku about latency."}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
    print(final.stop_reason, final.usage)
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
const stream = client.messages.stream({
  model: 'claude-sonnet-5', max_tokens: 1024,
  messages: [{ role: 'user', content: 'Write a haiku about latency.' }],
});
stream.on('text', (t) => process.stdout.write(t));
const final = await stream.finalMessage();
console.log(final.stop_reason, final.usage);
```
  </TabItem>
</Tabs>

## Tool use round-trip

```json
// 1. Request with tools
{ "model": "claude-opus-5", "max_tokens": 1024,
  "tools": [{
    "name": "get_order",
    "description": "Look up one order by ID. Use when the user references an order number. Returns status, items and total. Does not modify anything.",
    "input_schema": { "type": "object",
      "properties": { "order_id": { "type": "string", "description": "Order ID, e.g. ORD-12345" } },
      "required": ["order_id"] },
    "strict": true
  }],
  "messages": [{ "role": "user", "content": "Where is ORD-12345?" }] }

// 2. Response: stop_reason = "tool_use"
{ "content": [{ "type": "tool_use", "id": "toolu_01", "name": "get_order", "input": { "order_id": "ORD-12345" } }],
  "stop_reason": "tool_use" }

// 3. Follow-up with tool_result (in a USER message)
{ "messages": [
  { "role": "user", "content": "Where is ORD-12345?" },
  { "role": "assistant", "content": [{ "type": "tool_use", "id": "toolu_01", "name": "get_order", "input": { "order_id": "ORD-12345" } }] },
  { "role": "user", "content": [{ "type": "tool_result", "tool_use_id": "toolu_01",
      "content": "{\"status\":\"shipped\",\"eta\":\"2026-09-17\"}" }] }
] }

// Error result – structured, never empty-success
{ "type": "tool_result", "tool_use_id": "toolu_01", "is_error": true,
  "content": "{\"category\":\"not_found\",\"retryable\":false,\"message\":\"No order ORD-12345\"}" }
```

### The agentic loop (Python)

```python
def run(messages, tools, model="claude-opus-5"):
    while True:
        r = client.messages.create(model=model, max_tokens=4096, tools=tools, messages=messages)
        messages.append({"role": "assistant", "content": r.content})
        if r.stop_reason == "tool_use":
            results = []
            for block in r.content:
                if block.type == "tool_use":
                    try:
                        out = TOOLS[block.name](**block.input)
                        results.append({"type": "tool_result", "tool_use_id": block.id, "content": json.dumps(out)})
                    except ToolError as e:
                        results.append({"type": "tool_result", "tool_use_id": block.id, "is_error": True,
                                        "content": json.dumps({"category": e.category, "retryable": e.retryable, "message": str(e)})})
            messages.append({"role": "user", "content": results})
            continue
        if r.stop_reason == "max_tokens":
            messages.append({"role": "user", "content": "Continue."}); continue
        if r.stop_reason == "pause_turn":
            continue
        if r.stop_reason == "refusal":
            return handle_refusal(r)
        return r  # end_turn / stop_sequence
```

### tool_choice

| Value | Behaviour | Fable 5.1 |
| --- | --- | --- |
| `{"type":"auto"}` | Model decides (default) | ✓ |
| `{"type":"any"}` | Must call some tool | **400** |
| `{"type":"tool","name":"x"}` | Must call tool x | **400** |
| `{"type":"none"}` | No tools this turn | ✓ |
| `disable_parallel_tool_use: true` | One tool per turn | ✓ |

## Structured outputs

```json
{ "model": "claude-sonnet-5", "max_tokens": 1024,
  "output_config": { "format": { "type": "json_schema", "schema": {
    "type": "object",
    "properties": {
      "vendor": { "type": "string" },
      "total": { "type": "number" },
      "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] },
      "line_items": { "type": "array", "items": { "type": "object",
        "properties": { "sku": { "type": "string" }, "qty": { "type": "integer" } },
        "required": ["sku", "qty"], "additionalProperties": false } }
    },
    "required": ["vendor", "total", "currency", "line_items"],
    "additionalProperties": false } } },
  "messages": [{ "role": "user", "content": [
    { "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "…" } },
    { "type": "text", "text": "Extract the invoice." } ] }] }
```

Always validate downstream and run a **validation-retry** loop that feeds the specific error back.

## Prompt caching

```json
{ "system": [{ "type": "text", "text": "<20k-token style guide>", "cache_control": { "type": "ephemeral" } }],
  "tools": [ … ],
  "messages": [ … dynamic content last … ] }
```

- Order: `tools` → `system` → `messages`; cache breakpoints mark the end of a stable prefix.
- Minimum ~1024 tokens (2048 on Haiku 4.5). Up to 4 breakpoints.
- 5-minute TTL default (write 1.25×); 1-hour option (write 2×). Reads 0.1×.
- `usage.cache_read_input_tokens` confirms hits.

## Message Batches

```python
batch = client.messages.batches.create(requests=[
    {"custom_id": f"doc-{i}", "params": {"model": "claude-haiku-4-5", "max_tokens": 512,
     "messages": [{"role": "user", "content": doc}]}} for i, doc in enumerate(docs)])
# poll batch.processing_status until "ended", then stream results
for res in client.messages.batches.results(batch.id):
    if res.result.type == "succeeded": ...
```

50% discount; results within 24 h; per-item success/error; ideal for "overnight, cost matters".

## Errors and retries

| HTTP | Type | Retry? |
| --- | --- | --- |
| 400 | `invalid_request_error` | No – fix the request (e.g., forced `tool_choice` on Fable 5.1, `budget_tokens` on Opus 5) |
| 401 | `authentication_error` | No – key |
| 403 | `permission_error` | No – entitlement |
| 404 | `not_found_error` | No – model/resource |
| 413 | `request_too_large` | No – shrink |
| 429 | `rate_limit_error` | Yes – backoff, honour `retry-after` |
| 500 | `api_error` | Yes – backoff |
| 529 | `overloaded_error` | Yes – backoff, consider fallback model |

Exponential backoff with jitter; idempotency keys for side-effecting tools; log `request_id` from response headers.

## Other inputs and features

| Feature | Shape |
| --- | --- |
| Vision | Content block `type: image` with a `source` of type `base64` or `url` |
| PDF | Content block `type: document` with a `base64` or `url` source and `media_type: application/pdf` |
| Files API | Upload once, reference by `file_id` in a content block |
| Citations | Enable on documents to get source spans |
| Server-side tools | `web_search`, `code_execution`, `text_editor`, `bash`, `memory`, `computer` |
| Tool search | `tool_search` tool plus `defer_loading: true` on catalogue tools |
| MCP connector | Top-level `mcp_servers` array of objects with `type: url`, `url` and `name` |
| Context editing | `context_management` strategies that clear old tool results server-side |
| Compaction | Server-side summarisation preserving narrative |
| Memory tool | Persistent file-like store across sessions |

```json
// Vision and PDF content blocks
{ "type": "image",    "source": { "type": "url", "url": "https://example.com/chart.png" } }
{ "type": "document", "source": { "type": "base64", "media_type": "application/pdf", "data": "…" } }

// MCP connector
{ "mcp_servers": [{ "type": "url", "url": "https://mcp.example.com/mcp", "name": "orders" }] }
```

## Full streaming event sequence

The wire format is SSE. A complete turn with one text block and one tool call looks like this (elided deltas marked `…`):

```text
event: message_start
data: {"type":"message_start","message":{"id":"msg_01","role":"assistant","model":"claude-opus-5","content":[],"stop_reason":null,"usage":{"input_tokens":1200,"output_tokens":1}}}

event: content_block_start
data: {"type":"content_block_start","index":0,"content_block":{"type":"thinking","thinking":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"thinking_delta","thinking":"Checking the order…"}}

event: content_block_delta
data: {"type":"content_block_delta","index":0,"delta":{"type":"signature_delta","signature":"Er8B…"}}

event: content_block_stop
data: {"type":"content_block_stop","index":0}

event: content_block_start
data: {"type":"content_block_start","index":1,"content_block":{"type":"text","text":""}}

event: content_block_delta
data: {"type":"content_block_delta","index":1,"delta":{"type":"text_delta","text":"Looking that up"}}

event: content_block_stop
data: {"type":"content_block_stop","index":1}

event: content_block_start
data: {"type":"content_block_start","index":2,"content_block":{"type":"tool_use","id":"toolu_01","name":"get_order","input":{}}}

event: content_block_delta
data: {"type":"content_block_delta","index":2,"delta":{"type":"input_json_delta","partial_json":"{\"order_id\":\"ORD-12345\"}"}}

event: content_block_stop
data: {"type":"content_block_stop","index":2}

event: message_delta
data: {"type":"message_delta","delta":{"stop_reason":"tool_use","stop_sequence":null},"usage":{"output_tokens":57}}

event: message_stop
data: {"type":"message_stop"}

event: ping        (may arrive at any time; ignore)
```

| Event | Carries | Notes |
| --- | --- | --- |
| `message_start` | Shell message, input usage | `content` is empty; `stop_reason` null |
| `content_block_start` | Block type at an `index` | One per text/thinking/tool_use block |
| `content_block_delta` | `text_delta`, `thinking_delta`, `signature_delta`, `input_json_delta` | Tool input streams as partial JSON – buffer per index and parse at stop |
| `content_block_stop` | Block finished | – |
| `message_delta` | **`stop_reason`** and cumulative output usage | Branch here, not on prose |
| `message_stop` | Turn complete | – |
| `ping` | Keep-alive | Ignore |
| `error` | `overloaded_error` etc. mid-stream | Handle like the HTTP error |

:::caution[Streaming pitfall]
Tool input arrives as `input_json_delta` fragments; never `JSON.parse` a partial. Accumulate `partial_json` per block index and parse only after `content_block_stop`. The authoritative `stop_reason` is on `message_delta`.
:::

## tool_result with images

A tool can return an image (e.g., a rendered chart) alongside text. The `content` of a `tool_result` accepts a block array:

```json
{ "role": "user", "content": [
  { "type": "tool_result", "tool_use_id": "toolu_09", "content": [
    { "type": "text", "text": "Chart rendered for Q3 revenue." },
    { "type": "image", "source": { "type": "base64", "media_type": "image/png", "data": "iVBORw0KGgo…" } }
  ] }
]}
```

## Parallel tool calls – full round trip

The model can emit several `tool_use` blocks in one turn. Execute them concurrently and return **all** `tool_result` blocks in the next single user message.

```json
// Assistant turn: three parallel calls (stop_reason: tool_use)
{ "role": "assistant", "content": [
  { "type": "tool_use", "id": "toolu_a", "name": "get_weather", "input": { "city": "Paris" } },
  { "type": "tool_use", "id": "toolu_b", "name": "get_weather", "input": { "city": "Tokyo" } },
  { "type": "tool_use", "id": "toolu_c", "name": "get_fx",      "input": { "pair": "EURJPY" } }
]}

// Your next user turn: all results, matched by tool_use_id, order-independent
{ "role": "user", "content": [
  { "type": "tool_result", "tool_use_id": "toolu_a", "content": "{\"c\":18}" },
  { "type": "tool_result", "tool_use_id": "toolu_b", "content": "{\"c\":26}" },
  { "type": "tool_result", "tool_use_id": "toolu_c", "content": "{\"rate\":171.2}" }
]}
```

Use `disable_parallel_tool_use: true` in `tool_choice` to force one call per turn when calls have side effects that must be ordered.

## Structured output with validation-retry

```python
import json, jsonschema
from anthropic import Anthropic

client = Anthropic()
SCHEMA = { "type": "object",
  "properties": { "vendor": {"type": "string"}, "total": {"type": "number"},
                  "currency": {"type": "string", "enum": ["USD","EUR","GBP"]} },
  "required": ["vendor","total","currency"], "additionalProperties": False }

def extract(doc_text, max_attempts=3):
    messages = [{"role": "user", "content": f"Extract the invoice.\n<doc>{doc_text}</doc>"}]
    for attempt in range(max_attempts):
        r = client.messages.create(
            model="claude-sonnet-5", max_tokens=1024,
            output_config={"format": {"type": "json_schema", "schema": SCHEMA}},
            messages=messages)
        text = r.content[0].text
        try:
            data = json.loads(text)
            jsonschema.validate(data, SCHEMA)      # schema + business rules
            if data["total"] < 0:
                raise ValueError("total must be non-negative")
            return data
        except (json.JSONDecodeError, jsonschema.ValidationError, ValueError) as e:
            messages += [{"role": "assistant", "content": text},
                         {"role": "user", "content": f"That failed validation: {e}. Return corrected JSON only."}]
    raise RuntimeError("extraction failed after retries")   # escalate, never silently return bad data
```

The loop feeds the **specific** error back and caps attempts, then escalates — never returns an empty or unvalidated object (silent-failure anti-pattern).

## Prompt caching – multi-breakpoint layout

Up to four breakpoints. Order most-stable → least-stable so the longest possible prefix stays cached when only the tail changes.

```json
{
  "system": [
    { "type": "text", "text": "<static company style guide, 15k tokens>", "cache_control": { "type": "ephemeral" } }
  ],
  "tools": [
    { "name": "search_kb", "description": "…", "input_schema": { }, "cache_control": { "type": "ephemeral" } }
  ],
  "messages": [
    { "role": "user", "content": [
      { "type": "text", "text": "<retrieved policy docs, changes per session>", "cache_control": { "type": "ephemeral", "ttl": "1h" } },
      { "type": "text", "text": "Now: the user's current question (never cached)." }
    ]}
  ]
}
```

| Breakpoint | Content | TTL | Rationale |
| --- | --- | --- | --- |
| 1 | System style guide | 5-min | Never changes; deepest prefix |
| 2 | Tool definitions | 5-min | Stable across the app |
| 3 | Session documents | 1-hour | Reused all session; longer TTL earns the 2× write |
| — | Current question | none | Unique per request |

Rule: a breakpoint caches everything **before** it. Putting a volatile block early invalidates all deeper caching.

## Batch lifecycle states

```text
create → in_progress ──► (per request: succeeded | errored | canceled | expired)
                     └─► ended   (all requests terminal; results retrievable)
        cancel ─► canceling ─► ended
```

| `processing_status` | Meaning |
| --- | --- |
| `in_progress` | Still running; poll `request_counts` |
| `canceling` | Cancellation requested |
| `ended` | Terminal; fetch results stream |

| Per-request `result.type` | Handling |
| --- | --- |
| `succeeded` | Use `result.message` |
| `errored` | Inspect `result.error`; may re-submit that item |
| `canceled` | Batch was canceled before this ran |
| `expired` | Not completed within the 24 h window; re-submit |

Results are available for 29 days. Match items by `custom_id`; do not assume order.

## Error handling with backoff

<Tabs>
  <TabItem label="Python">
```python
import time, random
from anthropic import Anthropic, APIStatusError, RateLimitError, APIConnectionError

client = Anthropic()
RETRYABLE = {429, 500, 502, 503, 529}

def call_with_retry(**params):
    for attempt in range(6):
        try:
            return client.messages.create(**params)
        except RateLimitError as e:
            wait = float(e.response.headers.get("retry-after", 0)) or min(60, 2 ** attempt)
            time.sleep(wait + random.uniform(0, 0.5))
        except APIStatusError as e:
            if e.status_code in RETRYABLE:
                time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
            else:
                raise                              # 400/401/403/404/413 → fix, don't retry
        except APIConnectionError:
            time.sleep(min(60, 2 ** attempt) + random.uniform(0, 0.5))
    raise RuntimeError("exhausted retries")
```
  </TabItem>
  <TabItem label="TypeScript">
```typescript
import Anthropic, { APIError } from '@anthropic-ai/sdk';

const client = new Anthropic();
const RETRYABLE = new Set([429, 500, 502, 503, 529]);
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

async function callWithRetry(params: Anthropic.MessageCreateParamsNonStreaming) {
  for (let attempt = 0; attempt < 6; attempt++) {
    try {
      return await client.messages.create(params);
    } catch (err) {
      if (err instanceof APIError && (RETRYABLE.has(err.status ?? 0))) {
        const retryAfter = Number(err.headers?.['retry-after']) || Math.min(60, 2 ** attempt);
        await sleep((retryAfter + Math.random() * 0.5) * 1000);
        continue;
      }
      throw err; // 4xx client errors → fix the request
    }
  }
  throw new Error('exhausted retries');
}
```
  </TabItem>
</Tabs>

The SDKs retry automatically with backoff; a custom loop matters when you tune the ceiling, add jitter, or switch to a fallback model on repeated 529.

## Rate-limit headers and idempotency

| Header | Meaning |
| --- | --- |
| `anthropic-ratelimit-requests-remaining` | RPM budget left |
| `anthropic-ratelimit-input-tokens-remaining` | ITPM budget left |
| `anthropic-ratelimit-output-tokens-remaining` | OTPM budget left |
| `anthropic-ratelimit-*-reset` | When each bucket refills (RFC 3339) |
| `retry-after` | Seconds to wait after a 429/529 – honour it |
| `request-id` | Log this for every request; include in support tickets |

Proactively throttle when a `*-remaining` header approaches zero rather than waiting for the 429.

```python
# Idempotency: safe retries for side-effecting requests
client.messages.create(**params, extra_headers={"idempotency-key": f"charge-{order_id}"})
```

Reuse the same key on retries so a duplicate delivery does not double-charge. Design side-effecting **tools** the same way (accept an idempotency key argument).

## Files API and citations

```python
# Upload once, reference by file_id across many requests
f = client.files.upload(file=("contract.pdf", open("contract.pdf","rb"), "application/pdf"))

r = client.messages.create(
    model="claude-sonnet-5", max_tokens=1024,
    messages=[{"role": "user", "content": [
        {"type": "document", "source": {"type": "file", "file_id": f.id},
         "citations": {"enabled": True}},
        {"type": "text", "text": "What is the termination notice period? Cite the clause."}
    ]}])
```

With citations enabled, text blocks carry a `citations` array of source spans:

```json
{ "type": "text", "text": "The notice period is 60 days.",
  "citations": [{ "type": "page_location", "cited_text": "…sixty (60) days…",
                  "document_index": 0, "start_page_number": 4, "end_page_number": 4 }] }
```

Citations enable the provenance test: every claim points back to a span you can open.

## Context editing and compaction request shapes

```json
// Context editing: clear stale tool results server-side, keep the turn valid
{ "model": "claude-opus-5", "max_tokens": 4096,
  "context_management": {
    "edits": [{ "type": "clear_tool_uses", "trigger": { "type": "input_tokens", "value": 100000 },
                "keep": { "type": "tool_uses", "value": 3 } }]
  },
  "messages": [ … ] }
```

```json
// Compaction: summarise older turns while preserving the narrative
{ "context_management": {
    "edits": [{ "type": "compact", "trigger": { "type": "input_tokens", "value": 150000 } }] } }
```

| Strategy | Removes | Keeps | Use when |
| --- | --- | --- | --- |
| Context editing (`clear_tool_uses`) | Verbose old tool results | Recent N tool uses, all text | Long tool-heavy agent runs |
| Compaction (`compact`) | Old turns → summary | Narrative continuity | Long conversational sessions |

Both run server-side, so they keep Fable 5.1's append-only history valid (client-side trimming would not).

## Memory tool

```json
{ "model": "claude-opus-5", "max_tokens": 2048,
  "tools": [{ "type": "memory_20250818", "name": "memory" }],
  "messages": [{ "role": "user", "content": "Remember I prefer metric units, then convert 5 miles." }] }
```

The model reads/writes a persistent file-like store (via `memory` tool calls you execute against your backing store) that survives compaction and new sessions. Use for durable preferences and state; do not stuff it into the system prompt.

## MCP connector (server-side)

```python
r = client.messages.create(
    model="claude-opus-5", max_tokens=1024,
    mcp_servers=[{ "type": "url", "url": "https://mcp.example.com/mcp", "name": "orders",
                   "authorization_token": user_scoped_token }],
    extra_headers={"anthropic-beta": "mcp-client-2025-04-04"},
    messages=[{"role": "user", "content": "Where is ORD-12345?"}])
```

Anthropic's infrastructure connects to the remote MCP server for you — no local client harness. Pass a **user-scoped** token so tools act as the end user, not a shared super-user.

## Managed Agents vs Agent SDK

| | Managed Agents | Agent SDK (`claude-agent-sdk`) |
| --- | --- | --- |
| Who runs the loop | Anthropic (hosted loop + sandbox) | You (your infra) |
| Ops burden | Minimal | You own scaling, sandboxing, secrets |
| Control over runtime/network/data locality | Limited | Full |
| Tools/MCP/hooks/subagents | Configured | Full programmatic control |
| Choose when | "least operational overhead", "don't want to host" | "control the runtime", "data must stay in our VPC", "custom harness" |

```python
# Agent SDK sketch – you host the harness
from claude_agent_sdk import ClaudeAgent
agent = ClaudeAgent(model="claude-opus-5", tools=[...], mcp_servers=[...],
                    permission_mode="acceptEdits")
result = agent.run("Refactor the auth module and run the tests.")
```

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "tool_result goes in an assistant message" | It goes in a **user** message, matched by `tool_use_id` | Wrong-role distractor |
| "Parse the text for 'done' to end the loop" | Branch on `stop_reason` | Prose-parsing anti-pattern |
| "`max_tokens` truncation is a finished answer" | It is truncation — continue or raise the cap | Silent-failure distractor |
| "Retry every error" | Only 429/5xx/529 with backoff; fix 4xx | Retry-storm distractor |
| "Structured output means you can skip validation" | Still validate + retry business rules | Over-trust distractor |
| "Caching a volatile block early saves money" | It invalidates every deeper cache; stable first | Caching-layout distractor |
| "Empty result is fine when a tool finds nothing" | Return structured `is_error`/`not_found` | Silent empty-success anti-pattern |

## Scenario walkthrough

A team runs an agent that calls 3–4 tools per turn (some parallel), on Opus 5, in a long session that occasionally hits 529s and grows past 150k tokens. Payments tools must not double-charge. What does a correct implementation look like?

<Steps>
1. **Loop control** — branch on `stop_reason`; on `tool_use` execute all parallel `tool_use` blocks concurrently and return every `tool_result` in one user message.
2. **Ordering side effects** — for the payment tool, set `disable_parallel_tool_use` when it must be sequenced, and pass an idempotency key so a retried call is safe.
3. **529 handling** — exponential backoff with jitter honouring `retry-after`; after repeated 529, fall back to a newer-or-equal model (Opus 5 → Fable 5.1 is up-safe; never down to an older model mid-session if thinking blocks matter).
4. **Context growth** — configure `clear_tool_uses` context editing at ~100k input tokens keeping the last 3 tool uses; server-side so history stays valid.
5. **Errors from tools** — structured `is_error` results with category/retryable, never empty success.
6. **Observability** — log `request-id` and `usage` per call; watch `anthropic-ratelimit-*-remaining` and throttle before 429.
</Steps>

Rejected alternatives: parsing prose to stop (prose-parsing), retrying 400s (retry-storm), client-side history trimming (breaks append-only), and self-reported "I finished" as the loop exit (self-report reliance).
