# D2 · Model Selection and Optimization

LLM fundamentals, model options and tiers, pricing and cost math, breaking changes and migration, token budgeting, caching, batching, routing and latency levers.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This domain is roughly **9 of 53 items**. It tests whether you can pick the cheapest model that meets the quality bar and then drive cost and latency down with caching, batching, routing and output control – while pinning versions and handling breaking changes across releases. Expect worked-cost scenarios and "which model / which lever" trade-off questions.

## Learning objectives

By the end of this page you should be able to:

1. Explain **LLM fundamentals**: tokens, tokenisation cost, context window, sampling and non-determinism.
2. Configure **model options**: fast mode, extended/adaptive thinking, effort levels, and where `budget_tokens` still applies.
3. Compare the **model tiers** (Haiku 4.5, Sonnet 5, Opus 5, Fable 5.1) on capability, price and latency, and do **worked cost calculations**.
4. Identify **breaking changes** across releases and plan **pinning and migration**.
5. Apply **cost levers**: token budgeting, prompt caching, batching, and **routing/cascading**.
6. Apply **latency levers**: streaming, smaller models, shorter output, caching.

---

## 2.1 LLM fundamentals

- **Tokens** are sub-word units; billing and limits are in tokens, not characters. Roughly 1 token ≈ 4 characters ≈ 0.75 words in English.
- **Tokenisation cost**: both input and output are metered in tokens; longer prompts and outputs cost more and add latency.
- **Context window** is the total tokens (input + output) the model can attend to in one request – 1M on Fable 5.1 / Opus 5 / Sonnet 5, 200k on Haiku 4.5. `max_tokens` caps output within that window.
- **Sampling**: the model outputs a probability distribution; `temperature` and `top_p` control how it is sampled. Lower = more deterministic.
- **Non-determinism**: even at `temperature: 0`, outputs are not guaranteed byte-identical across runs or model versions. Design for it – validate output, do not assume exact repeats.

:::tip[Exam signal]
"Why did the same prompt give a different answer?" → non-determinism; reduce with `temperature: 0` and a pinned snapshot, but never assume perfect repeatability.
:::

---

## 2.2 Model options: thinking, effort and fast mode

| Option | What it does | Where |
| --- | --- | --- |
| `thinking: {type: 'adaptive'}` | Model decides how much to reason | All current models |
| `thinking: {type: 'enabled', budget_tokens: N}` | Fixed reasoning budget | **Haiku 4.5 only** |
| Effort `low\|medium\|high\|xhigh` | Tune reasoning depth; `high` default | Opus 5 / Sonnet 5 / Fable 5.1 (not Haiku) |
| Fast mode | Latency-optimised responses | Latency-sensitive use |

- Use **`xhigh`** effort only for the hardest coding/agentic work (Opus 5 / Fable 5.1) – it costs more thinking tokens and adds latency.
- **`budget_tokens`** returns `400` on Fable 5.x / Opus 5 / Sonnet 5. Only Haiku 4.5 still accepts it, and Haiku has **no** `effort` parameter.
- **Fable 5.1** always thinks; you cannot disable it.

---

## 2.3 Model tiers, pricing and trade-offs

| Model | ID | Context / max out | In / Out per MTok | Cache read | Use when |
| --- | --- | --- | --- | --- | --- |
| Fable 5.1 | `claude-fable-5-1` | 1M / 128k | $10 / $50 | $0.25 | Most capable; thinking always on; append-only harness |
| Opus 5 | `claude-opus-5` | 1M / 128k | $5 / $25 | — | Complex agentic coding, enterprise work |
| Sonnet 5 | `claude-sonnet-5` | 1M / 128k | $2 / $10 | — | Best speed/intelligence balance (default) |
| Haiku 4.5 | `claude-haiku-4-5` | 200k / 64k | $1 / $5 | — | Fastest/cheapest; high-volume simple tasks |

Rule of thumb: **start at Sonnet 5**; drop to Haiku 4.5 for high-volume simple work; move to Opus 5 for hard agentic/coding tasks; reserve Fable 5.1 for the most demanding reasoning (accepting its constraints and 30-day retention).

### Worked cost calculation

A task sends 3,000 input tokens and generates 500 output tokens, 20,000 times/day.

| Model | Input cost/day | Output cost/day | Total/day |
| --- | --- | --- | --- |
| Haiku 4.5 | 3,000 × 20,000 × $1 / 1e6 = $60 | 500 × 20,000 × $5 / 1e6 = $50 | **$110** |
| Sonnet 5 | $120 | $100 | **$220** |
| Opus 5 | $300 | $250 | **$550** |

If Haiku 4.5 meets quality, it is **5×** cheaper than Opus 5 here. Never default to the most capable model for simple, high-volume work.

:::tip[Exam signal]
"High volume + simple task + cost-sensitive" → Haiku 4.5. "Complex multi-step agentic coding" → Opus 5 (or Fable 5.1 for the hardest). "Balanced default" → Sonnet 5.
:::

---

## 2.4 Breaking changes across releases

| Change | Effect | Mitigation |
| --- | --- | --- |
| `budget_tokens` removed (Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8) | `400` if sent | Use `thinking: {type: 'adaptive'}` |
| Fable 5.1: `tool_choice: 'any'` and forced `{type:'tool'}` return `400` | Cannot force a tool | Use `auto` + instruction, `strict: true`, or structured outputs |
| Fable 5.1: thinking-binding | Thinking blocks readable only by producing model or newer; silent drop on fallback | Do not silently fall back mid-conversation |
| Fable 5.1: append-only | Editing/reordering earlier turns invalidates later thinking | Freeze `system`/`tools`; trim via context editing/compaction |
| Opus 4.1 retired 2026-08-05 | Requests fail | Migrate to a current snapshot |

:::danger[Fable 5.1 tool_choice]
On Fable 5.1, `tool_choice: 'any'` and forcing a specific tool both return `400`. If your extraction relies on forcing a tool, migrate to structured outputs or `strict: true`, or keep `tool_choice: 'auto'` with a clear instruction.
:::

---

## 2.5 Pinning and migration

- **Pin a snapshot** in production so behaviour does not shift under you. Model IDs from 4.6 onward are dateless but are still pinned snapshots.
- Store the model ID in **one** env-driven place so migration is a single change.
- **Migration checklist:** re-run the golden eval set on the new model; check for removed parameters (`budget_tokens`), changed `tool_choice` behaviour, latency and cost deltas; roll out behind a flag.
- Use `client.models.list()` / `.retrieve(id)` for live limits and availability.

```python
import os
MODEL = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5")   # pinned, one place
```

---

## 2.6 Token budgeting

Control tokens on both sides:

- **Input**: trim context, cache stable prefixes, prune old tool results, summarise long histories.
- **Output**: set a realistic `max_tokens`, ask for concise/structured output, stop early with `stop_sequences`.
- **Thinking**: adaptive by default; use effort levels rather than paying for `xhigh` everywhere.

Every token in the context window is billed each turn, so long conversations get expensive – budget with caching, editing and compaction (Domain 4).

---

## 2.7 Cost levers: caching, batching, routing

<Tabs>
  <TabItem label="Prompt caching">
Cache a stable prefix (system + tools + long docs). Cache read ≈ 0.1× input; write ≈ 1.25× (5-min) or 2× (1-hour). Worked example: a 20k-token cached prefix on Sonnet 5 costs 20k×$2/1e6 = $0.04 uncached vs ~$0.004 on a cache hit – **10×** cheaper on the prefix.
  </TabItem>
  <TabItem label="Batching">
Message Batches API: **50% off** input and output, results within 24h. Ideal for classification, extraction, evals, backfills – anything not real-time.
  </TabItem>
  <TabItem label="Routing / cascading">
Send easy requests to a cheap model and escalate only hard ones. A classifier or the cheap model's own confidence/failure routes to a bigger model. Most traffic stays cheap; only the tail pays for Opus/Fable.
  </TabItem>
</Tabs>

```text
Request ─► Haiku 4.5 (cheap, fast)
             │
             ├─ confident / simple ─► return
             └─ hard / low-confidence ─► Sonnet 5 ─► (rare) Opus 5
```

:::tip[Exam signal]
"Most requests are simple, a few are hard, minimise cost" → routing/cascading cheap→expensive. "Reduce cost without changing quality on repeated shared context" → caching. "Bulk, not real-time" → batching.
:::

---

## 2.8 Latency levers

| Lever | Effect | Trade-off |
| --- | --- | --- |
| **Streaming** | Improves *perceived* latency; tokens appear immediately | No change to total time |
| **Smaller model** | Lower time-to-completion | May reduce quality |
| **Shorter output** | Fewer output tokens = faster + cheaper | Less detail |
| **Prompt caching** | Skips reprocessing the prefix | Only helps stable prefixes |
| **Lower/adaptive thinking effort** | Less reasoning time | May reduce quality on hard tasks |
| **Fast mode** | Latency-optimised | For latency-sensitive use |

:::note[Streaming vs latency]
Streaming does not make total generation faster – it makes the *first token* arrive sooner, improving the user's experience. If total wall-clock time is the constraint, use a smaller model or shorter output.
:::

---

## 2.9 Worked cost comparison: a 10k-document job

Consider a batch extraction job over **10,000 documents**. Each request sends a **2,000-token shared prefix** (system + schema + instructions) plus **1,000 tokens of unique document text** (3,000 input tokens total) and generates **500 output tokens**. Below is the arithmetic for each tier, first synchronous and uncached, then with caching, then via the Batches API.

### Baseline (synchronous, no caching)

Cost per request = (input tokens × in-price + output tokens × out-price) ÷ 1,000,000, times 10,000 documents.

| Model | Input: 3,000 × 10k × price | Output: 500 × 10k × price | Total |
| --- | --- | --- | --- |
| Haiku 4.5 | 30M × $1 / 1e6 = $30 | 5M × $5 / 1e6 = $25 | **$55** |
| Sonnet 5 | 30M × $2 / 1e6 = $60 | 5M × $10 / 1e6 = $50 | **$110** |
| Opus 5 | 30M × $5 / 1e6 = $150 | 5M × $25 / 1e6 = $125 | **$275** |

### With prompt caching on the 2,000-token shared prefix

The 2,000-token prefix is written to cache once (≈1.25× the base input price for the 5-minute TTL) and read at ≈0.1× thereafter. The 1,000 unique tokens per document are never cacheable. For Sonnet 5:

- Cache write (first request): 2,000 × $2 × 1.25 / 1e6 = $0.005 (once, negligible).
- Prefix cache reads (remaining 9,999 requests): ≈2,000 × 10k × $2 × 0.1 / 1e6 = **$4** (versus $40 uncached for the prefix portion).
- Unique input (never cached): 1,000 × 10k × $2 / 1e6 = **$20**.
- Output unchanged: **$50**.
- Sonnet 5 with caching ≈ **$74** versus $110 — the shared prefix drops from $40 to about $4.

:::caution[Caching only helps a *stable, repeated* prefix]
The 2,000-token prefix must be byte-identical and marked with `cache_control` for hits. The 1,000 unique document tokens vary every request, so they are never cached. Caching is worthless if the prefix is below the minimum (~1024 tokens; 2048 on Haiku) or changes each call.
:::

### With the Message Batches API (50% off input and output)

Batching halves every line above and returns results within 24 hours — ideal because this job is latency-tolerant.

| Model | Sync (uncached) | Batch (uncached) | Batch + caching |
| --- | --- | --- | --- |
| Haiku 4.5 | $55 | **$27.50** | ≈ $25 |
| Sonnet 5 | $110 | **$55** | ≈ $37 |
| Opus 5 | $275 | **$137.50** | ≈ $100 |

:::tip[Exam signal]
'10,000 documents overnight, cost-sensitive' → cheapest model that clears the quality bar + **Batches (50%)** + **cache the shared prefix**. Stacking the three levers on Sonnet 5 takes the job from $110 to roughly $37. Do not reach for Opus 5 unless Sonnet fails the quality bar.
:::

---

## 2.10 Latency levers in depth

Total latency has two visible parts: **time-to-first-token (TTFT)** and **total generation time**. Different levers move different parts. Confusing them is a classic exam trap.

| Lever | Moves TTFT? | Moves total time? | Trade-off / note |
| --- | --- | --- | --- |
| **Streaming (SSE)** | Improves *perceived* start | No change to total | Tokens render as produced; user sees progress sooner |
| **Shorter output (`max_tokens`, concise instruction)** | Slight | Yes — fewer output tokens = faster + cheaper | Less detail |
| **Smaller model tier** | Yes | Yes | Haiku 4.5 fastest; may reduce quality |
| **Prompt caching** | Yes (skips prefix reprocessing) | Yes on prefix | Only helps a stable, repeated prefix |
| **Lower / adaptive thinking effort** | Yes | Yes | `low`/`medium` vs `high`/`xhigh`; may hurt hard tasks |
| **Fast mode** | Yes | Yes | Latency-optimised path for latency-sensitive use |

```text
Request ──► [cache read prefix] ──► [thinking effort] ──► [generate output]
   │              (caching)            (effort/fast)         (max_tokens)
   └── stream deltas to user as soon as generation starts (perceived latency)
```

:::note[Streaming is not a total-latency lever]
Streaming improves the *feel* by getting the first token out sooner; it does not reduce total wall-clock generation time. If total time is the hard constraint, use a **smaller model**, **shorter output**, **lower thinking effort**, or **fast mode** — not streaming alone.
:::

---

## 2.11 Pinning and migration checklists for breaking changes

Pin a snapshot in production and drive the model ID from one env-backed constant. When you migrate — or when a release ships a breaking change — walk a checklist rather than swapping IDs blind.

### General migration checklist

<Steps>

1. Change the single model constant behind a feature flag (do not edit IDs across many files).

2. Re-run the golden eval set on the candidate model; compare quality, latency p50/p95, and cost per task.

3. Scan the request payload for **removed parameters** (see below) and **changed behaviours** (`tool_choice`, thinking binding).

4. Roll out to a small traffic slice, watch error rates (400 spikes signal a rejected parameter), then ramp.

</Steps>

### Breaking-change specifics

| Breaking change | Symptom | Fix |
| --- | --- | --- |
| `budget_tokens` removed (Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8) | `400 invalid_request` when the field is sent | Replace with `thinking: \{type: 'adaptive'\}`; use effort levels; keep `budget_tokens` only on Haiku 4.5 |
| Fable 5.1 rejects `tool_choice: 'any'` / forced `\{type:'tool'\}` | `400` on every forced-tool request | Use `tool_choice: 'auto'` + instruction, `strict: true`, or `output_config.format` structured outputs |
| Fable 5.1 thinking binding | Thinking blocks silently dropped when a request falls back to an older model | Do not silently fall back mid-conversation; pin one model per session |
| Fable 5.1 append-only harness | Later thinking blocks invalidated after editing/reordering earlier turns | Freeze `system`/`tools`; put mid-session changes in a `role: 'system'` message; trim server-side via context editing/compaction |
| Opus 4.1 retired 2026-08-05 | `404 not_found` / request failure | Migrate to a current snapshot |

:::danger[Fable 5.1 is append-only]
On Fable 5.1 you cannot edit, reorder, or remove earlier turns without invalidating later thinking blocks. Harnesses must be **append-only**: freeze `system` and `tools`, express mid-session changes as `role: 'system'` messages, and trim only via server-side context editing or compaction. A harness that rewrites history will break on Fable 5.1.
:::

---

## 2.12 A routing / cascade implementation

Routing keeps most traffic on a cheap model and escalates only the hard tail. The decision to escalate should come from a **structured signal** (a classifier label, a validation failure, or a `refusal`/low-confidence tool result) — never from parsing prose.

```python
import os
from anthropic import Anthropic

client = Anthropic()

CHEAP = "claude-haiku-4-5"
MID = "claude-sonnet-5"
HARD = "claude-opus-5"

def answer(messages, tools=None):
    """Cascade: try cheap, escalate on refusal/low confidence or validation failure."""
    resp = client.messages.create(model=CHEAP, max_tokens=1024, messages=messages, tools=tools or [])
    if resp.stop_reason == "refusal" or not passes_quality_gate(resp):
        resp = client.messages.create(model=MID, max_tokens=2048, messages=messages, tools=tools or [])
        if resp.stop_reason == "refusal" or not passes_quality_gate(resp):
            resp = client.messages.create(model=HARD, max_tokens=4096, messages=messages, tools=tools or [])
    return resp

def passes_quality_gate(resp) -> bool:
    """Programmatic gate: schema validity, required fields present, no empty answer.
    Do NOT use the model's self-reported confidence (anti-pattern #4)."""
    text = "".join(b.text for b in resp.content if getattr(b, "type", None) == "text")
    return len(text.strip()) > 0 and "I cannot" not in text
```

```text
Request ─► Haiku 4.5 ──(gate passes)──► return
              │
              └─(refusal / gate fails)─► Sonnet 5 ──(gate passes)──► return
                                            │
                                            └─(rare)─► Opus 5 ──► return
```

:::tip[Exam signal]
'Most requests simple, a few hard, minimise cost, do not hurt the hard cases' → **cascade cheap→expensive** with a *programmatic* escalation gate. If the stem escalates on the model's own confidence score, that is anti-pattern #4 (self-report reliance) and wrong.
:::

---

## 2.13 A cost model you can compute in your head

The exam expects fast, defensible arithmetic. Memorise the per-MTok prices and the four multipliers, then estimate.

| Quantity | Formula |
| --- | --- |
| Base input cost | `input_tokens × in_price / 1e6` |
| Base output cost | `output_tokens × out_price / 1e6` |
| Cache **write** (5-min) | base input × **1.25** |
| Cache **write** (1-hour) | base input × **2** |
| Cache **read** (hit) | base input × **0.1** |
| Batch discount | every line × **0.5** |

Prices per MTok: Haiku 4.5 **$1/$5**, Sonnet 5 **$2/$10**, Opus 5 **$5/$25**, Fable 5.1 **$10/$50**.

**Worked drill.** 4,000 input + 800 output tokens, 30,000×/day, on Haiku vs Sonnet:

| Model | Input/day | Output/day | Total/day |
| --- | --- | --- | --- |
| Haiku 4.5 | 4000×30000×$1/1e6 = **$120** | 800×30000×$5/1e6 = **$120** | **$240** |
| Sonnet 5 | **$240** | **$240** | **$480** |

Haiku is exactly half here — but only choose it if it clears the quality bar. Then, if the workload is latency-tolerant, halve again with Batches, and cut the shared-prefix portion to ~10% with caching.

:::tip[Exam signal]
When a stem gives you token counts and a volume, it wants arithmetic. Compute `input×price + output×price`, then apply ×0.5 for Batches and ×0.1 for cached-prefix hits. The cheapest *adequate* model wins — never default to the most capable.
:::

---

## 2.14 Choosing thinking effort deliberately

Effort and thinking mode trade quality for latency and token cost. The exam tests picking the *minimum* that clears the bar.

| Setting | Cost / latency | Use when | Avoid when |
| --- | --- | --- | --- |
| `adaptive` (default reasoning) | Model decides depth | General agentic/reasoning work | — |
| effort `low`/`medium` | Cheaper, faster | Simple or latency-sensitive tasks | Hard multi-step proofs/coding |
| effort `high` (default) | Balanced | Most non-trivial reasoning | Trivial classification (wasteful) |
| effort `xhigh` | Most tokens, slowest | The hardest coding/agentic work (Opus 5 / Fable 5.1) | Anything routine — it just burns tokens |
| `budget_tokens` (Haiku 4.5 only) | Fixed budget | Bounding Haiku reasoning cost | Any other model → 400 |

:::caution[`xhigh` and `budget_tokens` are narrow]
`xhigh` is for the hardest work only; setting it everywhere inflates cost and latency for no quality gain. `budget_tokens` is **Haiku-4.5-only** — sending it to Sonnet 5 / Opus 5 / Fable 5.1 returns `400`. Haiku 4.5 in turn has **no** `effort` parameter, so 'Haiku with `xhigh`' is an invalid combination the exam uses as a distractor.
:::

---

## 2.15 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| The most capable model is the safe default | Start at Sonnet 5; use the cheapest model that clears the bar | Over-defaulting to Opus 5/Fable 5.1 is the top cost distractor |
| `temperature: 0` gives identical output | It reduces variance but is not byte-deterministic; pin a snapshot too | Reproducibility questions hinge on this nuance |
| Streaming reduces total latency | It only improves time-to-first-token | Separates perceived from actual latency |
| `budget_tokens` works if it is small enough | It is Haiku-4.5-only; others `400` regardless of value | A recurring breaking-change trap |
| Caching helps any repeated context | Only a byte-identical prefix above the minimum caches | Per-request values in the prefix kill the cache |
| Batching just adds latency for no benefit | It is 50% off for latency-tolerant bulk work | The cheapest path for offline jobs |
| A bigger model fixes rate limits | Rate limits are about token/request volume, not model choice | Model choice is the wrong lever for `429` |
| Escalate to a bigger model on the model's confidence score | Escalate on programmatic signals, not self-report (anti-pattern #4) | Cascade-design questions test this |

---

## 2.16 Scenario walkthrough: taming a runaway model bill

**Scenario.** A team runs everything on Opus 5 'to be safe'. The workload is 95% short FAQ-style classifications and 5% genuinely hard multi-step reasoning. The bill is 5× budget. Latency for the classifications must stay interactive; the hard cases can take longer. They also run a nightly 20,000-item eval on the same Opus 5. A proposal on the table: 'escalate to Fable 5.1 whenever Opus 5 reports confidence below 0.9.' You must cut cost without hurting the hard cases or the eval's fidelity.

**Expert reasoning trace.**

1. **Split the traffic by difficulty.** 95% is trivial → move it to **Haiku 4.5** (cheapest, fast enough to stay interactive). Keep a path to a bigger model for the 5%. Running everything on Opus 5 is the constraint-blind default that caused the overspend.
2. **Design the escalation gate correctly.** Escalate on a **programmatic signal** — a classifier label, a schema-validation failure, or a `refusal` — not on the model's **self-reported confidence** (anti-pattern #4). The proposed 0.9-confidence gate is the wrong design.
3. **Pick the escalation target by need, not prestige.** Most hard cases clear on **Sonnet 5**; reserve **Opus 5** for the rare hardest tail. Jumping straight to Fable 5.1 is over-engineered and adds its 30-day-retention and append-only constraints for no benefit here.
4. **Handle the nightly eval separately.** It is latency-tolerant and bulk → run it via the **Batches API** (50% off) on the cheapest model that clears the eval's quality bar. Do not confuse the eval's model with the production classifier's.
5. **Verify with arithmetic.** If the 95% moves from Opus ($5/$25) to Haiku ($1/$5), the bulk cost drops ~5×; only the 5% tail pays Sonnet/Opus rates. That alone likely brings the bill within budget.
6. **Reject the tempting alternatives.** 'Keep Opus 5 but lower `max_tokens`' — trims output only, not the model-tier overspend. 'Use Fable 5.1 for the hard cases because it is most capable' — over-engineered and adds retention/append-only constraints. 'Escalate on confidence' — self-report reliance.

**Correct decision.** Cascade Haiku 4.5 → Sonnet 5 → (rare) Opus 5 with a *programmatic* escalation gate; run the nightly eval on the cheapest adequate model via Batches. Pin snapshots and gate the change behind the golden eval set before rollout.

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Default to Opus 5 / Fable 5.1 for everything | Wastes cost; start at Sonnet 5, drop to Haiku for simple high-volume work |
| Assume `temperature: 0` gives identical outputs | Reduces variance but is not byte-deterministic across runs/versions |
| Send `budget_tokens` to Sonnet 5 / Opus 5 / Fable 5.1 | Returns `400`; Haiku-4.5-only |
| Force a tool on Fable 5.1 with `tool_choice` | `any`/forced tool return `400`; use structured outputs / `strict` |
| Treat streaming as a total-latency reduction | It only improves time-to-first-token |
| Cache a prefix that changes every request | No cache hits |
| Use synchronous calls for bulk latency-tolerant work | Batches give 50% off |
| Hard-code model IDs across many files | Makes migration error-prone; pin in one place |
| Escalate to a bigger model on the model's self-reported confidence | Self-report reliance (anti-pattern #4); gate on programmatic validation |
| Rewrite/reorder history in a Fable 5.1 harness | Invalidates later thinking blocks; harness must be append-only |
| Reach for Opus 5 on a latency-tolerant bulk job before trying Batches + caching | Wastes cost; Batches (50%) + prefix caching often make Sonnet 5 cheaper than needed |
| Treat the whole 3,000-token input as cacheable when only 2,000 tokens are a stable prefix | The unique per-document tokens never cache; only the repeated prefix does |
| Setting `xhigh` effort on routine tasks | Burns thinking tokens and latency for no quality gain; reserve it for the hardest work |
| Combining Haiku 4.5 with `effort: 'xhigh'` | Invalid — Haiku has no `effort` parameter; a common distractor |
| Lowering `max_tokens` to fix a model-tier overspend | Trims output only; the real lever is the cheaper model + caching + batching |
| Choosing a bigger model to resolve `429` rate limits | Rate limits track token/request volume, not model tier |
| Confusing a production model choice with the eval's model choice | Evals are latency-tolerant → cheapest adequate model + Batches, independent of prod |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A pipeline classifies 100,000 short support tickets per night with a simple label. Latency is irrelevant. Which choice minimises cost? (Select one)">
    A. Opus 5, synchronous, high concurrency.
    B. Sonnet 5, streaming.
    C. Haiku 4.5 via the Message Batches API.
    D. Fable 5.1 with `xhigh` effort.

    **Answer: C.** Simple, high-volume, latency-tolerant → cheapest model (Haiku 4.5) plus Batches (50% off). Opus/Fable (A, D) are over-engineered and costly; streaming (B) does not cut cost.
  </AccordionItem>

  <AccordionItem title="Q2 · A developer sends `budget_tokens: 1500` to Opus 5 and gets a 400. What is the fix? (Select one)">
    A. Lower the budget to 1000.
    B. Use `thinking: {type: 'adaptive'}`; `budget_tokens` is only valid on Haiku 4.5.
    C. Disable thinking entirely.
    D. Increase `max_tokens`.

    **Answer: B.** `budget_tokens` was removed on Opus 5; current models use adaptive thinking with effort levels. Only Haiku 4.5 still uses `budget_tokens`.
  </AccordionItem>

  <AccordionItem title="Q3 · Most requests to a service are trivial; a small fraction need deep reasoning. Cost must be minimised without hurting the hard cases. What design fits? (Select one)">
    A. Run everything on Opus 5.
    B. Run everything on Haiku 4.5.
    C. Route: handle simple requests on Haiku 4.5 and escalate hard/low-confidence ones to Sonnet 5 or Opus 5.
    D. Randomly assign models.

    **Answer: C.** Cascading routing keeps the bulk cheap and only pays for a bigger model on the hard tail. Uniform Opus (A) is wasteful; uniform Haiku (B) fails the hard cases; random (D) is nonsense.
  </AccordionItem>

  <AccordionItem title="Q4 · Each request reuses the same 15,000-token instruction-and-schema prefix on Sonnet 5. Which change cuts input cost the most? (Select one)">
    A. Switch to Haiku 4.5.
    B. Mark the stable prefix with `cache_control` and reuse it within the TTL.
    C. Set `temperature: 0`.
    D. Enable streaming.

    **Answer: B.** Prompt caching drops the prefix cost to ~10% on hits. Switching model (A) changes quality; temperature (C) and streaming (D) do not affect input cost.
  </AccordionItem>

  <AccordionItem title="Q5 · A user complains the app feels slow even though total time is fine for the task. Which lever best improves the experience with no quality change? (Select one)">
    A. Streaming so tokens render as they are produced.
    B. Switch to Opus 5.
    C. Increase `max_tokens`.
    D. Enable `xhigh` effort.

    **Answer: A.** Streaming improves perceived latency (time-to-first-token) without changing output. A bigger model (B), more output (C) and higher effort (D) would add latency.
  </AccordionItem>

  <AccordionItem title="Q6 · An extraction service forces a specific tool with `tool_choice` and is migrating to Fable 5.1. What breaks and what is the fix? (Select one)">
    A. Nothing changes.
    B. Forced `tool_choice` returns 400 on Fable 5.1; migrate to structured outputs or `strict: true`, or use `auto` with an instruction.
    C. Fable 5.1 does not support tools.
    D. Increase `max_tokens`.

    **Answer: B.** Fable 5.1 rejects `tool_choice: 'any'` and forced tools with 400. Structured outputs / `strict` / `auto`+instruction are the supported paths. Fable does support tools (C).
  </AccordionItem>

  <AccordionItem title="Q7 · Which TWO statements about the context window and `max_tokens` are correct? (Select two)">
    A. The context window includes both input and output tokens.
    B. `max_tokens` sets the maximum output tokens, within the context window.
    C. `max_tokens` is the size of the context window.
    D. Haiku 4.5 has a 1M context window.
    E. Sonnet 5 has a 1M context window.

    **Answer: A and B.** Context window = input + output; `max_tokens` caps output. Haiku 4.5 is 200k, not 1M (D wrong); Sonnet 5 is 1M (E correct but the pair A+B is the intended answer set) — select A and B as the two statements about context window and `max_tokens`.
  </AccordionItem>

  <AccordionItem title="Q8 · A team pins `claude-sonnet-5` and wants to trial Opus 5 safely. What is the correct migration practice? (Select one)">
    A. Swap the ID everywhere and ship.
    B. Change the single env-driven model constant behind a flag, re-run the golden eval set, check for removed parameters and cost/latency deltas, then roll out.
    C. Let each service pick its own model.
    D. Only change it in production to get real signal.

    **Answer: B.** Centralised, flagged, eval-gated migration is correct. Swapping everywhere (A) or per-service drift (C) or prod-only changes (D) are risky.
  </AccordionItem>

  <AccordionItem title="Q9 · A latency-tolerant nightly eval of 5,000 prompts must be as cheap as possible. Which combination is best? (Select two)">
    A. Message Batches API for the 50% discount.
    B. Streaming each request.
    C. The cheapest model that meets the quality bar.
    D. `xhigh` effort on Fable 5.1.
    E. Opus 5 for every prompt.

    **Answer: A and C.** Batching (50% off) plus the cheapest adequate model minimises cost for offline work. Streaming (B) does not cut cost; high-effort Fable (D) and blanket Opus (E) are expensive.
  </AccordionItem>

  <AccordionItem title="Q10 · A 10,000-document extraction job runs overnight and is cost-sensitive. Sonnet 5 clears the quality bar. Each request has a 2,000-token stable prefix and 1,000 unique tokens. Which combination minimises cost? (Select one)">
    A. Opus 5, synchronous, no caching, for maximum quality.
    B. Sonnet 5 via the Batches API, with `cache_control` on the 2,000-token shared prefix.
    C. Sonnet 5 synchronous with caching only.
    D. Haiku 4.5 synchronous with `xhigh` effort.

    **Answer: B.** The job is latency-tolerant, so stack all three levers: cheapest model that clears the bar (Sonnet 5), Batches (50% off), and caching on the stable prefix — roughly $37 versus $110 baseline. Opus (A) is over-engineered; caching alone (C) leaves the 50% batch discount on the table; Haiku with `xhigh` (D) is invalid because Haiku has no `effort` parameter.
  </AccordionItem>

  <AccordionItem title="Q11 · A team migrates a production service to Fable 5.1. Their harness edits earlier turns to trim context and forces a specific extraction tool. What two problems will they hit? (Select two)">
    A. Forced `tool_choice` returns 400 on Fable 5.1.
    B. Editing earlier turns invalidates later thinking blocks; the harness must be append-only.
    C. Fable 5.1 does not support tools at all.
    D. Fable 5.1 requires `budget_tokens`.
    E. Fable 5.1 has no context window.

    **Answer: A and B.** Fable 5.1 rejects forced tools (use `output_config.format` / `strict` / `auto`+instruction) and requires an append-only harness (trim only via server-side context editing/compaction). Fable does support tools (C); `budget_tokens` is removed on Fable, not required (D); it has a 1M context window (E).
  </AccordionItem>

  <AccordionItem title="Q12 · Users report the app 'feels sluggish' but the total task time is acceptable, and quality must not drop. Which single change best helps? (Select one)">
    A. Enable streaming so tokens render as they are produced.
    B. Switch every request to Opus 5.
    C. Raise `max_tokens` to 8,000.
    D. Set effort to `xhigh`.

    **Answer: A.** The complaint is perceived latency (time-to-first-token), and quality must be preserved — streaming shows progress immediately without changing output. A bigger model (B), more output (C), and higher effort (D) all *increase* latency and change behaviour.
  </AccordionItem>

  <AccordionItem title="Q13 · A service routes most traffic to Haiku 4.5 and escalates hard cases to Sonnet 5. A developer proposes escalating whenever Claude says its own confidence is below 0.8. Why is this wrong, and what is the fix? (Select one)">
    A. It is correct; self-reported confidence is reliable.
    B. It relies on self-reported confidence (anti-pattern #4); escalate on a programmatic signal — schema-validation failure, `refusal`, or a classifier label.
    C. Routing is never appropriate; use one model.
    D. Escalate on output length instead.

    **Answer: B.** Self-reported confidence is not a trustworthy escalation trigger (anti-pattern #4). A cascade should escalate on structured, programmatic signals such as a failed validation, a `refusal` stop reason, or a classifier decision. Single-model (C) defeats the cost goal; output length (D) is not a quality signal.
  </AccordionItem>

  <AccordionItem title="Q14 · A task sends 4,000 input + 800 output tokens, 30,000×/day. Both Haiku 4.5 ($1/$5) and Sonnet 5 ($2/$10) clear the bar. What is the daily cost difference, and which is cheaper? (Select one)">
    A. They cost the same because the token counts are identical.
    B. Haiku 4.5 (~$240/day) is cheaper than Sonnet 5 (~$480/day) by about $240/day.
    C. Sonnet 5 is cheaper because larger models batch better.
    D. Haiku 4.5 ~$480/day; Sonnet 5 ~$240/day.

    **Answer: B.** Haiku: 4000×30000×$1/1e6=$120 + 800×30000×$5/1e6=$120 = $240; Sonnet: $240+$240=$480. Same tokens but different per-token prices (A wrong); there is no 'batches better' effect for a bigger model (C); D reverses the arithmetic.
  </AccordionItem>

  <AccordionItem title="Q15 · A developer sets `effort: 'xhigh'` on a high-volume, trivial classification task to 'be safe'. What is the effect? (Select one)">
    A. Higher quality at no extra cost.
    B. More thinking tokens and higher latency for no quality gain on a trivial task; use `low`/`medium` or `adaptive` instead.
    C. It disables sampling and caches results.
    D. It is required for classification.

    **Answer: B.** `xhigh` is for the hardest work; on trivial tasks it just burns tokens and time. It is not free (A), does not cache (C), and is not required (D).
  </AccordionItem>

  <AccordionItem title="Q16 · Which combination is invalid because of a model-specific parameter restriction? (Select one)">
    A. Sonnet 5 with `thinking: {type: 'adaptive'}`.
    B. Haiku 4.5 with `effort: 'xhigh'`.
    C. Opus 5 with `effort: 'high'`.
    D. Haiku 4.5 with `budget_tokens: 1024`.

    **Answer: B.** Haiku 4.5 has no `effort` parameter, so `effort: 'xhigh'` on Haiku is invalid. Adaptive on Sonnet 5 (A), `effort: high` on Opus 5 (C) and `budget_tokens` on Haiku 4.5 (D) are all valid.
  </AccordionItem>

  <AccordionItem title="Q17 · A team runs 95% trivial classifications and 5% hard reasoning entirely on Opus 5 and the bill is 5× budget. Which redesign cuts cost without hurting the hard cases? (Select one)">
    A. Lower `max_tokens` on all Opus 5 calls.
    B. Cascade: Haiku 4.5 for the trivial bulk, escalating hard/low-quality cases (via a programmatic gate) to Sonnet 5 and rarely Opus 5.
    C. Move everything to Fable 5.1 because it is most capable.
    D. Enable streaming on every call.

    **Answer: B.** Routing the trivial bulk to Haiku and reserving bigger models for the hard tail cuts the dominant cost while protecting hard cases. Lowering `max_tokens` (A) trims output only; Fable 5.1 everywhere (C) is over-engineered and adds retention/append-only constraints; streaming (D) does not cut cost.
  </AccordionItem>

  <AccordionItem title="Q18 · A nightly 20,000-item eval runs on the same Opus 5 as production. How should the eval's cost be minimised independently? (Select two)">
    A. Run the eval via the Message Batches API for the 50% discount.
    B. Use the cheapest model that clears the eval's quality bar.
    C. Force `xhigh` effort for rigour.
    D. Stream every eval request.
    E. Keep it on Opus 5 to match production exactly.

    **Answer: A and B.** The eval is latency-tolerant and bulk, so Batches plus the cheapest adequate model minimise cost. `xhigh` (C) inflates cost; streaming (D) does not cut cost; matching production on Opus 5 (E) is unnecessary and expensive for an offline eval.
  </AccordionItem>

  <AccordionItem title="Q19 · Which TWO statements about the current lineup's context windows are correct? (Select two)">
    A. Sonnet 5 and Opus 5 each have a 1M context window.
    B. Haiku 4.5 has a 200k context window.
    C. Haiku 4.5 has a 1M context window.
    D. Fable 5.1 has a 128k context window.
    E. Sonnet 5 has a 200k context window.

    **Answer: A and B.** Sonnet 5 and Opus 5 are 1M; Haiku 4.5 is 200k. Haiku is not 1M (C); Fable 5.1's 128k is its MAX OUTPUT, not context (D); Sonnet 5 is 1M, not 200k (E).
  </AccordionItem>
</Accordions>

## Key takeaways

- Tokens drive cost, latency and limits; the context window holds input + output, `max_tokens` caps output only.
- `temperature: 0` reduces variance but is not byte-deterministic; pin a snapshot for reproducibility.
- Start at Sonnet 5; Haiku 4.5 for cheap high-volume, Opus 5 for hard agentic work, Fable 5.1 for the hardest.
- `budget_tokens` is Haiku-4.5-only; effort levels apply to Opus 5 / Sonnet 5 / Fable 5.1.
- Know the breaking changes: `budget_tokens` removal, Fable 5.1 `tool_choice`/thinking-binding/append-only, Opus 4.1 retirement.
- Cost levers: caching (~10% on hits), batching (50% off), routing/cascading cheap→expensive.
- Latency levers: streaming (perceived), smaller model / shorter output (actual), caching, fast mode.
- Compute cost in your head: `input×price + output×price`, then ×0.5 for Batches and ×0.1 for cached-prefix hits; the cheapest *adequate* model wins.
- Pick the minimum thinking effort that clears the bar; `xhigh` is for the hardest work only, `budget_tokens` is Haiku-4.5-only, and 'Haiku + `effort`' is invalid.
- Rate limits are about token/request volume — fix them with caching, trimming, batching or a higher tier, not by changing model tier.
- Treat a production model choice and an offline eval's model choice as independent decisions; evals are latency-tolerant, so use the cheapest adequate model via Batches.
