Domains
D2 · Model Selection and Optimization
LLM fundamentals, model options and tiers, pricing and cost math, breaking changes and migration, token budgeting, caching, batching, routing and latency levers.
This domain is roughly 9 of 53 items. It tests whether you can pick the cheapest model that meets the quality bar and then drive cost and latency down with caching, batching, routing and output control – while pinning versions and handling breaking changes across releases. Expect worked-cost scenarios and “which model / which lever” trade-off questions.
Learning objectives
By the end of this page you should be able to:
- Explain LLM fundamentals: tokens, tokenisation cost, context window, sampling and non-determinism.
- Configure model options: fast mode, extended/adaptive thinking, effort levels, and where
budget_tokensstill applies. - Compare the model tiers (Haiku 4.5, Sonnet 5, Opus 5, Fable 5.1) on capability, price and latency, and do worked cost calculations.
- Identify breaking changes across releases and plan pinning and migration.
- Apply cost levers: token budgeting, prompt caching, batching, and routing/cascading.
- Apply latency levers: streaming, smaller models, shorter output, caching.
2.1 LLM fundamentals
- Tokens are sub-word units; billing and limits are in tokens, not characters. Roughly 1 token ≈ 4 characters ≈ 0.75 words in English.
- Tokenisation cost: both input and output are metered in tokens; longer prompts and outputs cost more and add latency.
- Context window is the total tokens (input + output) the model can attend to in one request – 1M on Fable 5.1 / Opus 5 / Sonnet 5, 200k on Haiku 4.5.
max_tokenscaps output within that window. - Sampling: the model outputs a probability distribution;
temperatureandtop_pcontrol how it is sampled. Lower = more deterministic. - Non-determinism: even at
temperature: 0, outputs are not guaranteed byte-identical across runs or model versions. Design for it – validate output, do not assume exact repeats.
Exam signal
“Why did the same prompt give a different answer?” → non-determinism; reduce with temperature: 0 and a pinned snapshot, but never assume perfect repeatability.
2.2 Model options: thinking, effort and fast mode
| Option | What it does | Where |
|---|---|---|
thinking: {type: 'adaptive'} | Model decides how much to reason | All current models |
thinking: {type: 'enabled', budget_tokens: N} | Fixed reasoning budget | Haiku 4.5 only |
Effort low|medium|high|xhigh | Tune reasoning depth; high default | Opus 5 / Sonnet 5 / Fable 5.1 (not Haiku) |
| Fast mode | Latency-optimised responses | Latency-sensitive use |
- Use
xhigheffort only for the hardest coding/agentic work (Opus 5 / Fable 5.1) – it costs more thinking tokens and adds latency. budget_tokensreturns400on Fable 5.x / Opus 5 / Sonnet 5. Only Haiku 4.5 still accepts it, and Haiku has noeffortparameter.- Fable 5.1 always thinks; you cannot disable it.
2.3 Model tiers, pricing and trade-offs
| Model | ID | Context / max out | In / Out per MTok | Cache read | Use when |
|---|---|---|---|---|---|
| Fable 5.1 | claude-fable-5-1 | 1M / 128k | $10 / $50 | $0.25 | Most capable; thinking always on; append-only harness |
| Opus 5 | claude-opus-5 | 1M / 128k | $5 / $25 | — | Complex agentic coding, enterprise work |
| Sonnet 5 | claude-sonnet-5 | 1M / 128k | $2 / $10 | — | Best speed/intelligence balance (default) |
| Haiku 4.5 | claude-haiku-4-5 | 200k / 64k | $1 / $5 | — | Fastest/cheapest; high-volume simple tasks |
Rule of thumb: start at Sonnet 5; drop to Haiku 4.5 for high-volume simple work; move to Opus 5 for hard agentic/coding tasks; reserve Fable 5.1 for the most demanding reasoning (accepting its constraints and 30-day retention).
Worked cost calculation
A task sends 3,000 input tokens and generates 500 output tokens, 20,000 times/day.
| Model | Input cost/day | Output cost/day | Total/day |
|---|---|---|---|
| Haiku 4.5 | 3,000 × 20,000 × $1 / 1e6 = $60 | 500 × 20,000 × $5 / 1e6 = $50 | $110 |
| Sonnet 5 | $120 | $100 | $220 |
| Opus 5 | $300 | $250 | $550 |
If Haiku 4.5 meets quality, it is 5× cheaper than Opus 5 here. Never default to the most capable model for simple, high-volume work.
Exam signal
“High volume + simple task + cost-sensitive” → Haiku 4.5. “Complex multi-step agentic coding” → Opus 5 (or Fable 5.1 for the hardest). “Balanced default” → Sonnet 5.
2.4 Breaking changes across releases
| Change | Effect | Mitigation |
|---|---|---|
budget_tokens removed (Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8) | 400 if sent | Use thinking: {type: 'adaptive'} |
Fable 5.1: tool_choice: 'any' and forced {type:'tool'} return 400 | Cannot force a tool | Use auto + instruction, strict: true, or structured outputs |
| Fable 5.1: thinking-binding | Thinking blocks readable only by producing model or newer; silent drop on fallback | Do not silently fall back mid-conversation |
| Fable 5.1: append-only | Editing/reordering earlier turns invalidates later thinking | Freeze system/tools; trim via context editing/compaction |
| Opus 4.1 retired 2026-08-05 | Requests fail | Migrate to a current snapshot |
Fable 5.1 tool_choice
On Fable 5.1, tool_choice: 'any' and forcing a specific tool both return 400. If your extraction relies on forcing a tool, migrate to structured outputs or strict: true, or keep tool_choice: 'auto' with a clear instruction.
2.5 Pinning and migration
- Pin a snapshot in production so behaviour does not shift under you. Model IDs from 4.6 onward are dateless but are still pinned snapshots.
- Store the model ID in one env-driven place so migration is a single change.
- Migration checklist: re-run the golden eval set on the new model; check for removed parameters (
budget_tokens), changedtool_choicebehaviour, latency and cost deltas; roll out behind a flag. - Use
client.models.list()/.retrieve(id)for live limits and availability.
import osMODEL = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5") # pinned, one place2.6 Token budgeting
Control tokens on both sides:
- Input: trim context, cache stable prefixes, prune old tool results, summarise long histories.
- Output: set a realistic
max_tokens, ask for concise/structured output, stop early withstop_sequences. - Thinking: adaptive by default; use effort levels rather than paying for
xhigheverywhere.
Every token in the context window is billed each turn, so long conversations get expensive – budget with caching, editing and compaction (Domain 4).
2.7 Cost levers: caching, batching, routing
Cache a stable prefix (system + tools + long docs). Cache read ≈ 0.1× input; write ≈ 1.25× (5-min) or 2× (1-hour). Worked example: a 20k-token cached prefix on Sonnet 5 costs 20k×$2/1e6 = $0.04 uncached vs ~$0.004 on a cache hit – 10× cheaper on the prefix.
Message Batches API: 50% off input and output, results within 24h. Ideal for classification, extraction, evals, backfills – anything not real-time.
Send easy requests to a cheap model and escalate only hard ones. A classifier or the cheap model’s own confidence/failure routes to a bigger model. Most traffic stays cheap; only the tail pays for Opus/Fable.
Request ─► Haiku 4.5 (cheap, fast) │ ├─ confident / simple ─► return └─ hard / low-confidence ─► Sonnet 5 ─► (rare) Opus 5Exam signal
“Most requests are simple, a few are hard, minimise cost” → routing/cascading cheap→expensive. “Reduce cost without changing quality on repeated shared context” → caching. “Bulk, not real-time” → batching.
2.8 Latency levers
| Lever | Effect | Trade-off |
|---|---|---|
| Streaming | Improves perceived latency; tokens appear immediately | No change to total time |
| Smaller model | Lower time-to-completion | May reduce quality |
| Shorter output | Fewer output tokens = faster + cheaper | Less detail |
| Prompt caching | Skips reprocessing the prefix | Only helps stable prefixes |
| Lower/adaptive thinking effort | Less reasoning time | May reduce quality on hard tasks |
| Fast mode | Latency-optimised | For latency-sensitive use |
Streaming vs latency
Streaming does not make total generation faster – it makes the first token arrive sooner, improving the user’s experience. If total wall-clock time is the constraint, use a smaller model or shorter output.
2.9 Worked cost comparison: a 10k-document job
Consider a batch extraction job over 10,000 documents. Each request sends a 2,000-token shared prefix (system + schema + instructions) plus 1,000 tokens of unique document text (3,000 input tokens total) and generates 500 output tokens. Below is the arithmetic for each tier, first synchronous and uncached, then with caching, then via the Batches API.
Baseline (synchronous, no caching)
Cost per request = (input tokens × in-price + output tokens × out-price) ÷ 1,000,000, times 10,000 documents.
| Model | Input: 3,000 × 10k × price | Output: 500 × 10k × price | Total |
|---|---|---|---|
| Haiku 4.5 | 30M × $1 / 1e6 = $30 | 5M × $5 / 1e6 = $25 | $55 |
| Sonnet 5 | 30M × $2 / 1e6 = $60 | 5M × $10 / 1e6 = $50 | $110 |
| Opus 5 | 30M × $5 / 1e6 = $150 | 5M × $25 / 1e6 = $125 | $275 |
With prompt caching on the 2,000-token shared prefix
The 2,000-token prefix is written to cache once (≈1.25× the base input price for the 5-minute TTL) and read at ≈0.1× thereafter. The 1,000 unique tokens per document are never cacheable. For Sonnet 5:
- Cache write (first request): 2,000 × $2 × 1.25 / 1e6 = $0.005 (once, negligible).
- Prefix cache reads (remaining 9,999 requests): ≈2,000 × 10k × $2 × 0.1 / 1e6 = $4 (versus $40 uncached for the prefix portion).
- Unique input (never cached): 1,000 × 10k × $2 / 1e6 = $20.
- Output unchanged: $50.
- Sonnet 5 with caching ≈ $74 versus $110 — the shared prefix drops from $40 to about $4.
Caching only helps a stable, repeated prefix
The 2,000-token prefix must be byte-identical and marked with cache_control for hits. The 1,000 unique document tokens vary every request, so they are never cached. Caching is worthless if the prefix is below the minimum (~1024 tokens; 2048 on Haiku) or changes each call.
With the Message Batches API (50% off input and output)
Batching halves every line above and returns results within 24 hours — ideal because this job is latency-tolerant.
| Model | Sync (uncached) | Batch (uncached) | Batch + caching |
|---|---|---|---|
| Haiku 4.5 | $55 | $27.50 | ≈ $25 |
| Sonnet 5 | $110 | $55 | ≈ $37 |
| Opus 5 | $275 | $137.50 | ≈ $100 |
Exam signal
‘10,000 documents overnight, cost-sensitive’ → cheapest model that clears the quality bar + Batches (50%) + cache the shared prefix. Stacking the three levers on Sonnet 5 takes the job from $110 to roughly $37. Do not reach for Opus 5 unless Sonnet fails the quality bar.
2.10 Latency levers in depth
Total latency has two visible parts: time-to-first-token (TTFT) and total generation time. Different levers move different parts. Confusing them is a classic exam trap.
| Lever | Moves TTFT? | Moves total time? | Trade-off / note |
|---|---|---|---|
| Streaming (SSE) | Improves perceived start | No change to total | Tokens render as produced; user sees progress sooner |
Shorter output (max_tokens, concise instruction) | Slight | Yes — fewer output tokens = faster + cheaper | Less detail |
| Smaller model tier | Yes | Yes | Haiku 4.5 fastest; may reduce quality |
| Prompt caching | Yes (skips prefix reprocessing) | Yes on prefix | Only helps a stable, repeated prefix |
| Lower / adaptive thinking effort | Yes | Yes | low/medium vs high/xhigh; may hurt hard tasks |
| Fast mode | Yes | Yes | Latency-optimised path for latency-sensitive use |
Request ──► [cache read prefix] ──► [thinking effort] ──► [generate output] │ (caching) (effort/fast) (max_tokens) └── stream deltas to user as soon as generation starts (perceived latency)Streaming is not a total-latency lever
Streaming improves the feel by getting the first token out sooner; it does not reduce total wall-clock generation time. If total time is the hard constraint, use a smaller model, shorter output, lower thinking effort, or fast mode — not streaming alone.
2.11 Pinning and migration checklists for breaking changes
Pin a snapshot in production and drive the model ID from one env-backed constant. When you migrate — or when a release ships a breaking change — walk a checklist rather than swapping IDs blind.
General migration checklist
-
Change the single model constant behind a feature flag (do not edit IDs across many files).
-
Re-run the golden eval set on the candidate model; compare quality, latency p50/p95, and cost per task.
-
Scan the request payload for removed parameters (see below) and changed behaviours (
tool_choice, thinking binding). -
Roll out to a small traffic slice, watch error rates (400 spikes signal a rejected parameter), then ramp.
Breaking-change specifics
| Breaking change | Symptom | Fix |
|---|---|---|
budget_tokens removed (Fable 5.x, Opus 5, Sonnet 5, Opus 4.7–4.8) | 400 invalid_request when the field is sent | Replace with thinking: \{type: 'adaptive'\}; use effort levels; keep budget_tokens only on Haiku 4.5 |
Fable 5.1 rejects tool_choice: 'any' / forced \{type:'tool'\} | 400 on every forced-tool request | Use tool_choice: 'auto' + instruction, strict: true, or output_config.format structured outputs |
| Fable 5.1 thinking binding | Thinking blocks silently dropped when a request falls back to an older model | Do not silently fall back mid-conversation; pin one model per session |
| Fable 5.1 append-only harness | Later thinking blocks invalidated after editing/reordering earlier turns | Freeze system/tools; put mid-session changes in a role: 'system' message; trim server-side via context editing/compaction |
| Opus 4.1 retired 2026-08-05 | 404 not_found / request failure | Migrate to a current snapshot |
Fable 5.1 is append-only
On Fable 5.1 you cannot edit, reorder, or remove earlier turns without invalidating later thinking blocks. Harnesses must be append-only: freeze system and tools, express mid-session changes as role: 'system' messages, and trim only via server-side context editing or compaction. A harness that rewrites history will break on Fable 5.1.
2.12 A routing / cascade implementation
Routing keeps most traffic on a cheap model and escalates only the hard tail. The decision to escalate should come from a structured signal (a classifier label, a validation failure, or a refusal/low-confidence tool result) — never from parsing prose.
import osfrom anthropic import Anthropic
client = Anthropic()
CHEAP = "claude-haiku-4-5"MID = "claude-sonnet-5"HARD = "claude-opus-5"
def answer(messages, tools=None): """Cascade: try cheap, escalate on refusal/low confidence or validation failure.""" resp = client.messages.create(model=CHEAP, max_tokens=1024, messages=messages, tools=tools or []) if resp.stop_reason == "refusal" or not passes_quality_gate(resp): resp = client.messages.create(model=MID, max_tokens=2048, messages=messages, tools=tools or []) if resp.stop_reason == "refusal" or not passes_quality_gate(resp): resp = client.messages.create(model=HARD, max_tokens=4096, messages=messages, tools=tools or []) return resp
def passes_quality_gate(resp) -> bool: """Programmatic gate: schema validity, required fields present, no empty answer. Do NOT use the model's self-reported confidence (anti-pattern #4).""" text = "".join(b.text for b in resp.content if getattr(b, "type", None) == "text") return len(text.strip()) > 0 and "I cannot" not in textRequest ─► Haiku 4.5 ──(gate passes)──► return │ └─(refusal / gate fails)─► Sonnet 5 ──(gate passes)──► return │ └─(rare)─► Opus 5 ──► returnExam signal
‘Most requests simple, a few hard, minimise cost, do not hurt the hard cases’ → cascade cheap→expensive with a programmatic escalation gate. If the stem escalates on the model’s own confidence score, that is anti-pattern #4 (self-report reliance) and wrong.
2.13 A cost model you can compute in your head
The exam expects fast, defensible arithmetic. Memorise the per-MTok prices and the four multipliers, then estimate.
| Quantity | Formula |
|---|---|
| Base input cost | input_tokens × in_price / 1e6 |
| Base output cost | output_tokens × out_price / 1e6 |
| Cache write (5-min) | base input × 1.25 |
| Cache write (1-hour) | base input × 2 |
| Cache read (hit) | base input × 0.1 |
| Batch discount | every line × 0.5 |
Prices per MTok: Haiku 4.5 $1/$5, Sonnet 5 $2/$10, Opus 5 $5/$25, Fable 5.1 $10/$50.
Worked drill. 4,000 input + 800 output tokens, 30,000×/day, on Haiku vs Sonnet:
| Model | Input/day | Output/day | Total/day |
|---|---|---|---|
| Haiku 4.5 | 4000×30000×$1/1e6 = $120 | 800×30000×$5/1e6 = $120 | $240 |
| Sonnet 5 | $240 | $240 | $480 |
Haiku is exactly half here — but only choose it if it clears the quality bar. Then, if the workload is latency-tolerant, halve again with Batches, and cut the shared-prefix portion to ~10% with caching.
Exam signal
When a stem gives you token counts and a volume, it wants arithmetic. Compute input×price + output×price, then apply ×0.5 for Batches and ×0.1 for cached-prefix hits. The cheapest adequate model wins — never default to the most capable.
2.14 Choosing thinking effort deliberately
Effort and thinking mode trade quality for latency and token cost. The exam tests picking the minimum that clears the bar.
| Setting | Cost / latency | Use when | Avoid when |
|---|---|---|---|
adaptive (default reasoning) | Model decides depth | General agentic/reasoning work | — |
effort low/medium | Cheaper, faster | Simple or latency-sensitive tasks | Hard multi-step proofs/coding |
effort high (default) | Balanced | Most non-trivial reasoning | Trivial classification (wasteful) |
effort xhigh | Most tokens, slowest | The hardest coding/agentic work (Opus 5 / Fable 5.1) | Anything routine — it just burns tokens |
budget_tokens (Haiku 4.5 only) | Fixed budget | Bounding Haiku reasoning cost | Any other model → 400 |
xhigh and budget_tokens are narrow
xhigh is for the hardest work only; setting it everywhere inflates cost and latency for no quality gain. budget_tokens is Haiku-4.5-only — sending it to Sonnet 5 / Opus 5 / Fable 5.1 returns 400. Haiku 4.5 in turn has no effort parameter, so ‘Haiku with xhigh’ is an invalid combination the exam uses as a distractor.
2.15 Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| The most capable model is the safe default | Start at Sonnet 5; use the cheapest model that clears the bar | Over-defaulting to Opus 5/Fable 5.1 is the top cost distractor |
temperature: 0 gives identical output | It reduces variance but is not byte-deterministic; pin a snapshot too | Reproducibility questions hinge on this nuance |
| Streaming reduces total latency | It only improves time-to-first-token | Separates perceived from actual latency |
budget_tokens works if it is small enough | It is Haiku-4.5-only; others 400 regardless of value | A recurring breaking-change trap |
| Caching helps any repeated context | Only a byte-identical prefix above the minimum caches | Per-request values in the prefix kill the cache |
| Batching just adds latency for no benefit | It is 50% off for latency-tolerant bulk work | The cheapest path for offline jobs |
| A bigger model fixes rate limits | Rate limits are about token/request volume, not model choice | Model choice is the wrong lever for 429 |
| Escalate to a bigger model on the model’s confidence score | Escalate on programmatic signals, not self-report (anti-pattern #4) | Cascade-design questions test this |
2.16 Scenario walkthrough: taming a runaway model bill
Scenario. A team runs everything on Opus 5 ‘to be safe’. The workload is 95% short FAQ-style classifications and 5% genuinely hard multi-step reasoning. The bill is 5× budget. Latency for the classifications must stay interactive; the hard cases can take longer. They also run a nightly 20,000-item eval on the same Opus 5. A proposal on the table: ‘escalate to Fable 5.1 whenever Opus 5 reports confidence below 0.9.’ You must cut cost without hurting the hard cases or the eval’s fidelity.
Expert reasoning trace.
- Split the traffic by difficulty. 95% is trivial → move it to Haiku 4.5 (cheapest, fast enough to stay interactive). Keep a path to a bigger model for the 5%. Running everything on Opus 5 is the constraint-blind default that caused the overspend.
- Design the escalation gate correctly. Escalate on a programmatic signal — a classifier label, a schema-validation failure, or a
refusal— not on the model’s self-reported confidence (anti-pattern #4). The proposed 0.9-confidence gate is the wrong design. - Pick the escalation target by need, not prestige. Most hard cases clear on Sonnet 5; reserve Opus 5 for the rare hardest tail. Jumping straight to Fable 5.1 is over-engineered and adds its 30-day-retention and append-only constraints for no benefit here.
- Handle the nightly eval separately. It is latency-tolerant and bulk → run it via the Batches API (50% off) on the cheapest model that clears the eval’s quality bar. Do not confuse the eval’s model with the production classifier’s.
- Verify with arithmetic. If the 95% moves from Opus ($5/$25) to Haiku ($1/$5), the bulk cost drops ~5×; only the 5% tail pays Sonnet/Opus rates. That alone likely brings the bill within budget.
- Reject the tempting alternatives. ‘Keep Opus 5 but lower
max_tokens’ — trims output only, not the model-tier overspend. ‘Use Fable 5.1 for the hard cases because it is most capable’ — over-engineered and adds retention/append-only constraints. ‘Escalate on confidence’ — self-report reliance.
Correct decision. Cascade Haiku 4.5 → Sonnet 5 → (rare) Opus 5 with a programmatic escalation gate; run the nightly eval on the cheapest adequate model via Batches. Pin snapshots and gate the change behind the golden eval set before rollout.
Exam traps in this domain
| Trap | Why it is wrong |
|---|---|
| Default to Opus 5 / Fable 5.1 for everything | Wastes cost; start at Sonnet 5, drop to Haiku for simple high-volume work |
Assume temperature: 0 gives identical outputs | Reduces variance but is not byte-deterministic across runs/versions |
Send budget_tokens to Sonnet 5 / Opus 5 / Fable 5.1 | Returns 400; Haiku-4.5-only |
Force a tool on Fable 5.1 with tool_choice | any/forced tool return 400; use structured outputs / strict |
| Treat streaming as a total-latency reduction | It only improves time-to-first-token |
| Cache a prefix that changes every request | No cache hits |
| Use synchronous calls for bulk latency-tolerant work | Batches give 50% off |
| Hard-code model IDs across many files | Makes migration error-prone; pin in one place |
| Escalate to a bigger model on the model’s self-reported confidence | Self-report reliance (anti-pattern #4); gate on programmatic validation |
| Rewrite/reorder history in a Fable 5.1 harness | Invalidates later thinking blocks; harness must be append-only |
| Reach for Opus 5 on a latency-tolerant bulk job before trying Batches + caching | Wastes cost; Batches (50%) + prefix caching often make Sonnet 5 cheaper than needed |
| Treat the whole 3,000-token input as cacheable when only 2,000 tokens are a stable prefix | The unique per-document tokens never cache; only the repeated prefix does |
Setting xhigh effort on routine tasks | Burns thinking tokens and latency for no quality gain; reserve it for the hardest work |
Combining Haiku 4.5 with effort: 'xhigh' | Invalid — Haiku has no effort parameter; a common distractor |
Lowering max_tokens to fix a model-tier overspend | Trims output only; the real lever is the cheaper model + caching + batching |
Choosing a bigger model to resolve 429 rate limits | Rate limits track token/request volume, not model tier |
| Confusing a production model choice with the eval’s model choice | Evals are latency-tolerant → cheapest adequate model + Batches, independent of prod |
Practice questions
Q1 · A pipeline classifies 100,000 short support tickets per night with a simple label. Latency is irrelevant. Which choice minimises cost? (Select one)
A. Opus 5, synchronous, high concurrency.
B. Sonnet 5, streaming.
C. Haiku 4.5 via the Message Batches API.
D. Fable 5.1 with xhigh effort.
Answer: C. Simple, high-volume, latency-tolerant → cheapest model (Haiku 4.5) plus Batches (50% off). Opus/Fable (A, D) are over-engineered and costly; streaming (B) does not cut cost.
Q2 · A developer sends `budget_tokens: 1500` to Opus 5 and gets a 400. What is the fix? (Select one)
A. Lower the budget to 1000.
B. Use thinking: {type: 'adaptive'}; budget_tokens is only valid on Haiku 4.5.
C. Disable thinking entirely.
D. Increase max_tokens.
Answer: B. budget_tokens was removed on Opus 5; current models use adaptive thinking with effort levels. Only Haiku 4.5 still uses budget_tokens.
Q3 · Most requests to a service are trivial; a small fraction need deep reasoning. Cost must be minimised without hurting the hard cases. What design fits? (Select one)
A. Run everything on Opus 5. B. Run everything on Haiku 4.5. C. Route: handle simple requests on Haiku 4.5 and escalate hard/low-confidence ones to Sonnet 5 or Opus 5. D. Randomly assign models.
Answer: C. Cascading routing keeps the bulk cheap and only pays for a bigger model on the hard tail. Uniform Opus (A) is wasteful; uniform Haiku (B) fails the hard cases; random (D) is nonsense.
Q4 · Each request reuses the same 15,000-token instruction-and-schema prefix on Sonnet 5. Which change cuts input cost the most? (Select one)
A. Switch to Haiku 4.5.
B. Mark the stable prefix with cache_control and reuse it within the TTL.
C. Set temperature: 0.
D. Enable streaming.
Answer: B. Prompt caching drops the prefix cost to ~10% on hits. Switching model (A) changes quality; temperature (C) and streaming (D) do not affect input cost.
Q5 · A user complains the app feels slow even though total time is fine for the task. Which lever best improves the experience with no quality change? (Select one)
A. Streaming so tokens render as they are produced.
B. Switch to Opus 5.
C. Increase max_tokens.
D. Enable xhigh effort.
Answer: A. Streaming improves perceived latency (time-to-first-token) without changing output. A bigger model (B), more output (C) and higher effort (D) would add latency.
Q6 · An extraction service forces a specific tool with `tool_choice` and is migrating to Fable 5.1. What breaks and what is the fix? (Select one)
A. Nothing changes.
B. Forced tool_choice returns 400 on Fable 5.1; migrate to structured outputs or strict: true, or use auto with an instruction.
C. Fable 5.1 does not support tools.
D. Increase max_tokens.
Answer: B. Fable 5.1 rejects tool_choice: 'any' and forced tools with 400. Structured outputs / strict / auto+instruction are the supported paths. Fable does support tools (C).
Q7 · Which TWO statements about the context window and `max_tokens` are correct? (Select two)
A. The context window includes both input and output tokens.
B. max_tokens sets the maximum output tokens, within the context window.
C. max_tokens is the size of the context window.
D. Haiku 4.5 has a 1M context window.
E. Sonnet 5 has a 1M context window.
Answer: A and B. Context window = input + output; max_tokens caps output. Haiku 4.5 is 200k, not 1M (D wrong); Sonnet 5 is 1M (E correct but the pair A+B is the intended answer set) — select A and B as the two statements about context window and max_tokens.
Q8 · A team pins `claude-sonnet-5` and wants to trial Opus 5 safely. What is the correct migration practice? (Select one)
A. Swap the ID everywhere and ship. B. Change the single env-driven model constant behind a flag, re-run the golden eval set, check for removed parameters and cost/latency deltas, then roll out. C. Let each service pick its own model. D. Only change it in production to get real signal.
Answer: B. Centralised, flagged, eval-gated migration is correct. Swapping everywhere (A) or per-service drift (C) or prod-only changes (D) are risky.
Q9 · A latency-tolerant nightly eval of 5,000 prompts must be as cheap as possible. Which combination is best? (Select two)
A. Message Batches API for the 50% discount.
B. Streaming each request.
C. The cheapest model that meets the quality bar.
D. xhigh effort on Fable 5.1.
E. Opus 5 for every prompt.
Answer: A and C. Batching (50% off) plus the cheapest adequate model minimises cost for offline work. Streaming (B) does not cut cost; high-effort Fable (D) and blanket Opus (E) are expensive.
Q10 · A 10,000-document extraction job runs overnight and is cost-sensitive. Sonnet 5 clears the quality bar. Each request has a 2,000-token stable prefix and 1,000 unique tokens. Which combination minimises cost? (Select one)
A. Opus 5, synchronous, no caching, for maximum quality.
B. Sonnet 5 via the Batches API, with cache_control on the 2,000-token shared prefix.
C. Sonnet 5 synchronous with caching only.
D. Haiku 4.5 synchronous with xhigh effort.
Answer: B. The job is latency-tolerant, so stack all three levers: cheapest model that clears the bar (Sonnet 5), Batches (50% off), and caching on the stable prefix — roughly $37 versus $110 baseline. Opus (A) is over-engineered; caching alone (C) leaves the 50% batch discount on the table; Haiku with xhigh (D) is invalid because Haiku has no effort parameter.
Q11 · A team migrates a production service to Fable 5.1. Their harness edits earlier turns to trim context and forces a specific extraction tool. What two problems will they hit? (Select two)
A. Forced tool_choice returns 400 on Fable 5.1.
B. Editing earlier turns invalidates later thinking blocks; the harness must be append-only.
C. Fable 5.1 does not support tools at all.
D. Fable 5.1 requires budget_tokens.
E. Fable 5.1 has no context window.
Answer: A and B. Fable 5.1 rejects forced tools (use output_config.format / strict / auto+instruction) and requires an append-only harness (trim only via server-side context editing/compaction). Fable does support tools (C); budget_tokens is removed on Fable, not required (D); it has a 1M context window (E).
Q12 · Users report the app 'feels sluggish' but the total task time is acceptable, and quality must not drop. Which single change best helps? (Select one)
A. Enable streaming so tokens render as they are produced.
B. Switch every request to Opus 5.
C. Raise max_tokens to 8,000.
D. Set effort to xhigh.
Answer: A. The complaint is perceived latency (time-to-first-token), and quality must be preserved — streaming shows progress immediately without changing output. A bigger model (B), more output (C), and higher effort (D) all increase latency and change behaviour.
Q13 · A service routes most traffic to Haiku 4.5 and escalates hard cases to Sonnet 5. A developer proposes escalating whenever Claude says its own confidence is below 0.8. Why is this wrong, and what is the fix? (Select one)
A. It is correct; self-reported confidence is reliable.
B. It relies on self-reported confidence (anti-pattern #4); escalate on a programmatic signal — schema-validation failure, refusal, or a classifier label.
C. Routing is never appropriate; use one model.
D. Escalate on output length instead.
Answer: B. Self-reported confidence is not a trustworthy escalation trigger (anti-pattern #4). A cascade should escalate on structured, programmatic signals such as a failed validation, a refusal stop reason, or a classifier decision. Single-model (C) defeats the cost goal; output length (D) is not a quality signal.
Q14 · A task sends 4,000 input + 800 output tokens, 30,000×/day. Both Haiku 4.5 ($1/$5) and Sonnet 5 ($2/$10) clear the bar. What is the daily cost difference, and which is cheaper? (Select one)
A. They cost the same because the token counts are identical.
B. Haiku 4.5 ($240/day) is cheaper than Sonnet 5 ($480/day) by about $240/day.
C. Sonnet 5 is cheaper because larger models batch better.
D. Haiku 4.5 ~$480/day; Sonnet 5 ~$240/day.
Answer: B. Haiku: 4000×30000×$1/1e6=$120 + 800×30000×$5/1e6=$120 = $240; Sonnet: $240+$240=$480. Same tokens but different per-token prices (A wrong); there is no ‘batches better’ effect for a bigger model (C); D reverses the arithmetic.
Q15 · A developer sets `effort: 'xhigh'` on a high-volume, trivial classification task to 'be safe'. What is the effect? (Select one)
A. Higher quality at no extra cost.
B. More thinking tokens and higher latency for no quality gain on a trivial task; use low/medium or adaptive instead.
C. It disables sampling and caches results.
D. It is required for classification.
Answer: B. xhigh is for the hardest work; on trivial tasks it just burns tokens and time. It is not free (A), does not cache (C), and is not required (D).
Q16 · Which combination is invalid because of a model-specific parameter restriction? (Select one)
A. Sonnet 5 with thinking: {type: 'adaptive'}.
B. Haiku 4.5 with effort: 'xhigh'.
C. Opus 5 with effort: 'high'.
D. Haiku 4.5 with budget_tokens: 1024.
Answer: B. Haiku 4.5 has no effort parameter, so effort: 'xhigh' on Haiku is invalid. Adaptive on Sonnet 5 (A), effort: high on Opus 5 (C) and budget_tokens on Haiku 4.5 (D) are all valid.
Q17 · A team runs 95% trivial classifications and 5% hard reasoning entirely on Opus 5 and the bill is 5× budget. Which redesign cuts cost without hurting the hard cases? (Select one)
A. Lower max_tokens on all Opus 5 calls.
B. Cascade: Haiku 4.5 for the trivial bulk, escalating hard/low-quality cases (via a programmatic gate) to Sonnet 5 and rarely Opus 5.
C. Move everything to Fable 5.1 because it is most capable.
D. Enable streaming on every call.
Answer: B. Routing the trivial bulk to Haiku and reserving bigger models for the hard tail cuts the dominant cost while protecting hard cases. Lowering max_tokens (A) trims output only; Fable 5.1 everywhere (C) is over-engineered and adds retention/append-only constraints; streaming (D) does not cut cost.
Q18 · A nightly 20,000-item eval runs on the same Opus 5 as production. How should the eval's cost be minimised independently? (Select two)
A. Run the eval via the Message Batches API for the 50% discount.
B. Use the cheapest model that clears the eval’s quality bar.
C. Force xhigh effort for rigour.
D. Stream every eval request.
E. Keep it on Opus 5 to match production exactly.
Answer: A and B. The eval is latency-tolerant and bulk, so Batches plus the cheapest adequate model minimise cost. xhigh (C) inflates cost; streaming (D) does not cut cost; matching production on Opus 5 (E) is unnecessary and expensive for an offline eval.
Q19 · Which TWO statements about the current lineup's context windows are correct? (Select two)
A. Sonnet 5 and Opus 5 each have a 1M context window. B. Haiku 4.5 has a 200k context window. C. Haiku 4.5 has a 1M context window. D. Fable 5.1 has a 128k context window. E. Sonnet 5 has a 200k context window.
Answer: A and B. Sonnet 5 and Opus 5 are 1M; Haiku 4.5 is 200k. Haiku is not 1M (C); Fable 5.1’s 128k is its MAX OUTPUT, not context (D); Sonnet 5 is 1M, not 200k (E).
Key takeaways
- Tokens drive cost, latency and limits; the context window holds input + output,
max_tokenscaps output only. temperature: 0reduces variance but is not byte-deterministic; pin a snapshot for reproducibility.- Start at Sonnet 5; Haiku 4.5 for cheap high-volume, Opus 5 for hard agentic work, Fable 5.1 for the hardest.
budget_tokensis Haiku-4.5-only; effort levels apply to Opus 5 / Sonnet 5 / Fable 5.1.- Know the breaking changes:
budget_tokensremoval, Fable 5.1tool_choice/thinking-binding/append-only, Opus 4.1 retirement. - Cost levers: caching (~10% on hits), batching (50% off), routing/cascading cheap→expensive.
- Latency levers: streaming (perceived), smaller model / shorter output (actual), caching, fast mode.
- Compute cost in your head:
input×price + output×price, then ×0.5 for Batches and ×0.1 for cached-prefix hits; the cheapest adequate model wins. - Pick the minimum thinking effort that clears the bar;
xhighis for the hardest work only,budget_tokensis Haiku-4.5-only, and ‘Haiku +effort’ is invalid. - Rate limits are about token/request volume — fix them with caching, trimming, batching or a higher tier, not by changing model tier.
- Treat a production model choice and an offline eval’s model choice as independent decisions; evals are latency-tolerant, so use the cheapest adequate model via Batches.
Last updated Sep 18, 2026