# Model Lineup & Pricing

Current Claude models, IDs, limits, prices, breaking changes and worked cost calculations as of September 2026.

import { Steps } from '@prosefly/astro-components';

## Current lineup (September 2026)

| Model | ID | Context | Max output | Input / Output per MTok | Cache read | Positioning |
| --- | --- | --- | --- | --- | --- | --- |
| Claude Fable 5.1 | `claude-fable-5-1` | 1M | 128k | $10 / $50 | $0.25 | Most capable; hardest reasoning and agentic work |
| Claude Opus 5 | `claude-opus-5` | 1M | 128k | $5 / $25 | $0.50 | Default for complex agentic coding and enterprise |
| Claude Sonnet 5 | `claude-sonnet-5` | 1M | 128k | $2 / $10 | $0.20 | Best speed / intelligence balance |
| Claude Haiku 4.5 | `claude-haiku-4-5` | 200k | 64k | $1 / $5 | $0.10 | Fastest and cheapest; classification, routing, extraction |

Cache writes cost ≈ 1.25× base input (5-minute TTL) or 2× (1-hour TTL). Batch API: 50% off input and output.

**Legacy, still available:** Fable 5, Opus 4.8, 4.7, 4.6, 4.5, Sonnet 4.6, 4.5. **Retired:** Opus 4.1 (2026-08-05).

**Published retirement floors:** Fable 5.1 not before 2027-09-01 · Opus 5 not before 2027-07-24 · Sonnet 5 not before 2027-06-30 · Haiku 4.5 not before 2026-10-15.

:::note[Verify live]
Model IDs from the 4.6 generation onward are dateless but still pinned snapshots. Query `client.models.list()` and `client.models.retrieve(id)` for live `max_input_tokens`, `max_tokens` and `capabilities` rather than trusting any static table – including this one.
:::

## Capability matrix

| Capability | Fable 5.1 | Opus 5 | Sonnet 5 | Haiku 4.5 |
| --- | --- | --- | --- | --- |
| Adaptive thinking `{"type":"adaptive"}` | ✓ (always on) | ✓ | ✓ | – |
| `budget_tokens` thinking | 400 error | 400 error | 400 error | ✓ (only model) |
| `effort` low/medium/high/xhigh | ✓ | ✓ | ✓ (no xhigh benefit) | – |
| `tool_choice: "any"` / forced tool | **400 error** | ✓ | ✓ | ✓ |
| Structured outputs `output_config.format` | ✓ | ✓ | ✓ | ✓ |
| `strict: true` tool schemas | ✓ | ✓ | ✓ | ✓ |
| Mid-conversation `role: "system"` messages | ✓ | ✓ | – | – |
| Task budgets | ✓ | ✓ | – | – |
| Thinking blocks portable to other models | – (bound) | ✓ | ✓ | n/a |
| Zero Data Retention | – (30-day required) | ✓ | ✓ | ✓ |
| Priority Tier | – | – | – | ✓ |
| Prompt caching | ✓ | ✓ | ✓ | ✓ (2048-token minimum) |
| Batch API | ✓ | ✓ | ✓ | ✓ |

## Fable 5.1 breaking changes (exam favourites)

1. **No forced tool use.** `tool_choice: "any"` and `{"type": "tool", "name": …}` return 400. Alternatives: `auto` plus an explicit instruction; `strict: true` on the tool schema; `output_config.format` structured output.
2. **Thinking-block binding.** Thinking blocks are readable only by the model that produced them or a newer one. A fallback or router switch to an older model silently drops them (unbilled) and the older model re-plans from scratch.
3. **Append-only history.** Editing, reordering or removing earlier turns invalidates every later thinking block. Freeze `system` and `tools`, deliver mid-session instruction changes as `role: "system"` messages, and trim context server-side (context editing / compaction) rather than on the client.
4. **Retention.** Requires 30-day data retention; not available to ZDR organisations unless authorised. Excluded from Priority Tier (as are Opus 5 and Sonnet 5).

## Selecting a model – decision table

| Signal in the scenario | Choose | Why |
| --- | --- | --- |
| "Classify", "route", "extract simple fields", "millions of items", "sub-second" | Haiku 4.5 | Cheapest, fastest; quality sufficient for narrow tasks |
| "General assistant", "customer-facing chat", "balanced cost and quality" | Sonnet 5 | Best trade-off; 1M context |
| "Complex agentic coding", "multi-step reasoning", "enterprise default" | Opus 5 | Default for hard agentic work |
| "Hardest research", "frontier reasoning", budget is secondary | Fable 5.1 | Most capable; accept breaking-change constraints |
| "Cost matters, latency does not" | Any tier **+ Batch API** | 50% discount |
| Same long prefix reused | Any tier **+ prompt caching** | ~90% saving on cached reads |
| Mixed difficulty stream | **Cascade**: Haiku → Sonnet → Opus | Cheap model handles the bulk; escalate on low confidence *from a validator*, not self-report |

## Worked cost calculations

### Example 1 – 10,000 documents, 3,000 input tokens each, 500 output tokens each

| Approach | Input cost | Output cost | Total |
| --- | --- | --- | --- |
| Haiku 4.5 realtime | 30M × $1 = $30 | 5M × $5 = $25 | **$55** |
| Sonnet 5 realtime | 30M × $2 = $60 | 5M × $10 = $50 | **$110** |
| Opus 5 realtime | 30M × $5 = $150 | 5M × $25 = $125 | **$275** |
| Sonnet 5 Batch (50%) | $30 | $25 | **$55** |
| Opus 5 Batch (50%) | $75 | $62.50 | **$137.50** |

### Example 2 – prompt caching on a 20,000-token system prompt + tools, 1,000 requests/day on Sonnet 5

| | Without caching | With caching (5-min TTL, ~1 write per 5 min ≈ 288 writes/day) |
| --- | --- | --- |
| Prefix tokens billed | 20M at $2 = **$40/day** | Writes: 288 × 20k × $2.50 = $14.40; Reads: 712 × 20k × $0.20 = $2.85 → **≈ $17.25/day** |
| Saving | – | ≈ 57% on the prefix; approaches 90% as request rate rises |

Rule: caching pays back after roughly two reads of the same prefix within the TTL.

### Example 3 – choosing effort

| Task | Effort | Rationale |
| --- | --- | --- |
| Subagent that renames files per a spec | `low` | Mechanical |
| Summarise a 50-page contract | `medium` | Moderate reasoning, cost-sensitive |
| Default production reasoning | `high` | Baseline |
| Multi-hour refactor across 40 files (Opus 5 / Fable 5.1) | `xhigh` | Hardest agentic work |

## Cloud availability

| Platform | Notes |
| --- | --- |
| Claude API (direct) | Full feature surface first; simplest |
| Amazon Bedrock | AWS IAM, data residency, FedRAMP High, existing AWS spend |
| Google Vertex AI | GCP IAM, residency, FedRAMP High |
| Microsoft Foundry | Azure ecosystem |

Feature availability can lag on cloud platforms; check per-feature docs before committing an architecture.

## Feature availability by cloud platform

The exam tests the reflex that **the direct Claude API gets new features first**, and that a governance or residency requirement can force a platform that lags on a feature you depend on. Treat this as a snapshot to reason with, not a live SLA.

| Feature | Claude API | Bedrock | Vertex AI | Foundry |
| --- | --- | --- | --- | --- |
| Newest model on launch day | ✓ first | Usually days–weeks later | Usually days–weeks later | Varies |
| Prompt caching | ✓ | ✓ | ✓ | Check |
| Message Batches | ✓ | ✓ (batch inference) | ✓ (batch prediction) | Check |
| MCP connector (server-side) | ✓ | Lags | Lags | Lags |
| Files API / citations | ✓ | Partial | Partial | Check |
| Structured outputs `output_config.format` | ✓ | ✓ | ✓ | Check |
| Context editing / compaction | ✓ | Lags | Lags | Lags |
| FedRAMP High authorisation | – | ✓ | ✓ | – |
| Data residency controls | Limited | ✓ (region) | ✓ (region) | ✓ (region) |
| IAM / enterprise identity | API keys | AWS IAM | GCP IAM | Entra ID |

:::tip[Exam signal]
"We are on AWS with FedRAMP High and IAM already" → Bedrock, even if it means waiting for a feature. "We need the MCP connector and context editing today" → direct Claude API. A stem that pairs a residency requirement with a bleeding-edge feature is testing whether you notice the platform lag.
:::

## Rate-limit tiers and dimensions

Rate limits apply on three dimensions simultaneously; you hit whichever binds first.

| Dimension | Meaning | Typical binding case |
| --- | --- | --- |
| RPM | Requests per minute | Many small classification calls |
| ITPM | Input tokens per minute | Large prompts / long context / uncached prefixes |
| OTPM | Output tokens per minute | Long generations, streaming many tokens |

| Tier | How you move up | Notes |
| --- | --- | --- |
| Tier 1–4 (standard) | Automatic with usage + payment history | Higher tiers raise RPM/ITPM/OTPM |
| Custom / enterprise | Sales agreement | Committed throughput |
| Priority Tier | Reserved capacity for latency-sensitive traffic | **Excludes Fable 5.1, Opus 5, Sonnet 5**; Haiku 4.5 eligible |

:::tip[Exam signal]
A 429 with lots of cached input still counts cached-read tokens against ITPM at a reduced rate but writes count in full — caching lowers cost more than it lowers ITPM pressure. If a stem says "we keep hitting 429 on a 300k-token prompt at low request volume", the binding limit is ITPM, not RPM: shrink the prompt, cache the prefix, or raise the tier.
:::

## Deprecation and migration lifecycle

Model retirement is a **planned lifecycle event**, not an emergency. The exam-correct posture: watch deprecation notices, keep an eval set, migrate behind that eval, pin IDs in production.

```text
Announced → Deprecated (still callable) → Retirement floor date → Retired (404)
   │              │                              │
   │              └─ start eval-gated migration  └─ published earliest date; often extended
   └─ appears in models.list() deprecation metadata / dashboard
```

| Reasoning step | What to do |
| --- | --- |
| Notice arrives | Read the retirement floor date; it is the earliest, not a promise |
| Pick the target | Newer-or-equal model; check the capability matrix for breaking changes |
| Guard the migration | Run the existing golden set on the new model; compare per-segment, not aggregate |
| Handle thinking blocks | Migrating **up** (older→newer) is safe; migrating a Fable 5.1 flow **down** drops thinking blocks |
| Cut over | Change the pinned ID; keep the old ID available for rollback until the floor date |

### Per-model migration checklists

<Steps>
1. **To Haiku 4.5** — remove any `effort` parameter (unsupported → 400); if you relied on `budget_tokens`, this is the *only* current model that keeps it, so a downgrade from adaptive-thinking models is where `budget_tokens` reappears. Re-tune prompts for the 200k/64k window; verify classification/extraction accuracy on the golden set.
2. **To Sonnet 5** — drop any mid-conversation `role: "system"` messages and task budgets (unsupported); confirm `xhigh` effort is not assumed (no benefit). Re-check latency budgets; it is the balanced default.
3. **To Opus 5** — safe target for most agentic upgrades; `xhigh` available. Remove `budget_tokens` (400) in favour of `{"type":"adaptive"}`. Re-baseline cost — 2.5× Sonnet input.
4. **To Fable 5.1** — the highest-friction migration. Remove forced `tool_choice` (`any`/named → 400); switch to `auto`+instruction, `strict: true`, or `output_config.format`. Freeze `system`/`tools` for append-only history. Confirm the org is **not** ZDR (30-day retention required) and does not depend on Priority Tier. Re-baseline cost at $10/$50.
</Steps>

## More worked cost scenarios

### Example 4 – caching + Batch combined

Nightly enrichment of 50,000 records on Sonnet 5. Each request: 12,000-token shared instruction/schema prefix (cacheable) + 800 unique input tokens + 300 output tokens. Run as one Batch job.

| Cost component | Calculation | Cost |
| --- | --- | --- |
| Cached prefix reads (Batch, 50% off the 0.1× read) | 50k × 12k × ($2 × 0.1 × 0.5)/1M = 600M tok × $0.10 | $60.00 |
| One cache write (first request seeds it) | 12k × ($2 × 1.25)/1M | $0.03 |
| Unique input (Batch 50%) | 50k × 800 × ($2 × 0.5)/1M = 40M × $1.00 | $40.00 |
| Output (Batch 50%) | 50k × 300 × ($10 × 0.5)/1M = 15M × $5.00 | $75.00 |
| **Total** | | **≈ $175.03** |

Naïve Sonnet 5 realtime, no cache: prefix 50k×12k=600M×$2=$1,200 + input 40M×$2=$80 + output 15M×$10=$150 = **$1,430**. Caching+Batch cuts it ≈ 88%. The prefix dominates, so caching it matters far more than the discount on the small unique portion.

:::caution[Cache scope in Batch]
Cache hits require the prefix to be identical and within the TTL window. In a large Batch the writes/reads interleave across the job; budget for a handful of writes, not one. The dominant saving still comes from reading the 12k prefix 50,000 times at 0.1×.
:::

### Example 5 – cascade routing math

A stream of 100,000 tickets. A Haiku 4.5 first pass answers 100% (1,500 in / 400 out each). An **external validator** flags 18% as low-confidence; those escalate to Opus 5 (same tokens). Compare to sending everything to Opus 5.

| Path | Input | Output | Cost |
| --- | --- | --- | --- |
| Haiku pass (all 100k) | 150M × $1 = $150 | 40M × $5 = $200 | $350 |
| Opus escalation (18k) | 27M × $5 = $135 | 7.2M × $25 = $180 | $315 |
| **Cascade total** | | | **$665** |
| Opus-only (all 100k) | 150M × $5 = $750 | 40M × $25 = $1,000 | **$1,750** |

Cascade saves ≈ 62%. The break-even escalation rate `e` where cascade = Opus-only solves `350 + 1750e = 1750` → `e ≈ 80%`. Below ~80% escalation, cascade wins; above it, the Haiku pass is pure overhead — send everything to Opus.

:::tip[Exam signal]
Escalation must be triggered by an **external validator or a downstream check**, never by the cheap model's self-reported confidence (self-report reliance, anti-pattern 4). A stem that routes on "if Haiku says it is unsure" is the distractor; "if a validator rejects the answer" is correct.
:::

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Fable 5.1 is just a bigger Opus 5, use it everywhere" | It is most capable but $10/$50, no forced tools, no ZDR, not Priority Tier | Cost-blind and constraint-blind distractor |
| "Prompt caching makes requests free after the first" | Reads are 0.1×, not 0; writes cost 1.25×/2×; TTL expires | Overstates savings; break-even ≈ 2 reads |
| "Batch is slower so it is worse" | Batch trades latency for 50% cost; ideal for offline/overnight | Latency-blind vs cost-sensitive stems |
| "Temperature 0 makes Claude deterministic and correct" | Reduces variance, not error; not a truth switch | Confuses variance with accuracy |
| "Pick the newest model to be safe" | Newest can add breaking changes (Fable 5.1) and cost 5–10× | Over-engineered / cost-blind |
| "Rate limits are just requests per minute" | RPM, ITPM and OTPM bind independently | A large-prompt 429 is ITPM, not RPM |
| "A retirement date means the model dies then" | It is the *earliest* floor; migrate behind an eval before it | Panic-migration distractor |

## Scenario walkthrough

A fintech runs a document-classification pipeline on 2M PDFs/month. Requirements: US-federal customer (FedRAMP High), personal data (needs regional residency and, ideally, ZDR), cost-sensitive, latency non-critical (nightly), classification quality must be measured per document type before switching models.

Expert reasoning trace:

<Steps>
1. **Residency + FedRAMP High** rules out the direct API for the regulated workload → Bedrock or Vertex. Pick the one matching existing cloud spend/IAM.
2. **ZDR desired + cheap + high-volume classification** → Haiku 4.5. It is Priority-Tier eligible (irrelevant here, latency non-critical) and supports ZDR. Fable 5.1 is *rejected*: no ZDR, 10× the price, and classification does not need frontier reasoning (constraint-blind + over-engineered).
3. **Latency non-critical, 2M/month** → Batch API (50% off). Realtime is rejected as needlessly expensive (cost-blind).
4. **Shared schema/instruction prefix** → prompt caching on the stable prefix; dynamic PDF last.
5. **"Measured per document type"** → per-segment eval on a golden set, not aggregate accuracy (aggregate-metric distractor).
6. **Model migration later** → pin `claude-haiku-4-5`, keep the eval set, migrate behind it when a successor ships.
</Steps>

Correct architecture: Bedrock + Haiku 4.5 + Batch + prompt caching + per-segment evals. Each tempting alternative (Fable everywhere, realtime for speed, aggregate accuracy, self-reported confidence routing) maps to a named distractor pattern.
