Appendix · OpenAI
Model Lineup & Pricing (OpenAI)
OpenAI's September 2026 API model line-up, IDs and aliases, reasoning efforts, prices, context windows, specialised model families, a selection table and worked cost calculations.
This is the single reference for OpenAI model IDs, prices and limits used across the OpenAI tracks. Everything here reflects the published API line-up as of 15 September 2026; prices are USD per million tokens (MTok) and are rollout-sensitive, so treat the table as a snapshot to reason with, not a live rate card. Re-verify against the API pricing and API models pages before you commit an architecture or a budget.
API line-up (September 2026)
All latest models take text and image input, emit text output, are multilingual and support vision. They are served through the Responses API and the client SDKs.
| Model | ID | Reasoning efforts | Input $/MTok | Output $/MTok | Context | Max output | Knowledge cutoff |
|---|---|---|---|---|---|---|---|
| GPT-6 Astra | gpt-6-astra | low, medium, high, xhigh, max | $10 | $50 | 1.05M | 128K | 30 Apr 2026 |
| GPT-5.6 Sol | gpt-5.6-sol (alias gpt-5.6) | none…max | $4 | $20 | 1.05M | 128K | 16 Feb 2026 |
| GPT-5.6 Terra | gpt-5.6-terra | none…max | $2 | $12 | 1.05M | 128K | 16 Feb 2026 |
| GPT-5.6 Luna | gpt-5.6-luna | none…max | $0.20 | $1.20 | 1.05M | 128K | 16 Feb 2026 |
Previous generation, still available: gpt-5.5 at $5 / $30 and gpt-5.5-pro at $30 / $180. Legacy models remain callable; new work should target the GPT-5.6 line or Astra.
Aliases move, pinned IDs do not
gpt-5.6 is an alias that resolves to gpt-5.6-sol. Aliases can be repointed by OpenAI; pin the explicit ID (gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, gpt-6-astra) in production so a server-side alias change never silently alters your cost or behaviour.
Reasoning effort
GPT-5.6 models accept efforts from none through max; Astra runs low, medium, high, xhigh, max. The rule straight from the docs is to use the lowest reasoning effort that gets the result — higher effort spends more output tokens (and therefore more money and latency) on internal reasoning. There is no exact mapping from GPT-5.5 efforts to GPT-5.6, so re-tune rather than porting an effort setting across a generation.
| Effort | Spend it on |
|---|---|
none (5.6 only) | Deterministic transforms where reasoning adds nothing |
low | Mechanical, well-specified tasks |
medium | Everyday reasoning; a sensible default |
high | Genuinely hard, open-ended problems |
xhigh | The hardest multi-step work; Astra and 5.6 |
max | Maximum deliberation on one task; use sparingly |
Model positioning
Selection guidance, straight from the docs:
| Signal in the scenario | Choose | Why |
|---|---|---|
| Hardest end-to-end work: sustained reasoning, judgment, multi-tool orchestration | GPT-6 Astra | The frontier model; accept the $10 / $50 price |
| Complex, open-ended, high-value work | GPT-5.6 Sol | Strong reasoning below Astra’s price |
| Pragmatic all-rounder; natural replacement for GPT-5.5 workloads | GPT-5.6 Terra | Balanced cost and quality |
| Clear, repeatable, high-volume tasks: extraction, classification, transformation, structured summaries | GPT-5.6 Luna | Cheapest by an order of magnitude |
| Migrating an existing GPT-5.5 workload | Terra first, benchmark, then decide | Closest cost/quality match |
capability ▲ │ GPT-6 Astra frontier, judgment, multi-tool │ GPT-5.6 Sol complex, open-ended, high-value │ GPT-5.6 Terra pragmatic all-rounder │ GPT-5.6 Luna high-volume, structured, cheap └───────────────────────────────► $ / MTok rises choose the lowest tier that clears the task, then the lowest effort that clears itChatGPT-side model names
The consumer/business product exposes different labels for the same underlying family. Do not assume an API ID exists for a ChatGPT label or vice versa.
| ChatGPT label | Notes |
|---|---|
| GPT-5.6 Sol | Default reasoning-capable model |
| GPT-5.6 Sol Pro | Higher-tier Sol |
| GPT-5.6 Terra | All-rounder |
| GPT-5.6 Luna | Fast, lightweight |
| GPT-5 Thinking Mini | Lightweight thinking model |
| GPT-6 Pro (powered by GPT-6 Astra) | On Pro $100 / Pro $200 / Business / Enterprise |
Legacy models remain available in ChatGPT. The exact model roster a user sees is plan-dependent — see the ChatGPT feature and plan reference.
Specialised model families
These are purpose-built lines, several of them gated to authorised or approved organisations. Never route general traffic to them.
| Family | Models | Purpose / access |
|---|---|---|
| Cyber | gpt-5.6-cyber, gpt-daybreak-red-latest, gpt-daybreak-blue-latest | Authorised security work only |
| Life sciences | gpt-rosalind-research | Life-sciences research, approved orgs |
| Image | gpt-image-2.5-sunburst, gpt-image-2.5-flare | Image generation |
| Realtime / GPT-Live | gpt-live-1, gpt-realtime-2.1, gpt-realtime-2.1-mini, gpt-realtime-2, gpt-realtime-translate, gpt-realtime-1.5 | Low-latency voice / realtime and live translation |
| Transcription | gpt-transcribe, gpt-live-transcribe, gpt-realtime-whisper, gpt-4o-transcribe, gpt-4o-mini-transcribe | Speech-to-text |
| TTS | gpt-4o-mini-tts | Text-to-speech |
Gated families
The Daybreak (red/blue) and gpt-5.6-cyber models are for authorised security work only, and gpt-rosalind-research is limited to approved life-sciences organisations. A scenario that reaches for one of these for ordinary application traffic is choosing the wrong tool.
Selecting a model — decision table
| The scenario says… | Model | Effort | Delivery |
|---|---|---|---|
| “extract fields from millions of documents, cheaply” | Luna | low/none | Batch |
| “customer-facing assistant, balanced” | Terra | medium | Realtime |
| “hard open-ended analysis, high value” | Sol | high | Realtime |
| “the hardest agentic coding / research task we have” | Astra | xhigh/max | Realtime or Agents API |
| “offline nightly job, latency does not matter” | any tier | as needed | Batch (discounted) |
| “the same long system prompt on every call” | any tier | as needed | prompt caching |
| “we are migrating our GPT-5.5 workload” | Terra | benchmark | compare before switching |
Worked cost calculations
Prices are per MTok. These illustrate order-of-magnitude differences; run your own numbers against live pricing before committing.
Example 1 — a single 50k-token RAG answer
A retrieval-augmented answer with 50,000 input tokens (retrieved context + prompt) and 1,000 output tokens, one call.
| Model | Input cost | Output cost | Total |
|---|---|---|---|
| Luna | 50k × $0.20/1M = $0.010 | 1k × $1.20/1M = $0.0012 | ≈ $0.0112 |
| Terra | 50k × $2/1M = $0.100 | 1k × $12/1M = $0.012 | ≈ $0.112 |
| Sol | 50k × $4/1M = $0.200 | 1k × $20/1M = $0.020 | ≈ $0.220 |
| Astra | 50k × $10/1M = $0.500 | 1k × $50/1M = $0.050 | ≈ $0.550 |
The input side dominates a RAG answer because retrieved context is large relative to the reply. Astra costs ~49× Luna for the same answer — so the model choice is a retrieval-quality decision, not a reasoning-power one, unless the synthesis is genuinely hard.
Example 2 — a 1M-token batch job
Classify or transform one million input tokens with 200k tokens of output, run offline.
| Model | Input | Output | Realtime total | With Batch (offline) |
|---|---|---|---|---|
| Luna | 1M × $0.20 = $0.20 | 0.2M × $1.20 = $0.24 | $0.44 | cheapest; Batch lowers it further |
| Terra | 1M × $2 = $2.00 | 0.2M × $12 = $2.40 | $4.40 | reserve for harder items |
| Sol | 1M × $4 = $4.00 | 0.2M × $20 = $4.00 | $8.00 | rarely justified for classification |
For structured high-volume work, Luna on Batch is the default; escalating a subset to Terra only where a validator flags low quality is a cascade, not a blanket upgrade.
Example 3 — an agent session
A single agent session that consumes 300k input tokens (accumulated context, tool results, compaction) and produces 40k output tokens across many turns.
| Model | Input | Output | Total |
|---|---|---|---|
| Terra | 300k × $2/1M = $0.60 | 40k × $12/1M = $0.48 | ≈ $1.08 |
| Sol | 300k × $4/1M = $1.20 | 40k × $20/1M = $0.80 | ≈ $2.00 |
| Astra | 300k × $10/1M = $3.00 | 40k × $50/1M = $2.00 | ≈ $5.00 |
Agent sessions accumulate context, so compaction and context summarisation (see the Agents cheat sheet) are the real cost levers: a session that lets context grow unbounded on Astra is where budgets disappear. Prompt caching a stable system prompt across turns also helps.
Assessment signal
A stem that pairs “cheapest”, “high-volume”, “extraction” or “classification” with a latency-tolerant window is pointing at Luna on Batch. A stem that says “hardest”, “sustained reasoning” or “multi-tool judgment” points at Astra. “Natural replacement for our GPT-5.5 workload” is the Terra tell.
How to re-verify
- Open developers.openai.com/api/docs/models for the current IDs, aliases, context windows and cutoffs.
- Open developers.openai.com/api/docs/pricing for live input/output rates and any Batch/cached-input discounts.
- Check the Codex models page for the Codex-specific roster and retirement dates.
- When a fact here disagrees with the docs, the docs win — this page is a study aid, not an authority.
Key facts to memorise
- Four current API models: Astra ($10/$50), Sol ($4/$20), Terra ($2/$12), Luna ($0.20/$1.20); all 1.05M context, 128K max output.
gpt-5.6aliases togpt-5.6-sol; pin explicit IDs in production.- Use the lowest tier and lowest effort that clears the task; there is no GPT-5.5 → 5.6 effort mapping.
- Terra is the pragmatic replacement for GPT-5.5 workloads; Luna is for high-volume structured tasks.
- Specialised families (cyber, Daybreak, Rosalind) are gated and off-limits for general traffic.
Last updated Sep 18, 2026