# Model Lineup & Pricing (OpenAI)

OpenAI's September 2026 API model line-up, IDs and aliases, reasoning efforts, prices, context windows, specialised model families, a selection table and worked cost calculations.

import { Steps } from '@prosefly/astro-components';

This is the single reference for OpenAI model IDs, prices and limits used across the [OpenAI tracks](/openai/). Everything here reflects the published API line-up as of **15 September 2026**; prices are USD per million tokens (MTok) and are rollout-sensitive, so treat the table as a snapshot to reason with, not a live rate card. Re-verify against the [API pricing](https://developers.openai.com/api/docs/pricing) and [API models](https://developers.openai.com/api/docs/models) pages before you commit an architecture or a budget.

## API line-up (September 2026)

All latest models take text and image input, emit text output, are multilingual and support vision. They are served through the **Responses API** and the client SDKs.

| Model | ID | Reasoning efforts | Input $/MTok | Output $/MTok | Context | Max output | Knowledge cutoff |
| --- | --- | --- | --- | --- | --- | --- | --- |
| GPT-6 Astra | `gpt-6-astra` | low, medium, high, xhigh, max | $10 | $50 | 1.05M | 128K | 30 Apr 2026 |
| GPT-5.6 Sol | `gpt-5.6-sol` (alias `gpt-5.6`) | none…max | $4 | $20 | 1.05M | 128K | 16 Feb 2026 |
| GPT-5.6 Terra | `gpt-5.6-terra` | none…max | $2 | $12 | 1.05M | 128K | 16 Feb 2026 |
| GPT-5.6 Luna | `gpt-5.6-luna` | none…max | $0.20 | $1.20 | 1.05M | 128K | 16 Feb 2026 |

**Previous generation, still available:** `gpt-5.5` at $5 / $30 and `gpt-5.5-pro` at $30 / $180. Legacy models remain callable; new work should target the GPT-5.6 line or Astra.

:::note[Aliases move, pinned IDs do not]
`gpt-5.6` is an alias that resolves to `gpt-5.6-sol`. Aliases can be repointed by OpenAI; pin the explicit ID (`gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-6-astra`) in production so a server-side alias change never silently alters your cost or behaviour.
:::

## Reasoning effort

GPT-5.6 models accept efforts from `none` through `max`; Astra runs `low, medium, high, xhigh, max`. The rule straight from the docs is to **use the lowest reasoning effort that gets the result** — higher effort spends more output tokens (and therefore more money and latency) on internal reasoning. There is **no exact mapping** from GPT-5.5 efforts to GPT-5.6, so re-tune rather than porting an effort setting across a generation.

| Effort | Spend it on |
| --- | --- |
| `none` (5.6 only) | Deterministic transforms where reasoning adds nothing |
| `low` | Mechanical, well-specified tasks |
| `medium` | Everyday reasoning; a sensible default |
| `high` | Genuinely hard, open-ended problems |
| `xhigh` | The hardest multi-step work; Astra and 5.6 |
| `max` | Maximum deliberation on one task; use sparingly |

## Model positioning

Selection guidance, straight from the docs:

| Signal in the scenario | Choose | Why |
| --- | --- | --- |
| Hardest end-to-end work: sustained reasoning, judgment, multi-tool orchestration | **GPT-6 Astra** | The frontier model; accept the $10 / $50 price |
| Complex, open-ended, high-value work | **GPT-5.6 Sol** | Strong reasoning below Astra's price |
| Pragmatic all-rounder; natural replacement for GPT-5.5 workloads | **GPT-5.6 Terra** | Balanced cost and quality |
| Clear, repeatable, high-volume tasks: extraction, classification, transformation, structured summaries | **GPT-5.6 Luna** | Cheapest by an order of magnitude |
| Migrating an existing GPT-5.5 workload | **Terra** first, benchmark, then decide | Closest cost/quality match |

```text
                 capability ▲
                            │  GPT-6 Astra      frontier, judgment, multi-tool
                            │  GPT-5.6 Sol      complex, open-ended, high-value
                            │  GPT-5.6 Terra    pragmatic all-rounder
                            │  GPT-5.6 Luna     high-volume, structured, cheap
                            └───────────────────────────────► $ / MTok rises
   choose the lowest tier that clears the task, then the lowest effort that clears it
```

## ChatGPT-side model names

The consumer/business product exposes different labels for the same underlying family. Do not assume an API ID exists for a ChatGPT label or vice versa.

| ChatGPT label | Notes |
| --- | --- |
| GPT-5.6 Sol | Default reasoning-capable model |
| GPT-5.6 Sol Pro | Higher-tier Sol |
| GPT-5.6 Terra | All-rounder |
| GPT-5.6 Luna | Fast, lightweight |
| GPT-5 Thinking Mini | Lightweight thinking model |
| GPT-6 Pro (powered by GPT-6 Astra) | On Pro $100 / Pro $200 / Business / Enterprise |

Legacy models remain available in ChatGPT. The exact model roster a user sees is **plan-dependent** — see the [ChatGPT feature and plan reference](/appendix/openai/chatgpt-features/).

## Specialised model families

These are purpose-built lines, several of them gated to authorised or approved organisations. Never route general traffic to them.

| Family | Models | Purpose / access |
| --- | --- | --- |
| Cyber | `gpt-5.6-cyber`, `gpt-daybreak-red-latest`, `gpt-daybreak-blue-latest` | Authorised security work only |
| Life sciences | `gpt-rosalind-research` | Life-sciences research, approved orgs |
| Image | `gpt-image-2.5-sunburst`, `gpt-image-2.5-flare` | Image generation |
| Realtime / GPT-Live | `gpt-live-1`, `gpt-realtime-2.1`, `gpt-realtime-2.1-mini`, `gpt-realtime-2`, `gpt-realtime-translate`, `gpt-realtime-1.5` | Low-latency voice / realtime and live translation |
| Transcription | `gpt-transcribe`, `gpt-live-transcribe`, `gpt-realtime-whisper`, `gpt-4o-transcribe`, `gpt-4o-mini-transcribe` | Speech-to-text |
| TTS | `gpt-4o-mini-tts` | Text-to-speech |

:::caution[Gated families]
The Daybreak (`red`/`blue`) and `gpt-5.6-cyber` models are for authorised security work only, and `gpt-rosalind-research` is limited to approved life-sciences organisations. A scenario that reaches for one of these for ordinary application traffic is choosing the wrong tool.
:::

## Selecting a model — decision table

| The scenario says… | Model | Effort | Delivery |
| --- | --- | --- | --- |
| "extract fields from millions of documents, cheaply" | Luna | `low`/`none` | Batch |
| "customer-facing assistant, balanced" | Terra | `medium` | Realtime |
| "hard open-ended analysis, high value" | Sol | `high` | Realtime |
| "the hardest agentic coding / research task we have" | Astra | `xhigh`/`max` | Realtime or Agents API |
| "offline nightly job, latency does not matter" | any tier | as needed | **Batch** (discounted) |
| "the same long system prompt on every call" | any tier | as needed | **prompt caching** |
| "we are migrating our GPT-5.5 workload" | Terra | benchmark | compare before switching |

## Worked cost calculations

Prices are per MTok. These illustrate order-of-magnitude differences; run your own numbers against live pricing before committing.

### Example 1 — a single 50k-token RAG answer

A retrieval-augmented answer with 50,000 input tokens (retrieved context + prompt) and 1,000 output tokens, one call.

| Model | Input cost | Output cost | Total |
| --- | --- | --- | --- |
| Luna | 50k × $0.20/1M = $0.010 | 1k × $1.20/1M = $0.0012 | **≈ $0.0112** |
| Terra | 50k × $2/1M = $0.100 | 1k × $12/1M = $0.012 | **≈ $0.112** |
| Sol | 50k × $4/1M = $0.200 | 1k × $20/1M = $0.020 | **≈ $0.220** |
| Astra | 50k × $10/1M = $0.500 | 1k × $50/1M = $0.050 | **≈ $0.550** |

The input side dominates a RAG answer because retrieved context is large relative to the reply. Astra costs ~49× Luna for the same answer — so the model choice is a retrieval-quality decision, not a reasoning-power one, unless the synthesis is genuinely hard.

### Example 2 — a 1M-token batch job

Classify or transform one million input tokens with 200k tokens of output, run offline.

| Model | Input | Output | Realtime total | With Batch (offline) |
| --- | --- | --- | --- | --- |
| Luna | 1M × $0.20 = $0.20 | 0.2M × $1.20 = $0.24 | **$0.44** | cheapest; Batch lowers it further |
| Terra | 1M × $2 = $2.00 | 0.2M × $12 = $2.40 | **$4.40** | reserve for harder items |
| Sol | 1M × $4 = $4.00 | 0.2M × $20 = $4.00 | **$8.00** | rarely justified for classification |

For structured high-volume work, Luna on Batch is the default; escalating a subset to Terra only where a validator flags low quality is a cascade, not a blanket upgrade.

### Example 3 — an agent session

A single agent session that consumes 300k input tokens (accumulated context, tool results, compaction) and produces 40k output tokens across many turns.

| Model | Input | Output | Total |
| --- | --- | --- | --- |
| Terra | 300k × $2/1M = $0.60 | 40k × $12/1M = $0.48 | **≈ $1.08** |
| Sol | 300k × $4/1M = $1.20 | 40k × $20/1M = $0.80 | **≈ $2.00** |
| Astra | 300k × $10/1M = $3.00 | 40k × $50/1M = $2.00 | **≈ $5.00** |

Agent sessions accumulate context, so **compaction and context summarisation** (see the [Agents cheat sheet](/appendix/openai/agents-cheatsheet/)) are the real cost levers: a session that lets context grow unbounded on Astra is where budgets disappear. Prompt caching a stable system prompt across turns also helps.

:::tip[Assessment signal]
A stem that pairs "cheapest", "high-volume", "extraction" or "classification" with a latency-tolerant window is pointing at **Luna on Batch**. A stem that says "hardest", "sustained reasoning" or "multi-tool judgment" points at **Astra**. "Natural replacement for our GPT-5.5 workload" is the Terra tell.
:::

## How to re-verify

<Steps>
1. Open [developers.openai.com/api/docs/models](https://developers.openai.com/api/docs/models) for the current IDs, aliases, context windows and cutoffs.
2. Open [developers.openai.com/api/docs/pricing](https://developers.openai.com/api/docs/pricing) for live input/output rates and any Batch/cached-input discounts.
3. Check the [Codex models](https://developers.openai.com/codex/models) page for the Codex-specific roster and retirement dates.
4. When a fact here disagrees with the docs, the docs win — this page is a study aid, not an authority.
</Steps>

## Key facts to memorise

- Four current API models: **Astra** ($10/$50), **Sol** ($4/$20), **Terra** ($2/$12), **Luna** ($0.20/$1.20); all 1.05M context, 128K max output.
- `gpt-5.6` aliases to `gpt-5.6-sol`; pin explicit IDs in production.
- Use the **lowest tier and lowest effort** that clears the task; there is no GPT-5.5 → 5.6 effort mapping.
- Terra is the pragmatic replacement for GPT-5.5 workloads; Luna is for high-volume structured tasks.
- Specialised families (cyber, Daybreak, Rosalind) are gated and off-limits for general traffic.
