Codex Path
D1 · Codex Fundamentals and Surfaces
The Codex surfaces and when to use each, the models available in Codex and how to control them, the reasoning ladder from Low to Ultra, and the September 2026 model retirements.
This domain is worth 20% of the mock — roughly 10 of 50 items. It tests whether you can place a coding task on the right Codex surface, pick a model and reasoning effort that fit the work, and reason about the model line-up and its recent changes. Almost every item comes down to one question: given this task, this environment and this deadline, which surface, model and effort do I reach for?
What you need to know
Codex is one agent behind many surfaces: the ChatGPT desktop app, ChatGPT Work on the web, the Codex CLI, the Codex IDE extension, Codex cloud, and Codex Micro, plus integrations into Slack, GitHub, GitLab (beta) and Linear. The surfaces share one configuration file, so the choice between them is about where the work happens, not about capability. Inside Codex you can run several models — gpt-6-astra, gpt-5.6-sol, gpt-5.6-terra, gpt-5.6-luna, and the gpt-5.3-codex-spark research preview — and control them with /model interactively or -m / --model on the command line. A reasoning ladder from Low to Ultra trades thinking time and parallelism for speed and cost. As of 31 August 2026 the gpt-5.4 family retired from Codex under ChatGPT sign-in, so knowing the current replacements matters.
Learning objectives
By the end of this page you should be able to:
- Select the right Codex surface for a task from the desktop app, ChatGPT Work web, CLI, IDE extension, cloud and Micro.
- Choose a Codex model and switch it with
/model,-mor--modelfor both interactive and non-interactive runs. - Apply the reasoning ladder Low → Medium → High → Extra High → Max → Ultra, and explain the Max-versus-Ultra distinction and subagents.
- Reason about the model line-up, including the retirement of
gpt-5.4/gpt-5.4-minifrom Codex and their replacements. - Distinguish interactive, non-interactive and cloud execution and match each to a task shape.
1.1 The Codex surfaces
Codex is the same agent everywhere; the surface decides where it runs and how you interact with it.
| Surface | What it is | Reach for it when |
|---|---|---|
| ChatGPT desktop app | Codex embedded in the desktop ChatGPT client | You want a chat-driven coding session on your machine alongside other ChatGPT work |
| ChatGPT Work (web) | Codex inside ChatGPT Work in the browser | You are on a managed workspace and want browser access with workspace controls |
| Codex CLI | A terminal agent (codex, codex exec) | You live in the terminal, want to script runs, or need non-interactive automation in CI |
| Codex IDE extension | Codex inside your editor | You want a tight edit-review loop with the diff in front of you in the IDE |
| Codex cloud | Codex running in a hosted sandbox | Long-running tasks, parallel workstreams, or work you do not want tying up your laptop |
| Codex Micro | A lightweight Codex surface | Small, quick, low-overhead tasks |
Assessment signal
Stems that say terminal, script, CI, or non-interactive point to the CLI. In my editor, see the diff as I go points to the IDE extension. Long-running, overnight, several tasks at once, without tying up my machine points to Codex cloud.
ONE CODEX AGENT one shared config.toml ┌──────────┬───────────┬──────────┬───────────┬─────────┬────────┐ │ Desktop │ ChatGPT │ CLI │ IDE │ Cloud │ Micro │ │ app │ Work web │ codex(1) │ extension │ sandbox │ │ └────┬─────┴─────┬─────┴────┬─────┴─────┬─────┴────┬────┴───┬────┘ │ │ │ │ │ │ chat on managed scripted tight edit long / quick, desktop browser & CI runs loop, diff parallel small1.2 Interactive, non-interactive and cloud execution
The same task can run three ways, and the exam expects you to match the shape of the work to the mode.
| Mode | How you invoke it | Best for |
|---|---|---|
| Interactive | codex (opens a session), or the IDE / desktop chat | Exploratory work, review-first loops, anything where you steer mid-task |
| Non-interactive | codex exec 'task' | Scripted, repeatable, unattended runs — CI jobs, batch fixes, scheduled work |
| Cloud | Codex cloud in ChatGPT / Work | Long-running or parallel work in a hosted sandbox, independent of your machine |
Worked example — the same “fix the failing test” task, three ways:
codex# then, in the session:# > fix the failing test in the payments module and show me the diffYou can steer, approve escalations and inspect the diff before anything is written.
codex exec -m gpt-5.6-terra "fix the failing test in the payments module"One shot, no prompts. Ideal in CI or a script; pair it with a sandbox and a permission mode so it cannot escalate silently.
Codex cloud → New task → "fix the failing test in the payments module"Runs in a hosted sandbox; follow progress, come back to the diff and evidence later.Right when the task is long-running or you want to start several at once.
1.3 The models available in Codex
Codex can run the current model line-up plus a Codex-specific research preview. Reach for the lowest-cost model that meets the task.
| Model | ID | Reach for it when |
|---|---|---|
| GPT-6 Astra | gpt-6-astra | The hardest end-to-end work — sustained reasoning, judgment, multi-tool, large refactors |
| GPT-5.6 Sol | gpt-5.6-sol (alias gpt-5.6) | Complex, open-ended, high-value coding work |
| GPT-5.6 Terra | gpt-5.6-terra | The pragmatic all-rounder; the natural replacement for GPT-5.5 workloads |
| GPT-5.6 Luna | gpt-5.6-luna | Clear, repeatable, high-volume tasks — small fixes, mechanical edits |
| GPT-5.3 Codex Spark | gpt-5.3-codex-spark | Near-instant iteration; text-only research preview, ChatGPT Pro |
gpt-5.5 and gpt-5.4 remain listed as “other models”. Use the model lineup appendix as the single source for prices and limits.
Assessment signal
Hardest, most complex end-to-end → Astra. High-value but open-ended → Sol. Replace my GPT-5.5 workload, sensible default → Terra. High-volume, mechanical, cheapest that works → Luna. Do not pick Astra by default; the exam rewards the lowest model that meets the task.
1.4 Controlling the model: /model, -m and --model
You choose the model per session, per command, or as a shared default.
# Interactive: switch mid-sessioncodex# > /model (opens the model picker)
# Launch on a specific modelcodex --model gpt-5.6codex -m gpt-5.6-terra
# Non-interactive on a specific modelcodex exec -m gpt-5.6 "refactor the auth middleware and run the tests"The shared default lives in config.toml:
model = "gpt-5.6"| Control | Scope | Use it for |
|---|---|---|
config.toml model = '…' | Every surface, by default | Team-wide or personal default |
codex --model / codex -m | This launch | One-off override |
/model in-session | The rest of this session | Switching after seeing how a task is going |
codex exec -m | This non-interactive run | Scripted or CI runs |
1.5 The reasoning ladder: Low to Ultra
Codex exposes a ladder of reasoning effort in the CLI. Higher effort means more thinking time (and cost); the top rung adds parallelism.
Low ─► Medium ─► High ─► Extra High ─► Max ─► Ultra(fast, (default) (harder) (deep) │ │ cheap) │ │ MORE THINKING DELEGATES TO ON ONE TASK SUBAGENTS IN (Max) PARALLEL (Ultra)- Low — quick, cheap; mechanical edits and simple lookups.
- Medium — the default; most day-to-day work.
- High — harder tasks needing more careful reasoning.
- Extra High — deep, multi-step reasoning on one hard task.
- Max — more thinking time on a single task; use when one hard problem needs the model to reason longer, not when the work splits into pieces.
- Ultra — automatic delegation to subagents in parallel; use when a task genuinely decomposes into independent sub-tasks that can run at once.
The Max-versus-Ultra distinction is a favourite item: Max deepens a single line of reasoning, Ultra spreads work across subagents. In GUI clients the levels read Light / Medium / High / Extra High, and the Astra rollout exposes Power options such as Terra Light, Sol Light, Sol Medium, Astra Light, Astra Medium and Astra Extra High. If Ultra is missing from the picker, enable it via Settings → Configuration → “Ultra in model picker slider”.
Assessment signal
One hard problem, think longer → Max. Several independent parts, run at once, parallel subagents → Ultra. Ultra is not in my picker → enable the Ultra slider in Settings → Configuration.
1.6 The September 2026 model retirements
Model availability in Codex changed under ChatGPT sign-in, and the exam tests the replacements, not the drama.
| Retired / deprecated (ChatGPT sign-in) | Date | Replace with |
|---|---|---|
gpt-5.4 | 31 Aug 2026 | gpt-5.6-terra |
gpt-5.4-mini | 31 Aug 2026 | gpt-5.6-luna |
gpt-5.2 | already deprecated | current line-up |
gpt-5.3-codex | already deprecated | current line-up |
Two facts matter for judgment items: the retirement applies to ChatGPT sign-in, and API-key sign-in is unaffected. So a team pinned to gpt-5.4 in a shared config that signs in with ChatGPT must migrate to Terra/Luna; a service that authenticates with an API key does not face the same deadline.
Assessment signal
gpt-5.4 stopped working, our shared config broke on 1 September → the ChatGPT-sign-in retirement; move to gpt-5.6-terra (or -luna for the mini). We sign in with an API key → unaffected.
Decision framework
Use SURFACE-MODEL-EFFORT (SME) to place any task in three moves.
| Step | Question | Rule |
|---|---|---|
| Surface | Where does the work happen? | Terminal/CI → CLI (codex exec if unattended); in-editor loop → IDE extension; long/parallel → cloud; chat on desktop → desktop app; managed browser → ChatGPT Work |
| Model | How hard is the task? | Mechanical/high-volume → Luna; sensible default / GPT-5.5 replacement → Terra; complex, high-value → Sol; hardest end-to-end → Astra |
| Effort | Does it need more thinking or more parallelism? | Simple → Low/Medium; hard single problem → High/Extra High/Max; decomposable → Ultra (subagents) |
Applied: “a scripted overnight batch that renames a deprecated API across many files” → CLI with codex exec (surface), Luna (model — mechanical, high-volume), Low or Medium effort (it is not hard, just large). “redesign the caching layer, one gnarly problem” → IDE extension, Sol or Astra, Max effort.
Common mistakes
| Mistake | Why it happens | What to do instead |
|---|---|---|
| Defaulting to Astra for everything | It is the strongest model, so it feels “safe” | Pick the lowest model that meets the task; Luna/Terra handle most work at a fraction of the cost |
| Using Ultra for a single hard problem | Confusing “harder” with “more parallel” | Use Max for depth on one task; Ultra only when work decomposes into independent parts |
| Running unattended work interactively | Habit of using codex for everything | Use codex exec for scripted/CI runs so it does not wait for prompts |
| Tying up the laptop with a long task | Not knowing cloud exists | Run long-running or parallel work in Codex cloud |
Assuming gpt-5.4 still works | It retired under ChatGPT sign-in on 31 Aug 2026 | Migrate the shared config to gpt-5.6-terra / -luna, or use API-key sign-in |
| Setting the model per session every time | Not knowing config.toml sets a default | Set model = 'gpt-5.6' once in the shared config |
| Reaching for the IDE for a CI job | The IDE is the comfortable surface | CI is non-interactive; use codex exec on the CLI |
| Treating desktop and CLI as different agents | The surfaces look different | They share one config.toml; capability is the same, interaction differs |
Scenario challenge
Scenario. Priya leads a four-person platform team. They have three pieces of work in flight this afternoon: (1) a mechanical rename of a deprecated logging call across ~180 files, (2) a genuinely hard redesign of the request-batching logic that one engineer will pair on, and (3) a set of four independent, small bug fixes in unrelated modules that all need to land today. The team signs in to Codex with their ChatGPT accounts, and their shared config.toml still pins model = "gpt-5.4". One engineer reports Codex “stopped picking up the model” this morning.
Expert reasoning trace.
- Fix the blocker first. The config pins
gpt-5.4, which retired from Codex under ChatGPT sign-in on 31 August. Since the whole team signs in with ChatGPT, the fix is to change the shared default togpt-5.6-terra(the direct replacement forgpt-5.4). API-key sign-in would be unaffected, but that is not how this team authenticates. - Place task 1 (the rename). It is mechanical and high-volume, and it can run unattended. Surface: CLI with
codex exec. Model: Luna. Effort: Low. No need to occupy anyone. - Place task 2 (the redesign). One hard problem, an engineer pairing. Surface: IDE extension for a tight review loop. Model: Sol or Astra. Effort: Max — more thinking time on one task, not Ultra, because it does not split into independent parts.
- Place task 3 (four independent fixes). These genuinely decompose. This is the Ultra case — automatic delegation to subagents in parallel — or four cloud tasks running at once. Either way, cloud is a good surface so the fixes do not block the two engineers.
- Do not over-model. Nothing here needs Astra except possibly the redesign; the rename on Astra would burn budget for no benefit.
Exam-correct decision: migrate the shared config to gpt-5.6-terra; run the rename non-interactively on the CLI with Luna at Low effort; do the redesign in the IDE with Sol/Astra at Max; fan the four independent fixes out with Ultra or as parallel cloud tasks. Not everything on Astra, not Ultra for the single redesign, not the rename in an interactive session.
Assessment traps
| Trap | Why it is tempting | The discriminator |
|---|---|---|
| “Use Astra for the important task” | Strongest model feels safest | Match model to difficulty, not importance; Terra/Sol usually suffice |
| “Use Ultra to think harder” | Ultra is the top rung | Ultra means parallel subagents; use Max to think harder on one task |
“gpt-5.4 is fine, it is recent” | It was current not long ago | It retired from Codex under ChatGPT sign-in on 31 Aug 2026 |
| “Run the CI fix interactively” | codex is the familiar command | CI is unattended; use codex exec |
| “Cloud and CLI are different agents” | The interfaces differ | Same agent, one shared config; the surface only changes interaction |
“Set the model with --effort” | Sounds plausible | Model is -m / --model / /model; effort is the Low–Ultra ladder |
| “Ultra is missing, so my plan is impossible” | It is absent from the picker | Enable the Ultra slider in Settings → Configuration |
Practice questions
Each item states how many responses to select. Attempt before revealing.
Q1 · A developer wants to run a repeatable code fix inside a CI job with no human at the keyboard. Which surface and mode fit BEST? (Select one)
A. The Codex IDE extension in interactive mode
B. The Codex CLI with codex exec
C. The ChatGPT desktop app chat
D. Codex Micro
Answer: B. CI is unattended and scripted, which is exactly what codex exec on the CLI is for. The IDE extension (A) and desktop chat (C) are interactive and wait for input. Micro (D) is for small quick tasks, not a scripted CI pipeline.
Q2 · What is the difference between the Max and Ultra reasoning levels? (Select one)
A. Max is cheaper than Ultra but otherwise identical B. Max gives more thinking time on one task; Ultra automatically delegates to subagents running in parallel C. Ultra uses a smaller model; Max uses the largest model D. They are two names for the same setting
Answer: B. Max deepens reasoning on a single task; Ultra spreads work across parallel subagents. Cost (A) is not the distinction. Neither level changes the model choice (C), and they are not the same setting (D).
Q3 · A team's shared `config.toml` pins `model = 'gpt-5.4'` and they sign in to Codex with ChatGPT. On 1 September the model stops working. What happened and what is the fix? (Select one)
A. A billing lapse; renew the subscription
B. gpt-5.4 retired from Codex under ChatGPT sign-in on 31 Aug 2026; change the default to gpt-5.6-terra
C. The config file is corrupt; regenerate it
D. Codex requires the newest CLI; upgrade it
Answer: B. gpt-5.4 (and gpt-5.4-mini) retired from Codex under ChatGPT sign-in on 31 August 2026, with gpt-5.6-terra and gpt-5.6-luna as replacements. Billing (A), a corrupt file (C) and the CLI version (D) do not explain a model-specific cut-off tied to that date and sign-in type.
Q4 · Which model is the pragmatic default and the natural replacement for previous GPT-5.5 workloads in Codex? (Select one)
A. gpt-6-astra
B. gpt-5.6-luna
C. gpt-5.6-terra
D. gpt-5.3-codex-spark
Answer: C. Terra is described as the pragmatic all-rounder and the natural replacement for GPT-5.5 workloads. Astra (A) is for the hardest work, Luna (B) for high-volume mechanical tasks, and Spark (D) is a text-only research preview for near-instant iteration.
Q5 · You are launching a non-interactive Codex run and want it on GPT-5.6 Sol. Which command is correct? (Select one)
A. codex --effort sol "refactor the parser"
B. codex exec -m gpt-5.6-sol "refactor the parser"
C. codex /model gpt-5.6-sol
D. codex run --sol "refactor the parser"
Answer: B. codex exec runs non-interactively and -m selects the model. --effort (A) is not how you pick a model; /model (C) is an in-session command, not a shell flag; codex run --sol (D) is not a real command form.
Q6 · An engineer wants a tight loop where they see and approve each diff inside their editor while pairing on a hard refactor. Which surface fits BEST? (Select one)
A. Codex cloud
B. The Codex CLI with codex exec
C. The Codex IDE extension
D. Codex Micro
Answer: C. The IDE extension puts the diff in front of you for a tight edit-review loop. Cloud (A) is for long-running or parallel work away from your machine; codex exec (B) is non-interactive; Micro (D) is for small quick tasks, not a hard refactor with review.
Q7 · Which TWO tasks are the strongest fit for the Ultra reasoning level rather than Max? (Select two)
A. Four independent bug fixes in unrelated modules that all need to land today B. One deep redesign of a single algorithm that needs longer reasoning C. A batch migration split into several independent, parallelisable sub-tasks D. A one-line typo fix E. Explaining what a function does
Answer: A and C. Ultra delegates to parallel subagents, so it fits work that decomposes into independent parts (A, C). A single deep problem (B) is a Max case. A trivial fix (D) and a read-only explanation (E) need no high effort at all.
Q8 · Where does the shared model default live so that the CLI, IDE extension and desktop app all use it? (Select one)
A. In each surface’s own separate settings
B. In a single shared config.toml with model = 'gpt-5.6'
C. It cannot be shared; you set it per session
D. In AGENTS.md
Answer: B. Codex surfaces share one config.toml; model = 'gpt-5.6' sets the default everywhere. Per-surface settings (A) and per-session only (C) contradict the shared-config design. AGENTS.md (D) carries repository guidance, not the model default.
Q9 · A team needs to run several long tasks overnight without occupying anyone's laptop. Which surface fits BEST? (Select one)
A. Codex Micro B. The Codex IDE extension C. Codex cloud D. The ChatGPT desktop app
Answer: C. Codex cloud runs in a hosted sandbox, ideal for long-running or parallel work independent of your machine. Micro (A) is for small quick tasks; the IDE (B) and desktop app (D) tie up the local machine and expect interaction.
Q10 · Ultra is not visible in a user's model picker. What is the correct step? (Select one)
A. Reinstall the Codex CLI B. Enable the ‘Ultra in model picker slider’ via Settings → Configuration C. Upgrade to GPT-6 Astra D. Ultra only exists in the cloud surface
Answer: B. When Ultra is absent from the picker it is enabled via Settings → Configuration → ‘Ultra in model picker slider’. Reinstalling (A) and upgrading a model (C) are unrelated; Ultra is not restricted to cloud (D).
Q11 · A developer is doing a small, high-volume, mechanical edit repeated across many files and wants the cheapest model that will do it. Which model fits BEST? (Select one)
A. gpt-6-astra
B. gpt-5.6-sol
C. gpt-5.6-luna
D. gpt-5.3-codex-spark
Answer: C. Luna is for clear, repeatable, high-volume tasks and is the cheapest current model. Astra (A) and Sol (B) are for hard, high-value work and would overspend here. Spark (D) is a Pro-only research preview for fast iteration, not the standard high-volume choice.
Q12 · A lead argues that because Codex behaves differently in the CLI and the desktop app, they must configure each one separately. What is the accurate correction? (Select one)
A. They are right; each surface is a separate agent
B. The surfaces share one config.toml; the difference is how you interact, not the underlying agent or configuration
C. Only the CLI can be configured
D. Desktop cannot run Codex at all
Answer: B. Codex is one agent behind many surfaces, sharing a single config.toml; the surfaces differ only in interaction model. They are not separate agents (A), the CLI is not the only configurable surface (C), and the desktop app does run Codex (D).
Q13 · Which TWO statements about the `gpt-5.4` retirement are correct? (Select two)
A. It retired from Codex under ChatGPT sign-in on 31 August 2026
B. It also retired for API-key sign-in on the same date
C. Its replacement for the mini variant is gpt-5.6-luna
D. It was replaced by gpt-5.5
E. Nothing needs to change in a shared config
Answer: A and C. gpt-5.4 retired from Codex under ChatGPT sign-in on 31 Aug 2026, and gpt-5.4-mini’s replacement is gpt-5.6-luna (with gpt-5.6-terra replacing gpt-5.4). API-key sign-in is unaffected, so B is wrong; the replacement is the 5.6 line, not gpt-5.5 (D); and a config pinning gpt-5.4 under ChatGPT sign-in must change (E).
Q14 · An engineer is stuck on one genuinely hard algorithmic problem and wants Codex to reason for longer on that single task. Which reasoning level fits BEST? (Select one)
A. Low B. Ultra C. Max D. Medium
Answer: C. Max gives more thinking time on a single task, which is exactly this case. Ultra (B) delegates to parallel subagents, which does not help a single indivisible problem. Low (A) and Medium (D) apply less reasoning than the hard problem warrants.
Key takeaways
- Codex is one agent behind many surfaces (desktop, ChatGPT Work web, CLI, IDE extension, cloud, Micro) that share a single
config.toml. - Choose the surface by where the work happens: CLI for terminal/CI (with
codex execfor unattended runs), IDE for tight review loops, cloud for long-running or parallel work. - Pick the lowest model that meets the task: Luna (high-volume), Terra (default / GPT-5.5 replacement), Sol (complex), Astra (hardest); Spark is a Pro-only research preview.
- Control the model with
/modelin-session,-m/--modelat launch, ormodel = '…'inconfig.toml. - The reasoning ladder is Low → Medium → High → Extra High → Max → Ultra; Max deepens one task, Ultra delegates to parallel subagents.
gpt-5.4/gpt-5.4-miniretired from Codex under ChatGPT sign-in on 31 Aug 2026 → replace withgpt-5.6-terra/gpt-5.6-luna; API-key sign-in is unaffected.
Last updated Sep 18, 2026