Codex Path
Codex · Mock Exam 2
A harder 50-item, domain-weighted independent mock exam for the OpenAI Academy Codex pathway, with full explanations and a readiness readout. Not an official OpenAI assessment.
This is the second full-length, domain-weighted independent mock exam for the Codex track. It is built from publicly available OpenAI learning objectives and is not an official OpenAI assessment. Mock Exam 2 is deliberately harder than Mock Exam 1: more multi-constraint stems and more FIRST / BEST / MOST cost-effective / TWO qualifiers, with more scenario framing. All 50 questions are new and distinct from Mock Exam 1 and the domain-page items. Use it timed as your go/no-go gate.
Instructions
- Time: 60 minutes (our design choice for a 50-item mock; the Academy assessments are shorter, randomised selections from a larger bank).
- Items: 50, multiple-choice and multiple-response. Each item states how many answers to select.
- Selection: for multiple-response items you must select all correct options and no incorrect ones; partial selections are marked wrong.
- No guessing penalty: answer every question.
- Target: aim for at least 80% raw (≈ 40/50) before taking the real Academy Codex assessment, which passes at ≥ 80%.
- Work each question before expanding the answer.
Domain distribution
| # | Domain | Items here |
|---|---|---|
| 1 | Codex Fundamentals and Surfaces | 10 |
| 2 | Core Coding Workflows | 12 |
| 3 | Extending and Configuring Codex | 10 |
| 4 | Team Adoption and Governance | 10 |
| 5 | Scaling Across Teams and Systems | 8 |
Total: 10 + 12 + 10 + 10 + 8 = 50 items.
Readiness interpretation
This is an independent readiness indicator, not a score and not a prediction of any official result.
| Raw score (of 50) | Band | Interpretation |
|---|---|---|
| 45–50 | Strong readiness | 90%+; strong across all domains |
| 40–44 | Assessment ready | 80–89%; at or above the Academy threshold |
| 35–39 | Building confidence | 70–79%; close, target your weak domains |
| under 35 | Keep learning | below 70%; revisit the domain pages before re-attempting |
The 80% line is deliberate: it matches the Academy badge threshold.
Take the mock exam
Two ways to use the questions below: the interactive mode runs a timed sitting one question at a time and ends with your score, a per-domain breakdown and a full correction; the review mode underneath lists every question with its options one per line and the answer hidden until you ask for it.
Interactive mode
Take the practice exam
50 questions · one at a time · 60-minute countdown · results with per-domain breakdown and full correction at the end. Your progress is saved in this browser if you leave the page.
By domain
| Domain | Correct | Score |
|---|
Correction
All questions (review mode)
Options are listed one per line. The answer and explanation stay hidden until you click Show answer. Use the interactive mode above for a timed sitting.
A one-file typo fix must ship in the next five minutes, cost is a concern, and it decomposes into no sub-tasks. Which model and reasoning effort are MOST cost-effective?
Show answer
Answer: B.
A trivial mechanical fix wants the cheapest model (Luna) at low effort. Astra at Ultra (A) burns budget and adds pointless parallelism, Sol at Max (C) over-thinks a typo, and Terra at Extra High (D) still applies far more effort than a typo warrants.
An engineer must decide between Max and Ultra for a job that is a single, indivisible constraint-solving problem that keeps timing out at High. Which is the BEST choice and why?
Show answer
Answer: B.
Max deepens reasoning on a single indivisible task, which is this case. Ultra (A, C) delegates to parallel subagents and does not help an indivisible problem, and neither Max nor Ultra changes the model choice (D).
A team's shared config pins gpt-5.4, everyone signs in with ChatGPT, and after 31 August 2026 the model stops resolving. Which single action FIRST restores work with the direct replacement?
Show answer
Answer: B.
gpt-5.4retired under ChatGPT sign-in and Terra is its direct replacement, so updating the shared default resolves it. A subscription renewal (A), a corrupt file (C) and a CLI upgrade (D) do not explain a model-specific cut-off tied to that date and sign-in type.A workload is high-volume, mechanical, latency-sensitive and cost-constrained, running unattended in CI on GitHub. Which TWO choices best fit it?
Show answer
Answer: A and B.
Unattended CI is the
codex execcase (A) and a high-volume mechanical job wants the cheapest model, Luna (B). The interactive IDE (C) waits for input, Astra (D) overspends, and Ultra (E) adds parallelism that mechanical volume does not need.A user cannot find Ultra in their model picker but their task genuinely decomposes into independent parts. What is the correct FIRST step before redesigning the task?
Show answer
Answer: B.
When Ultra is absent from the picker it is enabled via Settings then Configuration then the 'Ultra in model picker slider'. Reinstalling (A) and upgrading a model (C) are unrelated, and Ultra is not restricted to cloud (D).
Which statement most accurately distinguishes GPT-5.6 Sol from GPT-5.6 Terra for Codex work?
Show answer
Answer: B.
Sol targets complex, open-ended, high-value work while Terra is the pragmatic default and the natural GPT-5.5 replacement. The cheapest high-volume model is Luna not Sol (A), the text-only preview is Spark (C), and they are distinct models (D).
A lead insists that because the desktop app and CLI 'look and feel different', each must be a different underlying agent needing separate configuration. What is the MOST accurate correction?
Show answer
Answer: B.
Codex is one agent behind many surfaces sharing one
config.toml; the surfaces differ only in interaction. They are not separate agents (A), the desktop app is a real Codex surface not a viewer (C), and it does run Codex (D).Three tasks are queued: (1) an overnight mechanical rename across many files, (2) one hard indivisible algorithm redesign, (3) four unrelated independent bug fixes due today. Which allocation of reasoning effort is BEST?
Show answer
Answer: B.
The mechanical rename needs little effort, the single hard redesign wants Max's deeper single-task thinking, and the four independent fixes decompose so they fit Ultra's parallel subagents. Ultra everywhere (A) and Max everywhere (D) misapply effort, and C inverts the correct assignments.
Which surface should a developer reach for when they want the diff in front of them in their editor and to approve each change during a hard, interactive refactor?
Show answer
Answer: B.
The IDE extension puts the diff in front of you for a tight, interactive edit-review loop. Cloud (A) suits long-running or parallel work away from the machine,
codex exec(C) is non-interactive, and Micro (D) is for small quick tasks.Which is the correct way to select the model and effort separately in Codex?
Show answer
Answer: B.
Model is chosen with
-m,--modelor/model, while effort is the Low-to-Ultra ladder — two independent controls. Option A swaps the two,AGENTS.mdis repo guidance not the effort control (C), and there is no single--effortflag for both (D).Codex returns a 300-line diff touching four unrelated areas from the one-line prompt 'speed up the export', and claims a 40% improvement with green tests, with a demo in one hour. What should you do FIRST?
Show answer
Answer: B.
A sprawling diff from a vague prompt is a scoping failure; re-scoping and independent verification is the correct first move, and the deadline raises rather than lowers the cost of a bad merge. Merging on the claim (A) skips verification, a bigger model (C) still sprawls on a vague task, and cherry-picking (D) risks an inconsistent change.
A reviewer wants to trust a Codex PR in minutes without reconstructing the author's confidence. Which TWO pieces of evidence are MOST valuable?
Show answer
Answer: A and C.
A focused diff (A) and a passing test run with a regression test (C) let a reviewer judge scope and correctness quickly. The effort level (B), a claim about the model (D) and token counts (E) tell a reviewer nothing about whether the change is right.
A nightly maintenance job must apply a mechanical migration across a repo with no human present, on the lowest cost that works, and it must not be able to escalate silently. Which setup is BEST?
Show answer
Answer: B.
Unattended mechanical work is the
codex execcase with a cheap model (Luna) and restrictive permissions plus a sandbox so it cannot escalate silently. Interactive mode (A) waits for input, an open IDE (C) is not for unattended runs and broad auto-approve is unsafe, and Astra with broad auto-approve (D) both overspends and removes containment.A teammate argues that because a Codex change has green tests, reading the diff is unnecessary before merge. What is the strongest counter-argument?
Show answer
Answer: B.
Tests are necessary but not sufficient; the diff reveals scope creep and confirms the change is on-task. Green tests are not a full review (A), tests remain valuable (C), and the model ID is irrelevant to correctness (D).
Codex reports a bug is fixed and the suite is green, but the original reported repro was never added to the suite. What is the MOST likely hidden risk?
Show answer
Answer: B.
A green suite that never exercised the real repro can hide a lookalike fix, so reproduce the reported case and add a regression test. A green suite alone does not prove the specific fix (A), deleting tests (C) reduces coverage, and swapping models (D) does not verify anything.
Three engineers must build three independent features locally at once without their branches colliding, keeping each verifiable. Which approach is BEST?
Show answer
Answer: B.
Separate worktrees or branches (or parallel cloud tasks) let independent work proceed at once, verified before a controlled merge. Serialising (A) wastes the parallelism, committing to main (C) is unsafe, and disabling tests (D) removes verification.
Which practice most directly satisfies the Codex objective to 'verify changes and provide clear evidence for review'?
Show answer
Answer: B.
Independently verifying and attaching evidence is exactly what the objective asks for. Model choice (A) and prompt length (C) do not produce evidence, and merging quickly (D) skips verification.
A prompt says 'improve the auth code'. Which rewrite BEST turns it into a scoped, reviewable task?
Show answer
Answer: B.
Option B supplies a goal, boundaries and an executable definition of done. 'Better and faster overall' (A) and 'refactor everything however you see fit' (C) are still unscoped, and specifying a model (D) is not scoping the task.
A developer wants runtime signals — test output and stack traces — fed back into the loop to verify and debug a change. Which context source provides that BEST?
Show answer
Answer: B.
The integrated terminal surfaces runtime signals such as test output and stack traces for verifying and debugging in-loop. A PR description (A) and a source comment (C) are static text, and the analytics dashboard (D) reports adoption, not runtime signals.
For a bug reported as 'checkout occasionally double-charges', which is the MOST reliable definition of done to hand Codex?
Show answer
Answer: B.
A failing test that reproduces the reported behaviour is an executable definition of done, and the boundary protects the gateway contract. 'Somehow' (A) is unscoped, a full rewrite (C) is disproportionate scope creep, and adding logging (D) does not define done or fix the bug.
A developer habitually runs CI fix jobs by opening an interactive codex session on their laptop and typing the task. Why is this the wrong tool, and what is the fix?
Show answer
Answer: B.
CI runs without a human, so
codex execis the correct non-interactive form; an interactive session would block waiting for input. Interactive is not ideal for CI (A), the issue is the mode not hardware (C), and interactive sessions can run tests (D).Which TWO elements are core parts of a well-scoped Codex task?
Show answer
Answer: A and C.
Scope is goal plus boundaries plus definition of done plus constraints; boundaries (A) and a definition of done (C) are two of those. A pricier model (B), a longer prompt (D) and defaulting to Ultra (E) are not scoping and do not substitute for it.
A fintech team wants a nightly dependency-upgrade job that opens a PR, and proposes broad auto-approve, full cloud internet access, experimental context management, and auto-merge on auto-review — on a ChatGPT Enterprise workspace. Which single element is impossible on availability grounds alone?
Show answer
Answer: C.
Experimental context management is Plus/Pro sign-in only at launch, so it is unavailable on an Enterprise workspace regardless of its merits. Broad auto-approve (A), full internet access (B) and auto-merge on auto-review (D) are unsafe choices but are not blocked purely by availability.
For a risky unattended run, which TWO controls should be combined to limit both what happens without a human and how far any action can reach?
Show answer
Answer: A and C.
A restrictive permission mode (A) governs what needs approval while a sandbox (C) contains impact; together they limit the decision surface and the blast radius. Broad auto-approve (B), full internet access (D) and disabling approvals (E) increase risk rather than contain it.
A team needs Codex to reach an internal data warehouse, run a custom check before every write, AND reproduce a specific session for debugging. Which mapping of needs to extension points is correct?
Show answer
Answer: B.
External data is MCP, running your own logic at a lifecycle point is a hook, and reproducing a session is record and replay. Option A swaps MCP and hooks, a single plugin (C) does not map to these distinct needs, and D mismatches every mapping.
An engineer proposes relying on auto-review as the sole gate before a fintech dependency PR auto-merges. What is the MOST accurate objection?
Show answer
Answer: B.
Auto-review is evidence feeding a human gate, not a replacement for it, especially for a regulated change. It does not fully substitute for a reviewer (A), it is not limited to formatting (C), and it is available across Codex surfaces including cloud (D).
What is the correct way to enable experimental context management, and under what limit?
Show answer
Answer: B.
It is opt-in through
features.context_management.experimental_mode = trueand, at launch, limited to Plus/Pro sign-in. It is not automatic (A), not a per-command flag (C), and not enabled inAGENTS.mdnor Enterprise-scoped (D).A team keeps putting per-repository build commands into config.toml and wonders why behaviour differs across repos. Which correction is MOST accurate?
Show answer
Answer: B.
Repo-specific build commands and conventions go in
AGENTS.md, whileconfig.tomlsets machine or workspace defaults, which explains the divergence.config.tomlis not per-repo (A), build commands are configurable (C), and the files serve different purposes so need not match (D).A cloud dependency-fetch task needs network access to one package registry. What is the MOST appropriate internet-access configuration?
Show answer
Answer: B.
Least privilege means granting only the access the task needs and keeping the default restricted. Fully open (A) is over-permissive, blocking access the task requires (C) breaks the task, and switching sign-in to bypass controls (D) is a governance failure.
An engineer wants to reproduce and share an exact Codex session with a colleague for debugging. Which feature is designed for that?
Show answer
Answer: B.
Record and replay captures a session so it can be re-run or shared. Skills (A) are reusable abilities, MCP (C) provides external access, and auto-review (D) summarises a diff.
A team confuses hooks and rules. Which distinction is MOST accurate?
Show answer
Answer: B.
Hooks execute your own logic at lifecycle points, while rules constrain what the agent may do. They are not the same (A), external access is MCP not rules (C), and neither is surface-restricted as described (D).
Which single statement correctly separates the scope of config.toml from AGENTS.md?
Show answer
Answer: B.
config.tomlsets machine or workspace defaults whileAGENTS.mdholds per-repository guidance — the reverse of option A. They are not interchangeable (C), and neither is restricted to a single surface (D).A healthcare company rolling Codex out to 120 engineers has repos touching PHI and a nightly cloud job currently authenticated with one engineer's personal access token. Which single change should be the FIRST governance fix for that job?
Show answer
Answer: B.
A personal token ties an org automation to one person and breaks when they leave; workload identity or a service account is the correct org identity. More admin scope (A) worsens least-privilege, sharing the token (C) is a security failure, and interactive mode (D) defeats an unattended job.
A regulated PHI workload needs both correct configuration and a provable action history. Which TWO belong in the exam-correct answer?
Show answer
Answer: A and C.
A regulated workload needs the right configuration (HIPAA) and an audit trail (Compliance API and audit events). Broad auto-approve (B) increases risk, disabling analytics (D) removes useful oversight, and a shared admin account (E) destroys accountability.
An organisation wants to prevent Codex configuration drift across teams while enforcing which models may be used. Which pair of controls FIRST addresses both?
Show answer
Answer: B.
Managed configuration enforces consistent settings and workspace model availability restricts which models the workspace may use. PATs and record and replay (A), the reasoning ladder and
codex exec(C), and one repo'sAGENTS.mdwith Slack (D) do not enforce org-wide config or model policy.Which statement about Codex Security is MOST accurate for governing code produced across a team?
Show answer
Answer: B.
Codex Security spans the plugin, CLI and cloud, with scans, triage, fixes and hardening, and integrates into CI or GitLab CI so code is scanned before merge. It is not CLI-only or post-merge-manual (A), it is part of the Codex security surface (C), and it does more than reporting (D).
Leadership asks for repeatable monthly adoption numbers for a board update. Which approach is MOST appropriate?
Show answer
Answer: B.
The Analytics API gives repeatable programmatic figures and the dashboard complements it. Impressions (A), license-count estimates (C) and shell-history sampling (D) are unreliable and not repeatable adoption measures.
An admin argues that giving everyone admin avoids permission tickets and that PHI repos need no special handling because access is internal. Which correction addresses BOTH errors?
Show answer
Answer: B.
Least privilege fixes the universal-admin error and HIPAA configuration is required for PHI regardless of internal framing. Both claims are not fine (A), disabling analytics (C) removes oversight without fixing either error, and relying on the model to protect PHI (D) is not a compliance control.
Which authentication choice is correct for each actor: a GitHub Actions pipeline, a shared org automation bot, and one engineer's personal script?
Show answer
Answer: B.
A CI/cloud pipeline uses workload identity federation (no long-lived secret), a shared org automation uses a service account, and an individual's script uses a PAT. Option A and C mismatch the actors, and using the admin's personal account for everything (D) destroys accountability.
During a rollout, which TWO controls together ensure access is provisioned consistently and revoked promptly when engineers leave?
Show answer
Answer: A and C.
Automated groups and provisioning manage membership so leavers lose access (A), and least-privilege roles and permissions govern what each member may do (C). The model picker (B) selects models, Codex Security (D) scans code, and the Analytics API (E) reports adoption.
An admin wants Codex-produced code scanned before merge on every repo, with an auditable record of who did what. Which pairing is correct?
Show answer
Answer: B.
Codex Security in CI scans before merge and the Compliance API with audit events provides the who-did-what record. Manual scans and the Analytics API (A) neither scan reliably nor audit, record and replay with model availability (C) do neither job, and auto-review alone (D) is evidence, not a scanning or audit control.
Which is the correct characterisation of managed configuration versus a shared setup document?
Show answer
Answer: B.
Managed configuration enforces settings centrally, whereas a document merely describes them and depends on manual compliance. They are not equivalent (A), the roles are not reversed (C), and managed configuration can control extensions (D).
A 40-engineer, 25-repo group reports tasks-per-week tripled while the PR revert rate quietly climbed. Which signal should drive the go/no-go decision, and what is the FIRST structural fix?
Show answer
Answer: B.
The rising revert rate is the quality signal that matters, and restoring the review gate while measuring quality is the structural fix. Tasks-per-week is activity not value (A, C), and a climbing revert trend should not be ignored (D).
Which TWO failure modes emerge specifically when scaling Codex across many repositories rather than in single-task use?
Show answer
Answer: A and C.
Config drift across repos (A) and review lapsing under change volume (C) are classic at-scale failure modes. A clear single prompt (B), a correct single-task model choice (D) and adding one
AGENTS.md(E) are healthy single-task practices, not scaling failures.A platform team must map three needs to the right Codex integration: Codex steps in GitHub CI, embedding Codex in an internal developer portal, and acting on Linear issues. Which mapping is correct?
Show answer
Answer: B.
The GitHub Action embeds Codex in GitHub CI, the SDK builds Codex into your own tooling like a portal, and the Linear integration handles issue-driven work. Option A swaps the SDK and Action, C mismatches all three, and Slack alone (D) does none of these jobs.
A team on GitLab plans a business-critical launch next week relying on the Codex GitLab integration. What is the MOST responsible planning stance?
Show answer
Answer: B.
The GitLab integration is beta, so a critical launch plan should account for beta risk. It is not GA-equivalent to GitHub (A), GitLab is supported in beta rather than unsupported (C), and a full migration (D) is disproportionate.
When scaling Codex, which combination BEST captures whether the rollout is actually helping?
Show answer
Answer: B.
Pairing adoption with quality shows whether rising usage is producing good outcomes. Adoption alone can hide falling quality (A), quality alone ignores whether the tool is used (C), and effort levels (D) do not measure rollout value.
A manager proposes to 'keep scaling and fix problems reactively as they appear'. Why is this the wrong stance for known scaling failure modes?
Show answer
Answer: B.
Known scaling failure modes need structural mitigations because reactive fixes trail the damage they cause. Reactive-only is not efficient here (A), a larger model does not prevent process failures (C), and these problems do occur at scale (D).
To make parallel outputs from many repositories safe to integrate, which upstream practice matters MOST?
Show answer
Answer: B.
Standardising
AGENTS.mdand configuration means outputs are produced under the same rules, which is what makes parallel integration safe. Divergent conventions (A) make integration unsafe, a bigger model (C) does not fix inconsistent context, and a rushed monorepo merge (D) is disproportionate and risky.Security debt accumulates when vulnerabilities merge faster than they are found across 30 repositories. Which TWO practices form the correct structural mitigation?
Show answer
Answer: A and C.
Scanning every repo in CI before merge (A) and keeping the human review gate under volume (C) are the structural mitigations. Scanning one repo occasionally (B) leaves gaps, post-breach scanning (D) is too late, and skipping scanning (E) is the risk to avoid.
Last updated Sep 18, 2026