# D2 · Core Coding Workflows

Scoping a coding task, giving repository context, review-first loops, non-interactive runs, verifying changes with tests and diffs, evidencing review, git worktrees and long-running work.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This is the heaviest domain on the mock — **24%**, roughly **12 of 50 items**. It is the practical core of the Codex pathway: scope a task well, give Codex the context it needs, work review-first, verify the change with tests and diffs, and hand a reviewer clear evidence. Nearly every item hinges on one question: *how do I get a correct change I can defend to a reviewer, without doing the work by hand?*

## What you need to know

A good Codex workflow starts with a tightly scoped task and enough repository context — the relevant files, the build and test commands, the conventions — usually captured in `AGENTS.md`. You work **review-first**: let Codex plan, approve the plan, let it implement, then verify. Verification means running the tests, reading the diff, and collecting evidence a reviewer can trust, not "it seems to work". Non-interactive runs (`codex exec`) suit scripted and unattended work; the integrated terminal and code review keep the loop tight. For long-running or parallel work, a **git worktree** lets a task run on its own branch without disturbing your working tree.

## Learning objectives

By the end of this page you should be able to:

1. **Scope a coding task** so Codex has a clear goal, boundaries and a definition of done.
2. **Give repository context** through `AGENTS.md`, relevant files, and build/test commands.
3. **Run a review-first workflow**: plan, approve, implement, verify.
4. **Use `codex exec`** for non-interactive and scripted runs.
5. **Verify a change** with tests and diffs and produce evidence a reviewer can rely on.
6. **Use git worktrees** to isolate long-running or parallel Codex work.

---

## 2.1 Scoping a coding task

The single biggest lever on output quality is the scope you hand Codex. A vague task produces a vague or over-broad change; a scoped task produces a reviewable one.

| Element of a good scope | Weak version | Strong version |
| --- | --- | --- |
| **Goal** | "improve the auth code" | "fix the token-refresh race in `auth/refresh.ts`" |
| **Boundaries** | (none) | "do not touch the public API or the database schema" |
| **Definition of done** | (implied) | "the new test `refresh.race.test.ts` passes and existing tests still pass" |
| **Constraints** | (none) | "keep the change under ~50 lines; no new dependencies" |

```text
SCOPE = Goal + Boundaries + Definition-of-done + Constraints
        │        │              │                   │
     what to   what NOT      how you'll know      cost / risk
     change    to touch      it's correct         limits
```

:::tip[Assessment signal]
Stems where the task is a one-liner ("make it faster", "clean this up") and Codex produces a sprawling change are testing scoping. The correct answer narrows the goal, sets boundaries and states a definition of done — it does not just re-run at a higher reasoning effort.
:::

## 2.2 Giving repository context

Codex works better with the same context a new engineer would need. The durable place to put it is `AGENTS.md`.

```md
# AGENTS.md
## Build & test
- Install: `pnpm install`
- Build: `pnpm build`
- Test: `pnpm test` (Vitest); a single file: `pnpm test path/to/file`
- Lint: `pnpm lint`

## Conventions
- TypeScript strict; no `any`.
- Prefer pure functions in `src/core`; side effects live in `src/adapters`.

## Do not touch
- `src/generated/**` (code-generated).
- Public API types in `src/api/types.ts` without an ADR.
```

| Context source | What it gives Codex | When to use |
| --- | --- | --- |
| `AGENTS.md` | Durable build/test commands, conventions, no-go areas | Always; check it into the repo |
| Named files in the prompt | The exact code to change | For a specific, scoped task |
| The integrated terminal | Runtime signals — test output, stack traces | To verify and debug in-loop |
| A failing test | An executable definition of done | Whenever the goal is "make this pass" |

## 2.3 The review-first workflow

Review-first means you keep a human decision point before anything irreversible. The loop:

<Steps>

1. **Scope** the task (2.1) and make sure the context is present (2.2).

2. **Ask for a plan first.** Have Codex outline the change before writing it. Cheap to correct, and it surfaces misunderstandings.

3. **Approve or correct the plan.** This is your first review gate.

4. **Let Codex implement**, inside a sandbox and permission mode appropriate to the risk.

5. **Verify** — run the tests, read the diff (2.5), collect evidence (2.6).

6. **Merge only after review.** Your own review plus, for anything shared, a second reviewer.

</Steps>

The point is not distrust; it is that *you own the change*. Review-first keeps the human where the risk is highest — before the merge, before a destructive command, before touching a protected path.

## 2.4 Non-interactive runs with `codex exec`

`codex exec` runs a task without an interactive session — the backbone of scripted and CI workflows.

<Tabs>
  <TabItem label="One-off scripted run">
    ```bash
    codex exec -m gpt-5.6-terra "add input validation to createUser and update its unit tests"
    ```
  </TabItem>
  <TabItem label="In a CI pipeline">
    ```bash
    # e.g. a nightly maintenance job
    codex exec -m gpt-5.6-luna "bump the deprecated logger calls to the new API across src/"
    ```
    Pair it with a restrictive permission mode and sandbox so an unattended run cannot escalate silently.
  </TabItem>
  <TabItem label="With explicit boundaries">
    ```bash
    codex exec "refactor the pricing module; do not change public types; run pnpm test"
    ```
  </TabItem>
</Tabs>

| Interactive (`codex`) | Non-interactive (`codex exec`) |
| --- | --- |
| Steer mid-task, approve escalations | One shot, no prompts |
| Exploratory and review-first loops | Scripted, repeatable, unattended |
| You are at the keyboard | CI, cron, batch jobs |

## 2.5 Verifying with tests and diffs

A change is not done because it compiles. Verify against the definition of done.

```text
VERIFY = run tests  +  read diff  +  reproduce the fix
         │             │             │
      does it work?  is the change  did it fix the ACTUAL
                     minimal and    reported problem, not a
                     on-scope?      lookalike?
```

- **Run the tests** Codex claims pass — do not take the claim on faith. If the goal was "make this test pass", run that test.
- **Read the diff.** Look for scope creep (files you did not expect), risky edits (config, migrations, generated code), and whether the change is minimal.
- **Reproduce the fix.** For a bug, confirm the original repro now passes and add a regression test if none exists.

:::tip[Assessment signal]
When a stem says Codex "reported all tests pass" or "says it fixed it", the item is testing whether you will independently run the tests and read the diff. The model's claim is a lead, not evidence.
:::

## 2.6 Evidence for reviewers

The Codex objective is explicit: *verify changes and provide clear evidence for review.* A reviewer should not have to reconstruct your confidence.

| Evidence | What it shows the reviewer | How to produce it |
| --- | --- | --- |
| Passing test run | The change meets the definition of done | Paste the `pnpm test` output, or the CI run link |
| A focused diff | The change is minimal and on-scope | The PR diff; call out anything unexpected |
| A new/updated test | The behaviour is now pinned | Add a regression test alongside the fix |
| An auto-review summary | Codex's own review of the diff | Turn on auto-review (see D3) and attach the summary |
| The repro before/after | The reported bug is actually fixed | Show the failing case now passing |

A pull request that says "Codex did it, LGTM" is the anti-pattern. A pull request with a scoped diff, a green test run, a regression test and a one-paragraph rationale is reviewable in minutes.

## 2.7 Git worktrees and long-running work

Long or parallel Codex work should not block your main working tree. A **git worktree** checks out a branch into a separate directory, so a task can run in isolation.

```bash
# Create an isolated worktree for a long refactor
git worktree add ../repo-refactor refactor/batching
# Run Codex there without disturbing your main checkout
codex exec -m gpt-5.6-sol "redesign request batching" 
# When done and merged:
git worktree remove ../repo-refactor
```

| Situation | Approach |
| --- | --- |
| One long-running task while you keep working | Run it in a worktree on its own branch, or in Codex cloud |
| Several independent tasks at once | Separate worktrees / branches, or Ultra / parallel cloud tasks |
| A quick fix on your current branch | Just run it in place; a worktree is overkill |

:::tip[Assessment signal]
`Without disturbing my working tree`, `on a separate branch`, `run several at once locally` → **git worktrees**. `Independent of my machine`, `overnight` → **Codex cloud**. The two solve overlapping problems; worktrees stay local, cloud offloads.
:::

## Decision framework

Use **SCOPE → CONTEXT → PLAN → VERIFY → EVIDENCE (SCPVE)** for every non-trivial task.

| Stage | Question | The move |
| --- | --- | --- |
| **Scope** | What exactly changes, and what must not? | Goal + boundaries + definition of done + constraints |
| **Context** | Does Codex know how to build, test and behave? | Ensure `AGENTS.md`, name the files, provide the failing test |
| **Plan** | Do we agree on the approach before code? | Ask for a plan; approve or correct it — first review gate |
| **Verify** | Is it actually correct and minimal? | Run the tests yourself; read the diff; reproduce the fix |
| **Evidence** | Can a reviewer trust this in minutes? | Attach test output, a focused diff, a regression test, an auto-review summary |

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Handing Codex a one-line vague task | It is faster to type | Add boundaries and a definition of done; scope drives quality |
| Trusting "all tests pass" without running them | The claim is convenient | Run the tests yourself; the claim is a lead, not evidence |
| Merging without reading the diff | The description sounds fine | Read the diff for scope creep and risky edits before merging |
| Running unattended work interactively | Habit | Use `codex exec` for scripted and CI runs |
| No `AGENTS.md`, so Codex guesses the build | It was never written | Check in an `AGENTS.md` with build/test commands and conventions |
| Long task blocks your working tree | Running it on your current branch | Use a git worktree or Codex cloud |
| A PR that says only "Codex did it" | The change felt obvious | Attach evidence: test run, focused diff, regression test |
| Skipping the plan step | Eager to see code | Ask for a plan first; it is the cheapest place to catch mistakes |
| Fixing a lookalike, not the reported bug | Not reproducing the original | Reproduce the repro before and after; add a regression test |

## Scenario challenge

**Scenario.** Sam is asked to "speed up the report export — it's too slow". They open Codex in the IDE and type exactly that. Codex produces a 300-line diff touching the export module, a caching layer, a database query and a config file, and reports "all tests pass, ~40% faster". The PR is due before a demo in an hour. A teammate says "the tests are green, just merge it".

**Expert reasoning trace.**

1. **The scope was too loose.** "Speed up the export" has no boundaries, no definition of done, no constraint on blast radius — hence the sprawling diff touching four areas. The first correction is to re-scope: which export, how slow is "too slow", what may and may not change (e.g., not the DB schema, not shared config).
2. **Do not trust "all tests pass".** Run the tests locally. If the goal is performance, the existing suite may not even measure it; a green suite says nothing about the 40% claim.
3. **Read the diff.** A 300-line change across caching, a query and config is far more than a targeted export speed-up needs. Scope creep and a config edit are exactly the risky signals to catch before a demo.
4. **Verify the actual claim.** "~40% faster" needs a measurement, not a model assertion — reproduce the slow case and time it before and after.
5. **Reject "green, just merge".** Green tests are one piece of evidence, not a review. A demo deadline raises, not lowers, the cost of a bad merge.
6. **Produce reviewable evidence.** Re-run scoped, attach the focused diff, the timing before/after, and a test that pins the improvement.

**Exam-correct decision:** re-scope the task with boundaries and a definition of done, ask for a plan, let Codex implement the narrow change, then run the tests and a timing measurement yourself, read the diff for scope creep, and open a PR with that evidence. **Not** merge on the green claim, **not** accept the 40% number unmeasured, **not** ship a four-area diff for a targeted speed-up.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Codex says the tests pass, so merge" | The claim is convenient | Run the tests yourself; the claim is a lead, not evidence |
| "The diff description is clear, no need to read the diff" | Descriptions are readable | The diff shows scope creep and risky edits the description hides |
| "Just raise the reasoning effort" for a bad result | Effort feels like the knob | A vague result usually needs better *scope*, not more effort |
| "Run the CI job in an interactive session" | `codex` is familiar | CI is unattended; use `codex exec` |
| "A long task must block my branch" | Not knowing worktrees | Use a git worktree or cloud to isolate it |
| "Green tests = reviewed" | Tests are objective | Tests are one evidence type; review also reads the diff and rationale |
| "No `AGENTS.md` is fine, Codex will figure out the build" | It often does | Guessing the build wastes runs; a checked-in `AGENTS.md` is durable context |

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · Codex reports 'all tests pass' after a change. What is the MOST appropriate next step before merging? (Select one)">
    A. Merge; the model would not claim tests pass if they did not
    B. Run the tests yourself and read the diff before merging
    C. Ask Codex in the same session whether it is sure
    D. Raise the reasoning effort and re-run

    **Answer: B.** The model's claim is a lead, not evidence; independently running the tests and reading the diff is the verification step. Trusting the claim (A) skips verification, a same-session self-check (C) reuses the same reasoning, and raising effort (D) does not verify anything.
  </AccordionItem>

  <AccordionItem title="Q2 · A one-line task 'clean up the utils module' produces a sprawling 400-line diff. What is the BEST corrective action? (Select one)">
    A. Accept it; more cleanup is better
    B. Re-scope the task with a specific goal, boundaries and a definition of done
    C. Switch to GPT-6 Astra and re-run the same one-liner
    D. Merge only the parts that look safe

    **Answer: B.** A sprawling diff is usually a scoping failure; a specific goal, boundaries and definition of done produce a reviewable change. More cleanup (A) is not the aim, a bigger model (C) will still sprawl on a vague task, and cherry-picking a diff (D) risks an inconsistent change.
  </AccordionItem>

  <AccordionItem title="Q3 · Where should durable build and test commands and 'do not touch' areas for a repository live so Codex uses them every time? (Select one)">
    A. In a comment in the main source file
    B. In `AGENTS.md` checked into the repository
    C. Only in each engineer's shell history
    D. In the PR description

    **Answer: B.** `AGENTS.md` is the durable, checked-in place for build/test commands, conventions and no-go areas. A source comment (A) is easy to miss, shell history (C) is per-machine and transient, and a PR description (D) is per-change, not durable repository context.
  </AccordionItem>

  <AccordionItem title="Q4 · Which command runs a scoped change non-interactively on GPT-5.6 Terra? (Select one)">
    A. `codex --interactive -m gpt-5.6-terra "..."`
    B. `codex exec -m gpt-5.6-terra "..."`
    C. `codex chat gpt-5.6-terra "..."`
    D. `codex plan -m gpt-5.6-terra "..."`

    **Answer: B.** `codex exec` runs non-interactively and `-m` selects the model. `--interactive` (A) is the opposite mode, and `codex chat` (C) and `codex plan` (D) are not the non-interactive execution form.
  </AccordionItem>

  <AccordionItem title="Q5 · A developer needs to run a long refactor on its own branch without disturbing their current working tree, staying on their machine. What fits BEST? (Select one)">
    A. A git worktree on a separate branch
    B. Delete the working tree and start over
    C. Force-push over main
    D. Disable tests to run faster

    **Answer: A.** A git worktree checks a branch out into a separate directory so a task runs in isolation without touching the main checkout. Deleting the tree (B) is destructive and unnecessary, force-pushing main (C) is dangerous and unrelated, and disabling tests (D) removes verification.
  </AccordionItem>

  <AccordionItem title="Q6 · What are the elements of a well-scoped Codex task? (Select two)">
    A. A clear goal and explicit boundaries on what not to touch
    B. The most expensive model available
    C. A definition of done, such as a test that must pass
    D. The longest possible prompt
    E. Ultra reasoning effort by default

    **Answer: A and C.** Scope is goal + boundaries + definition of done + constraints; A and C are two of those elements. A pricier model (B), a longer prompt (D) and defaulting to Ultra (E) are not scoping and do not substitute for it.
  </AccordionItem>

  <AccordionItem title="Q7 · A reviewer receives a PR that says only 'Codex implemented this, tests pass'. What evidence would make it reviewable in minutes? (Select two)">
    A. A focused diff with any unexpected changes called out
    B. A note that the reasoning effort was Max
    C. A pasted passing test run or CI link plus a regression test
    D. A statement that the model is very capable
    E. The model ID used

    **Answer: A and C.** Reviewable evidence is a focused diff plus a passing test run and a regression test that pins the behaviour. The effort level (B), a claim about the model (D) and the model ID (E) do not help a reviewer judge correctness or scope.
  </AccordionItem>

  <AccordionItem title="Q8 · The goal is 'make the failing test `checkout.test.ts` pass'. What is the cleanest way to give Codex a definition of done? (Select one)">
    A. Describe the bug in prose only
    B. Point Codex at the failing test as the executable definition of done and require the whole suite to still pass
    C. Ask for the largest change that could plausibly fix it
    D. Tell Codex to skip the test

    **Answer: B.** A failing test is an executable definition of done; requiring it to pass while the suite stays green scopes the task precisely. Prose only (A) is vaguer, a large speculative change (C) invites scope creep, and skipping the test (D) defeats the purpose.
  </AccordionItem>

  <AccordionItem title="Q9 · In a review-first workflow, what is the purpose of asking Codex for a plan before it writes code? (Select one)">
    A. To increase the token count
    B. To create a cheap, early review gate that surfaces misunderstandings before any code is written
    C. Because plans are required by the CLI
    D. To pick the model automatically

    **Answer: B.** The plan step is the cheapest place to catch a misunderstanding — you correct the approach before implementation. It is not about token count (A), it is not a CLI requirement (C), and it does not choose the model (D).
  </AccordionItem>

  <AccordionItem title="Q10 · Codex fixes a bug and the suite is green, but the original reported repro was never checked. What is the risk and the fix? (Select one)">
    A. No risk; a green suite proves the fix
    B. Codex may have fixed a lookalike issue, not the reported one; reproduce the original case before and after and add a regression test
    C. The suite is too slow; delete some tests
    D. The model is wrong; switch models

    **Answer: B.** A green suite that never exercised the actual repro can hide a lookalike fix; reproducing the reported case and adding a regression test closes the gap. A green suite alone does not prove the specific fix (A), deleting tests (C) reduces coverage, and swapping models (D) does not verify the fix.
  </AccordionItem>

  <AccordionItem title="Q11 · A nightly maintenance job should apply a mechanical migration across a repo with no human present. Which setup fits BEST? (Select one)">
    A. `codex` interactive with High effort
    B. `codex exec` with a low-cost model, a restrictive permission mode and a sandbox
    C. The IDE extension left open overnight
    D. A force-push to main after the change

    **Answer: B.** Unattended mechanical work is the `codex exec` case, paired with a cheap model and restrictive permissions/sandbox so it cannot escalate silently. Interactive mode (A) waits for input, an open IDE (C) is not designed for unattended runs, and force-pushing main (D) is destructive.
  </AccordionItem>

  <AccordionItem title="Q12 · When reading a Codex-produced diff, which signals most warrant closer scrutiny? (Select two)">
    A. Edits to config files or database migrations you did not expect
    B. Consistent code formatting
    C. Changes to generated files under a 'do not touch' path
    D. A helpful commit message
    E. Use of existing utility functions

    **Answer: A and C.** Unexpected config/migration edits and changes to protected generated paths are the risky, out-of-scope signals to scrutinise. Consistent formatting (B), a good commit message (D) and reusing utilities (E) are neutral-to-positive, not red flags.
  </AccordionItem>

  <AccordionItem title="Q13 · A teammate wants to merge a Codex change immediately because a demo is in an hour. Why is 'green tests, just merge' insufficient here? (Select one)">
    A. It is sufficient; green tests are a full review
    B. Green tests are one evidence type; a review also reads the diff for scope and confirms the change matches the intended, bounded task — and a deadline raises, not lowers, the cost of a bad merge
    C. Tests are unreliable, so ignore them
    D. Demos never need working code

    **Answer: B.** Tests are necessary but not a substitute for reading the diff and confirming scope, and time pressure increases the cost of shipping a bad change. Green tests are not a full review (A), tests are still valuable (C), and demos absolutely need working code (D).
  </AccordionItem>

  <AccordionItem title="Q14 · A developer keeps re-running the same vague prompt at higher and higher reasoning effort with poor results. What is the root cause and fix? (Select one)">
    A. The model is too weak; only Astra will work
    B. The task is under-scoped; adding a clear goal, boundaries and a definition of done will help more than more effort
    C. The CLI is broken; reinstall it
    D. Tests are missing; that is unrelated

    **Answer: B.** Poor results from a vague prompt are usually a scoping problem, not an effort problem; scoping the task fixes it. A bigger model (A) still sprawls on a vague task, the CLI is not implicated (C), and while tests help, the immediate cause is scope (D).
  </AccordionItem>

  <AccordionItem title="Q15 · A team wants three independent features built locally at the same time without their branches colliding. Which approach fits BEST? (Select one)">
    A. Do them one at a time on the same branch
    B. Use separate git worktrees (or branches) per feature, or parallel cloud tasks
    C. Commit all three to main directly
    D. Disable the tests to go faster

    **Answer: B.** Separate worktrees or branches (or parallel cloud tasks) let independent work proceed at once without collisions. Serialising (A) wastes the parallelism, committing to main (C) is unsafe, and disabling tests (D) removes verification.
  </AccordionItem>

  <AccordionItem title="Q16 · Which single practice most directly delivers on the Codex objective 'verify changes and provide clear evidence for review'? (Select one)">
    A. Choosing the most expensive model
    B. Running the tests yourself, reading the diff, and attaching the passing run plus a regression test to the PR
    C. Writing the longest possible prompt
    D. Merging quickly to keep momentum

    **Answer: B.** Independently verifying and attaching evidence (test run, focused diff, regression test) is exactly what the objective asks for. Model choice (A) and prompt length (C) do not produce evidence, and merging quickly (D) skips verification altogether.
  </AccordionItem>
</Accordions>

## Key takeaways

- Scope drives quality: goal + boundaries + definition of done + constraints beats a vague one-liner.
- Put durable context in `AGENTS.md`: build/test commands, conventions, and no-go areas.
- Work review-first: plan → approve → implement → verify → merge; keep the human before the irreversible step.
- `codex exec` runs non-interactive, scripted and CI work; `codex` interactive is for steering and review loops.
- Verify by running the tests yourself and reading the diff — the model's "tests pass" is a lead, not evidence.
- Give reviewers evidence: a focused diff, a passing test run, a regression test, an auto-review summary.
- Use git worktrees (or Codex cloud) to isolate long-running or parallel work from your main tree.
- Poor results from a vague prompt usually need better scope, not more reasoning effort.
