# D5 · Scaling Across Teams and Systems

Coordinating parallel workstreams and integrating them safely, standardising AGENTS.md across repos, the Codex SDK, App Server, GitHub Action and integrations, measuring adoption and quality, and the failure modes of scaling an agent.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

This domain is worth **16%** of the mock — roughly **8 of 50 items**. It is the "Scale Codex Across Governed Teams and Systems" course made concrete: coordinate parallel workstreams and integrate their outputs safely, standardise configuration across many repositories, wire Codex into the tools your teams already use, measure whether scaling is actually helping, and recognise the failure modes that only appear at scale. The exam rewards *safe integration and standardisation* over raw parallelism.

## What you need to know

Running one agent well (D2) is different from running many across an organisation. At scale you coordinate **parallel workstreams** and integrate their outputs **safely** — independent branches, verified changes, controlled merges — rather than letting agents collide. You standardise **`AGENTS.md`** across repositories so behaviour is predictable everywhere. You embed Codex in delivery through the **Codex SDK**, **App Server**, the **GitHub Action** and integrations with **GitHub, GitLab (beta), Slack and Linear**. And you measure both **adoption and quality**, because more agent activity is not automatically more value. The characteristic failure modes — merge conflicts between parallel agents, config drift, unreviewed changes at volume, and mistaking activity for outcomes — are all testable.

## Learning objectives

By the end of this page you should be able to:

1. **Coordinate parallel workstreams** and integrate their outputs safely.
2. **Standardise `AGENTS.md`** and configuration across many repositories.
3. **Integrate Codex** into delivery via the Codex SDK, App Server, GitHub Action, and GitHub/GitLab/Slack/Linear.
4. **Measure adoption and quality**, not just activity.
5. **Recognise and mitigate the failure modes** of scaling an agent across many repositories.

---

## 5.1 Parallel workstreams and safe integration

Scale means several Codex tasks running at once — across branches, repos or teams. The value is throughput; the risk is unsafe integration.

| Safe-integration practice | Why it matters at scale |
| --- | --- |
| One workstream per branch / worktree | Parallel agents do not overwrite each other |
| Verify each change independently (D2) | Volume multiplies the cost of one bad merge |
| Controlled merge (review + CI gates) | A human/automated gate stays between agent output and main |
| Integrate incrementally | Small, verified merges beat one giant reconciliation |

```text
        ┌── workstream A (branch) ──┐
task ──►┤── workstream B (branch) ──┤──► VERIFY each ──► CONTROLLED MERGE ──► main
        └── workstream C (branch) ──┘     (tests, diff)   (review + CI gate)
              parallel, isolated
```

:::tip[Assessment signal]
`Several agents at once`, `parallel`, `they overwrote each other` → isolate per branch/worktree and integrate through a controlled merge. Parallelism without isolation and a merge gate is the trap.
:::

## 5.2 Standardising `AGENTS.md` across repositories

One good `AGENTS.md` (D2) makes one repo predictable. At scale, *consistency across repos* is the lever: an engineer or an agent moving between repositories should meet the same conventions, build/test commands and no-go areas.

| Approach | Effect |
| --- | --- |
| A shared `AGENTS.md` template across repos | Predictable agent behaviour everywhere |
| Managed configuration (D4) alongside it | Consistent model, permissions and extensions |
| Periodic drift checks | Catch repos that have diverged |

Standardisation is what turns "Codex works on my repo" into "Codex works the same on all our repos", and it is what makes parallel workstreams safe to integrate — because each was produced under the same rules.

## 5.3 Programmatic and delivery integration

Codex plugs into how software is actually delivered.

| Integration | What it is | Reach for it when |
| --- | --- | --- |
| **Codex SDK** | Build Codex into your own tools/automation | You want programmatic control in your systems |
| **App Server** | Host Codex-backed applications | You are serving Codex capability to your org's apps |
| **GitHub Action** | Codex in a GitHub workflow | You want Codex steps in CI on GitHub |
| **GitHub** | Native GitHub integration | Repo, PR and workflow context on GitHub |
| **GitLab (beta)** | GitLab integration (beta) | You are on GitLab; note the beta status |
| **Slack** | Codex in Slack | Team-facing requests and notifications |
| **Linear** | Codex in Linear | Issue-driven work |

:::tip[Assessment signal]
`In our GitHub CI` → **GitHub Action**. `Build Codex into our own tooling` → **Codex SDK**. `We use GitLab` → GitLab integration, and remember it is **beta**. `From a Slack request` / `from a Linear issue` → the Slack / Linear integrations.
:::

## 5.4 Measuring adoption and quality

At scale, "more Codex activity" is not the goal — better, faster, safe delivery is. Measure both sides.

| Dimension | Example signals | Source |
| --- | --- | --- |
| **Adoption** | Active users, tasks run, teams onboarded | Workspace analytics / Analytics API (D4) |
| **Quality** | Review pass rate, defects, revert rate, security findings | CI, Codex Security, your delivery metrics |

```text
ACTIVITY  ≠  VALUE
│              │
tasks run   review pass rate, defects avoided,
per week    security findings resolved, cycle time

Measure BOTH; a rise in activity with falling quality is a warning, not a win.
```

:::tip[Assessment signal]
`Adoption is up, is it working?` → pair adoption metrics with **quality** metrics (review pass rate, defects, security findings). An answer that celebrates activity alone is the trap.
:::

## 5.5 Failure modes of scaling an agent

Some problems only appear once many agents run across many repos. Know the pattern and the mitigation.

| Failure mode | How it appears | Mitigation |
| --- | --- | --- |
| Parallel collisions | Two agents edit the same code and overwrite each other | One workstream per branch/worktree; controlled merge |
| Config drift | Repos diverge; Codex behaves inconsistently | Standardise `AGENTS.md`; managed configuration; drift checks |
| Unreviewed volume | Too many changes to review, so review lapses | Keep the review gate; scope smaller; use auto-review as evidence, not a replacement |
| Activity ≠ outcomes | Dashboards up, quality flat or down | Measure quality alongside adoption |
| Ungoverned auth/permissions at scale | Personal tokens, broad permissions proliferate | Workload identity/service accounts; least privilege (D4) |
| Security debt at volume | Vulnerabilities merge faster than they are found | Codex Security in CI on every repo |

## Decision framework

Use **SCALE SAFELY (SISM)**: **S**tandardise, **I**solate, **S**afe-merge, **M**easure.

| Step | Question | The move |
| --- | --- | --- |
| **Standardise** | Do all repos share the same rules? | A common `AGENTS.md` template + managed configuration + drift checks |
| **Isolate** | Can parallel work collide? | One workstream per branch/worktree; cloud or Ultra for parallelism |
| **Safe-merge** | Is every change verified before main? | Independent verification (D2) + a controlled merge with review and CI gates |
| **Measure** | Is scaling actually helping? | Adoption *and* quality metrics; treat activity-without-quality as a warning |

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Running parallel agents on the same branch | It seems simpler | One workstream per branch/worktree; integrate through a controlled merge |
| Letting each repo's `AGENTS.md` diverge | No one owns consistency | Standardise a template and run drift checks |
| Celebrating adoption numbers alone | Activity is easy to count | Pair adoption with quality (review pass rate, defects, security findings) |
| Skipping review because volume is high | There is too much to check | Keep the review gate; scope smaller; auto-review is evidence, not a replacement |
| Using the wrong integration surface | Reaching for the familiar one | GitHub Action for GitHub CI, SDK for your own tooling, GitLab (beta) on GitLab |
| Forgetting GitLab is beta | It looks like the others | Note the beta status when planning a GitLab rollout |
| Merging fast without CI security scans | Speed pressure at scale | Codex Security in CI on every repo before merge |
| Ungoverned auth spreading at scale | Copying the pilot's shortcuts | Workload identity/service accounts and least privilege everywhere (D4) |

## Scenario challenge

**Scenario.** A platform group of 40 engineers across 25 repositories has adopted Codex. Leadership is delighted: tasks-run-per-week has tripled. But three problems have surfaced: two parallel Codex tasks on the same feature branch overwrote each other last week; the 25 repos have drifted so Codex uses different build commands and conventions in each; and the revert rate on merged PRs has quietly climbed while nobody was watching quality. A manager proposes: "keep scaling — the activity numbers prove it is working — and we will fix issues as they come up".

**Expert reasoning trace.**

1. **The headline metric is misleading.** Tripled tasks-per-week is *activity*, not *value*. The rising revert rate is the quality signal that matters, and it is moving the wrong way — so "the numbers prove it is working" is exactly the activity-vs-outcomes trap.
2. **The collision is an isolation failure.** Two agents on one branch overwrote each other because the work was not isolated. The fix is one workstream per branch/worktree, integrated through a controlled merge, using cloud or Ultra for genuine parallelism.
3. **The drift is a standardisation failure.** 25 repos with different build commands and conventions make Codex behave inconsistently and make parallel outputs unsafe to integrate. Standardise a shared `AGENTS.md` template and managed configuration, with periodic drift checks.
4. **The revert rate is an unreviewed-volume / quality failure.** Climbing reverts while "nobody was watching" means the review gate and quality measurement both lapsed under volume. Reinstate the review gate (auto-review as evidence, not replacement) and measure quality alongside adoption.
5. **Reject "fix issues as they come".** At scale, reactive fixes lag the damage; the mitigations are structural (isolate, standardise, safe-merge, measure).

**Exam-correct decision:** treat the revert rate — not tasks-per-week — as the readiness signal; isolate parallel work per branch/worktree with a controlled merge; standardise `AGENTS.md` and configuration across the 25 repos with drift checks; reinstate the review gate and measure quality alongside adoption. **Not** "keep scaling because activity is up", **not** parallel work on shared branches, **not** letting each repo diverge, **not** reactive-only fixes.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Activity is up, so scaling works" | Activity is easy to see | Pair adoption with quality; rising activity with falling quality is a warning |
| "Run parallel agents on the same branch" | It seems simpler | Isolate per branch/worktree; integrate via controlled merge |
| "Each repo can keep its own conventions" | Local ownership feels fine | Standardise `AGENTS.md`; drift makes parallel outputs unsafe |
| "Too much volume to review, so trust the agent" | Review is a bottleneck | Keep the gate; scope smaller; auto-review is evidence, not a replacement |
| "GitLab works just like GitHub here" | The integrations look alike | GitLab is beta; plan accordingly |
| "Use the SDK to add Codex to our GitHub CI" | SDK sounds general | GitHub CI is the GitHub Action; the SDK is for your own tooling |
| "Fix scaling issues reactively" | It defers work | Scaling failures need structural mitigation, not reaction |

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · Two parallel Codex tasks overwrote each other's changes on one feature branch. What is the correct fix? (Select one)">
    A. Run fewer tasks
    B. Give each workstream its own branch or worktree and integrate through a controlled merge
    C. Merge directly to main to avoid the branch
    D. Disable tests to reduce conflicts

    **Answer: B.** Isolating each workstream on its own branch/worktree and merging through a controlled gate prevents collisions while keeping parallelism. Running fewer tasks (A) sacrifices throughput unnecessarily, merging to main (C) is unsafe, and disabling tests (D) removes verification.
  </AccordionItem>

  <AccordionItem title="Q2 · Codex behaves inconsistently across 20 repositories with different build commands. What is the BEST remedy? (Select one)">
    A. Accept the inconsistency
    B. Standardise a shared `AGENTS.md` template and managed configuration across the repos, with drift checks
    C. Use a larger model everywhere
    D. Give each engineer admin

    **Answer: B.** Standardising `AGENTS.md` and configuration makes agent behaviour predictable across repos and makes parallel outputs safe to integrate. Accepting drift (A) leaves the problem, a bigger model (C) does not fix inconsistent context, and universal admin (D) is a governance error.
  </AccordionItem>

  <AccordionItem title="Q3 · You want to add Codex steps to a GitHub CI workflow. Which integration fits BEST? (Select one)">
    A. The Codex SDK
    B. The GitHub Action
    C. The Slack integration
    D. The Linear integration

    **Answer: B.** The GitHub Action embeds Codex into a GitHub CI workflow. The SDK (A) is for building Codex into your own tooling, and Slack (C) and Linear (D) are team/issue integrations, not CI steps.
  </AccordionItem>

  <AccordionItem title="Q4 · Leadership reports Codex adoption tripled and concludes scaling is a success. What is the flaw? (Select one)">
    A. None; more usage is always better
    B. Activity is not value; quality metrics such as review pass rate, defects and revert rate must be measured alongside adoption
    C. Adoption cannot be measured
    D. They should have used a smaller model

    **Answer: B.** Rising activity without a quality view can hide falling outcomes; adoption must be paired with quality metrics. More usage is not automatically better (A), adoption is measurable (C), and model size (D) is unrelated to the measurement flaw.
  </AccordionItem>

  <AccordionItem title="Q5 · Which TWO practices make parallel Codex workstreams safe to integrate? (Select two)">
    A. Isolating each workstream on its own branch or worktree
    B. Merging every branch to main without review to save time
    C. Verifying each change independently before a controlled merge
    D. Sharing one branch across all agents
    E. Turning off CI to speed merges

    **Answer: A and C.** Isolation per branch/worktree and independent verification before a controlled merge are the safe-integration practices. Unreviewed merges (B), a shared branch (D) and turning off CI (E) all remove the safeguards that make parallelism safe.
  </AccordionItem>

  <AccordionItem title="Q6 · A team wants to build Codex capability into their own internal developer tool. Which integration fits BEST? (Select one)">
    A. The GitHub Action
    B. The Codex SDK
    C. The Slack integration
    D. Record & replay

    **Answer: B.** The Codex SDK gives programmatic control to build Codex into your own tools. The GitHub Action (A) is for GitHub CI, Slack (C) is a team surface, and record & replay (D) captures sessions rather than embedding capability.
  </AccordionItem>

  <AccordionItem title="Q7 · A team on GitLab plans to rely on the Codex GitLab integration for a critical launch next week. What should they weigh? (Select one)">
    A. Nothing; it is generally available and identical to GitHub's
    B. The GitLab integration is beta, so they should account for beta risk in a critical launch plan
    C. GitLab is unsupported entirely
    D. They must migrate to GitHub first

    **Answer: B.** The GitLab integration is in beta, which matters when planning a critical launch. It is not GA-equivalent to GitHub (A), GitLab is supported in beta rather than unsupported (C), and migrating to GitHub (D) is not required.
  </AccordionItem>

  <AccordionItem title="Q8 · The revert rate on merged Codex PRs is climbing while task volume rises. What does this indicate and what should be done? (Select one)">
    A. Success; more tasks mean more merges
    B. A quality/unreviewed-volume failure: reinstate the review gate, scope smaller, and treat the revert rate as a key readiness signal alongside adoption
    C. The model is too small; upgrade it
    D. Reverts are normal; ignore them

    **Answer: B.** A rising revert rate against rising volume signals that review lapsed under scale; the fix is to restore the gate, scope smaller and measure quality. It is not success (A), the cause is process not model size (C), and a climbing revert trend should not be ignored (D).
  </AccordionItem>

  <AccordionItem title="Q9 · How should security scanning be handled when Codex is used across 30 repositories? (Select one)">
    A. Manually, on the repos someone remembers
    B. Codex Security integrated into CI on every repository so code is scanned before merge
    C. Only on the largest repository
    D. Only after a breach

    **Answer: B.** At scale, Codex Security must run in CI on every repo so no repository merges unscanned code. Manual scanning (A) and scanning only one repo (C) leave gaps, and post-breach scanning (D) is too late.
  </AccordionItem>

  <AccordionItem title="Q10 · What is the relationship between adoption metrics and quality metrics at scale? (Select one)">
    A. Adoption alone proves value
    B. Both must be tracked; adoption shows usage while quality (review pass rate, defects, security findings) shows whether the usage produces good outcomes
    C. Quality replaces adoption entirely
    D. Neither is measurable

    **Answer: B.** Adoption and quality are complementary: usage without quality can mask declining outcomes, so both are tracked. Adoption alone does not prove value (A), quality does not replace adoption measurement (C), and both are measurable (D).
  </AccordionItem>

  <AccordionItem title="Q11 · A team wants Codex to act on requests raised in Linear issues. Which integration fits BEST? (Select one)">
    A. The GitHub Action
    B. The Linear integration
    C. The Codex SDK only
    D. Workload identity federation

    **Answer: B.** The Linear integration connects Codex to issue-driven work in Linear. The GitHub Action (A) is for GitHub CI, the SDK (C) is for custom tooling, and workload identity federation (D) is an auth mechanism, not an integration surface.
  </AccordionItem>

  <AccordionItem title="Q12 · Which TWO are characteristic failure modes that emerge specifically when scaling Codex across many repositories? (Select two)">
    A. Configuration drift causing inconsistent agent behaviour
    B. A single engineer writing a clear prompt
    C. Unreviewed change volume causing the review gate to lapse
    D. Choosing the correct model for one task
    E. Adding an `AGENTS.md` to one repo

    **Answer: A and C.** Config drift across repos and review lapsing under change volume are classic at-scale failure modes. A clear single prompt (B), a correct single-task model choice (D) and adding one `AGENTS.md` (E) are healthy single-task practices, not scaling failures.
  </AccordionItem>
</Accordions>

## Key takeaways

- Parallelism is throughput; safety is isolation plus a controlled merge. Give each workstream its own branch/worktree and verify before merging.
- Standardise `AGENTS.md` and configuration across repositories, with drift checks, so agent behaviour is predictable everywhere.
- Match the integration to the job: GitHub Action for GitHub CI, Codex SDK for your own tooling, App Server for hosting, and the GitHub/GitLab (beta)/Slack/Linear integrations for their platforms.
- Remember GitLab is beta when planning a critical rollout.
- Measure adoption *and* quality; rising activity with falling quality is a warning, not a win.
- Scaling failure modes — collisions, config drift, unreviewed volume, activity-as-value, ungoverned auth, security debt — need structural mitigation, not reactive fixes.
