# D5 · Review Points and Human Oversight

Placing human review where it pays off – mandatory review triggers, sampling-based review for volume work, and the cost of a misplaced gate.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is worth **16%** of the mock — roughly **8 of 50 items**. It tests where a human belongs in an otherwise automated workflow: not everywhere (that destroys the value) and not nowhere (that courts disaster), but at the points where a mistake would be expensive or irreversible. The recurring judgment is *placement* — putting the gate at the risky step, and choosing between gating every item and sampling a stream of them.

## What you need to know

Human oversight in a workflow means deciding **where** a person inspects or approves output, **when** review is mandatory versus optional, and **how** review scales when volume is high. A **mandatory review trigger** is any point where an error would be irreversible, regulated, external-facing, or touch sensitive data — those steps require a human gate regardless of how good the output looks. For **high-volume** work, gating every item is impractical, so you **sample**: review a representative fraction, track the error rate, and tighten sampling if quality drifts. A **misplaced gate** is costly in both directions — a gate on a trivial high-volume step throttles throughput for no benefit, while a missing gate on a risky step lets an expensive mistake through.

## Learning objectives

By the end of this page you should be able to:

1. **Identify** mandatory review triggers — irreversible, regulated, external, sensitive — in a workflow.
2. **Place** a review gate at the step where a mistake is most expensive, not by default everywhere.
3. **Design** sampling-based review for high-volume work and decide the sampling rate.
4. **Diagnose** the cost of a misplaced gate — both the throttling gate and the missing gate.
5. **Choose** the review mode — gate, sample, or spot-check — that matches the step's stakes and volume.

---

## 5.1 The three review modes

```text
GATE       every item is reviewed/approved before it proceeds
           ▸ use for irreversible / regulated / external / sensitive steps

SAMPLE     a representative fraction is reviewed; error rate is tracked
           ▸ use for high-volume, correctable, lower-consequence steps

SPOT-CHECK an occasional glance to confirm nothing has drifted
           ▸ use for low-stakes, stable, low-volume steps
```

| Mode | Coverage | Cost | Best for |
| --- | --- | --- | --- |
| **Gate** | 100% | High per item | Irreversible/regulated/external/sensitive |
| **Sample** | A fraction, tracked | Moderate, scales | High-volume, correctable work |
| **Spot-check** | Occasional | Low | Stable, low-stakes tasks |

The art is matching the mode to the step: a gate where a mistake is expensive, a sample where volume is high and mistakes are correctable, a spot-check where little is at stake.

:::tip[Assessment signal]
"Every item must be approved before sending/publishing/paying" → **gate**. "Thousands per day, occasional errors are caught later" → **sample**. A stem that puts a gate on a huge low-stakes stream, or removes the gate from an irreversible step, is describing a *misplaced* gate — the wrong-answer pattern.
:::

## 5.2 Mandatory review triggers

Some steps require a human gate no matter how confident or fluent the output. The triggers are the same ones that drove the risk axis in [Domain 1](/openai/applied-ai/domains/d1-finding-and-scoping-opportunities/).

| Trigger | Example | Why it is mandatory |
| --- | --- | --- |
| **Irreversible** | Sending an email, publishing, issuing a payment, deleting data | Cannot be undone once it goes |
| **Regulated** | Medical, legal, financial advice; HR decisions; disclosures | Legal/compliance exposure |
| **External-facing** | Customer, press, regulator, public communications | Reputational and contractual risk |
| **Sensitive data / people** | PII, protected characteristics, safety-critical instructions | Harm and fairness exposure |

If *any* trigger is present at a step, that step gets a gate. Fluency, speed and high aggregate accuracy do **not** exempt it — a 98%-accurate model still needs a human before an irreversible regulated action.

## 5.3 Placement — the gate goes at the risky step

Oversight is about *where*, not *whether*. The gate belongs immediately **before the risky action**, so the human reviews exactly what will go out.

```text
Draft reply ─► [review inside chat] ─► SEND   ← gate goes HERE, before send
                                       (irreversible/external)

Extract records ─► categorise ─► [sample check] ─► write to CRM
                                  ↑ sampling, not a gate on every record
```

Placing the gate too early (reviewing a draft that will change before the risky step) wastes the review; placing it too late (after the irreversible action) is no gate at all. The reviewer should see the *final* artefact at the *last reversible moment*.

## 5.4 Sampling-based review for volume work

When a step runs hundreds or thousands of times, gating every item is impossible without erasing the value. Sampling reviews a representative fraction and uses the observed error rate to manage risk.

| Design choice | Guidance |
| --- | --- |
| **Rate** | Higher when stakes are higher or the process is new; lower once it proves stable |
| **Selection** | Representative (and oversample the risky slices — e.g. high-value items) |
| **Response to drift** | If sampled error rate rises, tighten sampling or add a gate on the affected slice |
| **New-process burn-in** | Start with heavy sampling; relax as evidence of quality accumulates |

**Worked example.** A workflow categorises 800 receipts a week. Gating all 800 defeats the purpose. Instead: sample 5%, oversample receipts above a value threshold, review at week's end against acceptance criteria, and if the error rate on the sample exceeds the tolerance, raise the rate or gate the high-value slice. This is exactly the "automate categorisation, sample at reconciliation" pattern from Domain 1's scenario.

:::tip[Assessment signal]
"How do you keep quality on a high-volume automated step without reviewing everything?" is asking for sampling — a tracked fraction, oversampling the risky slice, tightening on drift. An answer that says "review every one" (throttles) or "trust it, no review" (reckless) is wrong.
:::

## 5.5 The cost of a misplaced gate

A gate is a cost as well as a control. Misplacing it hurts in one of two directions.

<Tabs>
  <TabItem label="Gate too tight (throttling)">
    A human approves every one of 5,000 low-stakes, correctable items a day. The workflow now moves at human speed, the value it was meant to create evaporates, and the reviewer rubber-stamps out of fatigue — so the gate is *both* slow and ineffective. **Fix:** switch to sampling; reserve gates for the risky slice.
  </TabItem>
  <TabItem label="Gate too loose (missing)">
    An irreversible customer email goes out with no human check because "the model is reliable". One wrong figure reaches thousands of customers; it cannot be recalled. **Fix:** add a mandatory gate before the irreversible/external step — the one place a gate is non-negotiable.
  </TabItem>
</Tabs>

The lesson: **more review is not safer if it is in the wrong place.** A single well-placed gate before the irreversible step plus sampling on the volume steps beats a blanket gate on everything.

## 5.6 Review fatigue and rubber-stamping

A gate only works if the reviewer actually reviews. Two failure modes erode gates over time.

| Failure | Cause | Countermeasure |
| --- | --- | --- |
| Rubber-stamping | Too many low-value approvals dull attention | Sample instead of gating; make each gate meaningful |
| Missing the real error | Reviewer checks format, not substance | Give the reviewer the acceptance criteria to check against |
| Alert overload | Everything is flagged, so nothing is | Flag only genuine review triggers; tune thresholds |

The design goal is *few, meaningful gates* the reviewer takes seriously — which is why over-gating (5.5) is self-defeating: it manufactures the fatigue that makes gates fail.

## 5.7 Oversight and the contract

Review checks output against something. That something is the **acceptance criteria** from [Domain 3](/openai/applied-ai/domains/d3-inputs-outputs-and-contracts/). A gate without criteria is a reviewer guessing; a gate *with* criteria is an objective check ("every question answered; no unverified account claim; figures reconcile"). Well-written contracts make oversight cheap and consistent; missing criteria make every review subjective and slow.

---

## Decision framework

**The GATE decision** — for each step, answer these to choose the review mode.

| # | Question | Implication |
| --- | --- | --- |
| **G** — Grave? | Is the step irreversible, regulated, external or sensitive? | If yes → **mandatory gate** before the action, full stop |
| **A** — Amount? | How high is the volume? | High volume + not grave → **sampling**, not a gate on each |
| **T** — Tolerance? | How correctable is an error? | Easily corrected → lighter review; irreversible → gate |
| **E** — Evidence? | Is the process new or proven? | New → heavier sampling/gating; proven → relax on evidence |

Read it as: **grave steps are gated; high-volume correctable steps are sampled; stable low-stakes steps are spot-checked; and you tighten wherever evidence shows drift.** Placement is always *immediately before the risky action*.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Gating every item on a high-volume, low-stakes step | "More review is safer" | Sample a tracked fraction; reserve gates for risky slices |
| No gate before an irreversible/external action | The model seems reliable | Add a mandatory gate at the last reversible moment |
| Placing the gate before the step that changes the output | Reviewing "early" feels proactive | Review the final artefact, immediately before the risky action |
| Gate after the irreversible action has happened | Review is treated as a formality | The gate must precede the point of no return |
| Reviewers rubber-stamping | Too many trivial approvals | Fewer, meaningful gates; sample the rest |
| Reviewing format instead of substance | Format is easy to check | Give reviewers the acceptance criteria to check against |
| Trusting high aggregate accuracy to skip a gate | 98% feels good enough | Aggregate accuracy doesn't cover the irreversible tail; gate it |
| Flagging everything for review | Caution defaults to "flag it" | Flag only genuine triggers; tune thresholds to avoid overload |

## Scenario challenge

**Scenario.** Aisha runs a workflow that handles inbound support at scale: it classifies each ticket, drafts a first reply, and — for a set of "simple" categories (password resets, shipping-status questions) — it can send automatically. For everything else it drafts and a human sends. Volume is ~2,000 tickets/day. Two incidents just happened. First, an auto-sent "shipping status" reply gave a customer a wrong delivery date pulled from a stale field, and it went out with no human check. Second, the team, alarmed, now wants a human to approve *every* reply — all 2,000/day — which would require doubling headcount and would still leave reviewers skimming. Aisha must redesign the oversight.

**Expert reasoning trace.**

1. **Locate the two failures precisely.** Incident one is a *missing gate* on an external-facing, irreversible action (a sent email) whose content depended on possibly-stale data — a mandatory-trigger step that had no gate. Incident two is the over-correction toward a *throttling gate* on all 2,000 items, which will manufacture rubber-stamping and destroy the throughput the automation exists to provide.
2. **Apply the mandatory-trigger test.** Sending an external reply is irreversible and external — a review trigger. But that does *not* mean gating all 2,000: it means the *auto-send* path must be reconsidered, because auto-send removed the human from an irreversible external step entirely.
3. **Separate the risky slice from the safe stream.** The problem is not the whole stream; it is that "shipping status" replies depend on live data that can be stale. So: keep auto-send only for categories whose replies are *self-contained and low-consequence* (a password-reset link is stable), and for any reply that quotes *data that could be stale or wrong* (delivery dates, balances), route to a human gate or verify the data before sending.
4. **Reject the blanket gate.** Gating all 2,000 is the throttling-gate anti-pattern: it doubles cost, invites rubber-stamping (which would have missed the stale date anyway if the reviewer skims), and punishes the 90% of safe tickets. More review in the wrong place is not safer.
5. **Design sampling for the auto-send stream.** For the categories that stay auto-send, add sampling: review a fraction daily, oversample any reply that pulled dynamic data, track the error rate, and tighten (or pull a category back to human-send) if drift appears. During this redesign — a *new* process — start with heavy sampling and relax on evidence.
6. **Wire in acceptance criteria.** Give the human-send reviewers and the samplers explicit criteria ("no unverified dynamic data; every customer question answered; on-brand tone"), so review checks substance, not format, and the stale-date class of error is specifically on the checklist.

**The decision:** *narrow* auto-send to genuinely self-contained low-consequence replies, add a **gate (or a data-verification step)** for any reply quoting dynamic data, and add **sampling** to the remaining auto-send stream — instead of either the original no-gate auto-send or the proposed gate-everything. The redesign puts the gate exactly where the risk is (replies with stale-able data) and uses sampling to hold quality on the safe high-volume stream, avoiding both the missing-gate incident and the throttling over-correction.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "Have a human approve every item to be safe" | More review sounds safer | On high-volume low-stakes work this throttles and breeds rubber-stamping; sample instead |
| "The model is reliable, so auto-send is fine" | High accuracy feels sufficient | Irreversible/external steps need a gate regardless of accuracy |
| "Review the draft early, before the rest of the workflow" | Early review feels proactive | Review the final artefact at the last reversible moment |
| "Add a review step after sending, just in case" | It looks like oversight | A gate after the irreversible action is not a gate |
| "Flag everything for review" | Caution defaults to flagging | Alert overload makes reviewers ignore flags; flag only real triggers |
| "Aggregate accuracy is 98%, skip the gate" | The headline looks strong | The dangerous errors live in the irreversible tail the average hides |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · A workflow sends customer emails automatically with no human check because 'the model is reliable'. What is the problem? (Select one)">
    A. Nothing — reliability justifies auto-send
    B. Sending is irreversible and external, a mandatory review trigger, so a human gate is required before send
    C. The model should be bigger
    D. The temperature is too high

    **Answer: B.** External, irreversible actions require a human gate regardless of model reliability. 'Reliability justifies it' (A) ignores the mandatory trigger. Model size (C) and temperature (D) don't address the missing gate.
  </AccordionItem>

  <AccordionItem title="Q2 · A step runs 3,000 times a day, errors are correctable, and stakes per item are low. Which review mode fits BEST? (Select one)">
    A. Gate every item
    B. Sampling a tracked fraction, oversampling risky slices
    C. No review at all
    D. Deep research on each item

    **Answer: B.** High-volume, correctable, low-stakes work is the textbook case for sampling with tracking. Gating all 3,000 (A) throttles and breeds rubber-stamping. No review (C) is reckless. Deep research (D) is unrelated to oversight.
  </AccordionItem>

  <AccordionItem title="Q3 · Where should a review gate be placed in a draft-then-send workflow? (Select one)">
    A. Before drafting begins
    B. Immediately before the send, so the reviewer sees the final artefact at the last reversible moment
    C. After the email has been sent
    D. It doesn't matter where

    **Answer: B.** The gate belongs just before the irreversible action, reviewing the final content. Before drafting (A) reviews nothing meaningful. After sending (C) is not a gate at all. Placement absolutely matters (D).
  </AccordionItem>

  <AccordionItem title="Q4 · Which of these is a MANDATORY review trigger? (Select one)">
    A. The output is longer than one page
    B. The output was generated quickly
    C. The step issues an irreversible payment
    D. The output uses bullet points

    **Answer: C.** Irreversible actions like issuing a payment mandate a human gate. Length (A), speed (B) and formatting (D) are irrelevant to whether a gate is required.
  </AccordionItem>

  <AccordionItem title="Q5 · A team reacts to one bad auto-sent reply by requiring human approval of all 20,000 daily replies. What is the risk of this over-correction? (Select one)">
    A. None — full review is always best
    B. It throttles throughput, doubles cost, and breeds rubber-stamping that misses real errors anyway
    C. It makes the model less accurate
    D. It violates OpenAI policy

    **Answer: B.** A blanket gate on high-volume work destroys the value and induces fatigue-driven rubber-stamping, so it's both slow and ineffective. Full review is not always best (A). It doesn't change model accuracy (C) or violate policy (D).
  </AccordionItem>

  <AccordionItem title="Q6 · An auto-send reply quoted a delivery date from a stale field and reached the customer. What is the BEST targeted fix? (Select one)">
    A. Gate every reply of every category
    B. Route replies that quote dynamic data (dates, balances) to a gate or verify the data before sending, while keeping self-contained replies automated
    C. Stop using AI for support entirely
    D. Lower the temperature

    **Answer: B.** The risk is specific to replies depending on stale-able data; gate or verify those while leaving genuinely self-contained replies automated. Gating everything (A) is the throttling over-correction. Abandoning AI (C) is disproportionate. Temperature (D) doesn't fix stale data.
  </AccordionItem>

  <AccordionItem title="Q7 · How should sampling rate change for a brand-new automated step? (Select one)">
    A. Start low and never change it
    B. Start with heavier sampling during burn-in, then relax as evidence of quality accumulates
    C. Never sample new steps
    D. Sample only after a customer complains

    **Answer: B.** New processes lack an evidence base, so you sample heavily at first and relax as quality is demonstrated. Starting low and never changing (A) ignores drift and burn-in. Not sampling new steps (C) is exactly backwards. Waiting for complaints (D) is reactive, not oversight.
  </AccordionItem>

  <AccordionItem title="Q8 · Why does over-gating undermine the very safety it seeks? (Select one)">
    A. It makes the model hallucinate
    B. Too many trivial approvals cause reviewer fatigue and rubber-stamping, so real errors slip through anyway
    C. Gates always slow the model's inference
    D. It increases token cost

    **Answer: B.** Excessive gates dull attention and turn review into rubber-stamping, defeating the control. It doesn't affect hallucination (A) or inference speed (C). Token cost (D) isn't the safety issue described.
  </AccordionItem>

  <AccordionItem title="Q9 · What makes a review gate an objective check rather than a reviewer guessing? (Select one)">
    A. A bigger model
    B. Explicit acceptance criteria from the step contract that the reviewer checks against
    C. A longer prompt
    D. Reviewing more items

    **Answer: B.** Acceptance criteria give the reviewer concrete, checkable conditions, turning review from subjective judgment into an objective check. Model size (A), prompt length (C) and volume (D) don't make a review objective.
  </AccordionItem>

  <AccordionItem title="Q10 · Which TWO steps in a workflow REQUIRE a mandatory gate rather than sampling? (Select two)">
    A. Publishing a press release externally
    B. Categorising internal receipts that reconcile monthly
    C. Filing a regulatory disclosure
    D. Drafting internal brainstorming notes
    E. Summarising a meeting for your own records

    **Answer: A and C.** External publication and regulatory filing are irreversible/regulated/external — mandatory gates. Receipt categorisation (B) is high-volume correctable work suited to sampling. Brainstorming (D) and personal summaries (E) are low-stakes, needing at most a spot-check.
  </AccordionItem>

  <AccordionItem title="Q11 · A reviewer is approving 500 items an hour and only glancing at formatting. What TWO changes improve oversight? (Select two)">
    A. Switch most of the stream to sampling so each remaining gate is meaningful
    B. Give the reviewer the acceptance criteria to check substance, not just format
    C. Increase the number of items reviewed to 1,000/hour
    D. Remove all review
    E. Use a bigger model so review isn't needed

    **Answer: A and B.** Fewer, meaningful gates (via sampling) plus criteria-driven substantive review fix rubber-stamping. Doubling the load (C) worsens fatigue. Removing review (D) is reckless. A bigger model (E) doesn't eliminate the need for oversight on risky steps.
  </AccordionItem>

  <AccordionItem title="Q12 · A workflow shows 97% aggregate accuracy, so a manager proposes removing the gate before an irreversible legal filing. What is the BEST response? (Select one)">
    A. Agree — 97% is high enough
    B. Keep the gate: aggregate accuracy hides the irreversible, regulated tail where a single error is catastrophic
    C. Remove the gate but add one after filing
    D. Raise the accuracy target to 99% and then remove the gate

    **Answer: B.** A high average does not cover the low-frequency, high-consequence errors on an irreversible regulated step, so the mandatory gate stays. 97% 'good enough' (A) ignores the tail risk. A gate after filing (C) is no gate. A higher target (D) still can't make an irreversible regulated action safe to leave ungated.
  </AccordionItem>
</Accordions>

## Key takeaways

- Oversight is about **placement**: gate the risky step, sample the high-volume stream, spot-check the stable low-stakes work.
- **Mandatory triggers** — irreversible, regulated, external, sensitive — require a human gate regardless of fluency, speed, or aggregate accuracy.
- Put the gate **immediately before the risky action**, reviewing the final artefact at the last reversible moment.
- For volume, **sample** a tracked fraction, oversample risky slices, and tighten on drift; sample heavily during a new process's burn-in.
- A **misplaced gate costs both ways**: gate-everything throttles and breeds rubber-stamping; a missing gate lets an irreversible mistake through.
- Fewer, meaningful gates beat blanket review; **acceptance criteria** turn a gate from guessing into an objective check.
- More review is not safer if it is in the wrong place.
