# D6 · Verifying and Evaluating Output

A verification protocol proportional to stakes, the provenance test, citation checking, and deciding when human review is required before you act on ChatGPT output.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is worth **14% of the mock – roughly 8 of 60 items**. It is where Domain 2's warning ("fluency is not accuracy") becomes a routine. It tests whether you treat every actionable output as a draft to verify, whether you scale verification to the stakes, and whether you can trace a claim to its source and check a citation properly. Nearly every item reduces to one question: *what could be wrong here, and how would I know?*

## What you need to know

Because the model generates probable text rather than verified truth, you own the accuracy of anything you act on. Verification should be **proportional to stakes**: a low-stakes draft needs a sanity read, a high-stakes external or regulated output needs line-by-line checking and a human review gate. The **provenance test** asks where each claim came from and chooses the cheapest sufficient check. **Citations must be opened**, not trusted on the strength of a real-sounding name. Internal-consistency errors (totals, contradictions) are the cheapest to catch and never need an external source.

## Learning objectives

By the end of this page you should be able to:

1. Apply a **verification protocol proportional to stakes**.
2. Use the **provenance test** to pick the cheapest sufficient check for a claim.
3. **Check citations** correctly instead of trusting their appearance.
4. Detect **internal-consistency** failures without external sources.
5. Decide when **human review** is mandatory before acting.
6. Avoid weak "verification" methods that only feel like checks.

---

## 6.1 Verification proportional to stakes

Not every output needs the same rigour. Match effort to consequences.

```text
Stakes ───────────────────────────────────────────────────────►
Low                      Medium                        High
brainstorm, first        internal memo, customer       external publication,
draft, personal note     email draft, slide deck       legal/medical/financial,
                                                        regulatory filing, pricing
Sanity-read; fix         Verify every number, name,    Line-by-line check;
obvious errors           date, citation; check         independent expert review;
                         internal consistency          documented sign-off
```

Signals in a stem that push toward **high rigour**: "will be published", "regulator", "client-facing", "contract", "compliance", "medical", "financial figures", "press release", "irreversible", "hiring decision".

:::tip[Assessment signal]
The correct answer picks the *cheapest check that is sufficient for the stakes*. Over-checking a throwaway draft and under-checking a board pack are both wrong; the exam rewards proportionality.
:::

## 6.2 The provenance test

For any claim that matters, ask **where did this come from?** – then use the cheapest check that fits the source.

| Source of the claim | Cheapest sufficient check |
| --- | --- |
| Quoted/paraphrased from a document you supplied | Locate the supporting sentence; confirm the paraphrase |
| Retrieved with a citation (Search / deep research) | Open the source; confirm it says what the model claims |
| From the model's own knowledge, no source | Treat as a lead; verify against an independent source |
| A number the model computed | Recompute with a tool or by hand |
| Internal (totals, cross-references) | Recompute or compare sections – no external source needed |

```text
a claim you will act on
   │
   ├─ came from supplied text? ── check the passage
   ├─ came with a citation? ───── open the source
   ├─ came from model memory? ─── verify independently
   ├─ was computed? ──────────── recompute with a tool
   └─ internal (totals)? ─────── recompute; no web needed
```

The provenance test prevents two opposite errors: reaching for the web to check arithmetic (over-engineered), and trusting a model-memory claim without any independent source (under-checked).

## 6.3 Checking citations

A hallucinated citation blends real-sounding elements – a plausible title, real-looking authors, a real journal name – around a claim the source may not support or that may not exist at all. The only valid check is to **open the source and confirm it exists and actually says what the model claims.**

| Weak "check" | Why it fails | Valid check |
| --- | --- | --- |
| The journal name is real | Fabrications reuse real names | Open the cited item |
| Ask for the DOI in the same chat | The DOI can also be fabricated | Look the source up independently |
| It appears twice in the answer | Repetition is not evidence | Open it once, properly |
| Regenerate and see if it recurs | Tests stability, not existence | Confirm it exists and supports the claim |

A single fabricated citation is a red flag for the **whole** output: once one fails, move from spot-checking to verifying every citation and any claim resting on one.

## 6.4 Internal consistency – the cheapest check

Some errors need no external source at all: you only compare the output with itself and its inputs.

| Inconsistency | How it hides | Detection |
| --- | --- | --- |
| Table rows do not sum to the stated total | The prose reads fine | Recompute the rows |
| Summary recommends A; body favours B | Both parts sound reasonable | Compare summary to body |
| A defined term rendered three ways | Each instance looks correct | Compare occurrences; supply a glossary |
| "Three risks" list has four items | Easy to skim past | Count against the claim |
| Timeline where effect precedes cause | Fluent narrative | Check the ordering |

Because these checks are free and fast, they are usually the **first** move on any output you will act on – catching them before spending effort on external verification.

## 6.5 When human review is required

Human review is **mandatory**, not optional, when any of these hold:

1. **Irreversibility** – sending, publishing, paying, deleting.
2. **Regulated domain** – legal, medical, financial advice, HR/hiring, safety-critical instructions.
3. **External audience** – customers, press, regulators, the public.
4. **Personal data or protected characteristics** – anything touching PII or fairness.
5. **Organisational policy says so** – many AI-use policies name review gates.
6. **Novel or high-uncertainty content** – the model is operating outside well-trodden ground.

Human review is **optional** for internal brainstorming, first drafts you will rewrite, and reformatting your own content.

:::tip[Assessment signal]
Options that skip review for an irreversible, regulated, external or personal-data output are wrong *regardless of how confident the output reads*. Look for "human-in-the-loop", "review gate", "expert sign-off" in correct answers.
:::

## 6.6 Weak checks that only feel like verification

Several popular "checks" do nothing and appear as distractors constantly.

<Tabs>
  <TabItem label="Same-chat self-review">
    Asking "are you sure?" in the same conversation reuses the reasoning that produced the error. Use an independent check: a fresh chat, a source, or a human.
  </TabItem>
  <TabItem label="Two tools agreeing">
    Two models giving the same answer is weak evidence – they can share the same wrong prior. Verify against an authoritative source, not a second model.
  </TabItem>
  <TabItem label="Regenerating">
    Re-running the prompt tests stability, not truth. A confidently repeated error is still an error.
  </TabItem>
  <TabItem label="Lowering temperature">
    Reduces variance, not ungroundedness. It does not make a claim true.
  </TabItem>
</Tabs>

## Decision framework

Use the **stakes-and-provenance protocol**: first set the stakes, then, for each claim, apply the provenance test at the cheapest sufficient level, and route high-stakes outputs through a human gate.

| Step | Action | Output |
| --- | --- | --- |
| 1. Set stakes | Is it irreversible, regulated, external or personal-data? | Low / medium / high rigour |
| 2. Free checks first | Recompute totals; compare summary vs body; count claims | Internal-consistency pass/fail |
| 3. Provenance per claim | Supplied text → passage; citation → open it; memory → independent source; computed → recompute | Each checkable claim verified |
| 4. Gate | High stakes → human/expert sign-off before acting | Documented review |
| 5. Record | Note what was verified if others rely on it | Auditable trail |

The value: it stops both under-checking (shipping unverified high-stakes output) and over-checking (web-searching to verify arithmetic), and it always starts with the free internal checks.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Forwarding polished output unverified | Fluency reads like accuracy | Verify anything you act on, proportional to stakes |
| Trusting a citation because the name looks real | Real-sounding elements | Open the source and confirm it supports the claim |
| Web-searching to check a table's arithmetic | Over-caution | Recompute internally; consistency needs no external source |
| Skipping review for an 'internal' figure feeding a decision | The 'internal' label | Stakes follow downstream use; verify it |
| Asking 'are you sure?' in the same chat | It feels like a check | Use an independent check |
| Treating two AIs agreeing as confirmation | Agreement feels strong | Verify against an authoritative source |
| Fixing only the paragraph with a bad citation | It seems contained | One fabrication invalidates the sample; check all |
| Over-checking a throwaway brainstorm | Caution habit | Match rigour to low stakes; sanity-read only |

## Scenario challenge

**Scenario.** Lena drafts a client-facing quarterly report with ChatGPT. It reads beautifully: a summary table of segment revenue, a recommendation to double down on Segment B, and four cited industry sources. She notices the segment rows sum to a different figure than the headline total, and one of the four citations is a report she cannot locate anywhere. A teammate says, "it reads great and we're late — just send it." The report goes to the client and informs their spend.

**Expert reasoning trace.**

1. **Set the stakes.** Client-facing, decision-driving, hard to retract — this is high rigour with a mandatory human review gate. "It reads great, send it" is the fluency trap and is rejected outright.
2. **Do the free checks first.** The rows not summing to the headline is an internal-consistency failure, checkable with no external source — recompute. That alone means the table cannot be trusted as-is.
3. **Apply the provenance test to citations.** The unlocatable report is a probable fabrication; because one citation failed, the sample-based trust collapses, so *every* citation and the claims resting on them must be checked — not just the one bad paragraph.
4. **Reject the weak checks.** Asking the chat "are you sure about the numbers?" reuses the biased reasoning; regenerating tests stability, not truth; lowering temperature changes variance, not grounding.
5. **Route the gate.** Even after correction, a client-facing report needs human/expert sign-off before release, and Lena should record what she verified because the client will rely on it.

**Exam-correct outcome:** recompute the totals, verify every citation (treating the whole output as unverified once one failed), correct or remove unsupported claims, and pass a human review gate before sending — not trusting the fluency, a same-chat self-check, or a regeneration.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "It's detailed and well-formatted, so send it" | Polish reads as quality | Fluency is not accuracy; verify before acting |
| "The citation's journal name is real, so it's valid" | Real names feel legitimate | Fabrications reuse real names; open the source |
| "Ask it in the same chat if it's sure" | Feels like verification | Same-chat self-review inherits the bias |
| "Two AIs agree, so it's confirmed" | Agreement feels strong | They can share a wrong prior; use an authoritative source |
| "It's internal, no need to check the figure" | The 'internal' label | Stakes follow downstream use |
| "Search the web to check the table math" | Being thorough | Internal consistency is recomputed, not searched |
| "Fix just the bad-citation paragraph" | Seems contained | One fabrication invalidates the sample; verify all |

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · ChatGPT returns five precise, unsourced market statistics for a niche segment. What is the BEST next step? (Select one)">
    A. Use them; they are internally consistent.
    B. Treat them as leads, ask for sources, and verify each against an independent source before use.
    C. Ask in the same chat whether it is confident.
    D. Lower the temperature and regenerate.

    **Answer: B.** Precise unsourced numbers on a niche topic are the hallucination signature; provenance says verify independently. A mistakes consistency for truth. C reuses the biased reasoning. D changes variance, not grounding.
  </AccordionItem>

  <AccordionItem title="Q2 · A contract summary says the notice period is 60 days. The contract was uploaded. What MOST appropriately validates this? (Select one)">
    A. Accept it; the file was supplied.
    B. Ask for the exact clause, then open the contract and confirm it.
    C. Ask a second AI the same question.
    D. Summarise the contract again and compare.

    **Answer: B.** Provenance test: a claim from supplied text is checked by locating the supporting passage. A over-trusts extraction. C is weak cross-model agreement. D tests stability, not correctness.
  </AccordionItem>

  <AccordionItem title="Q3 · Which check is FREE and should usually come FIRST on an output you will act on? (Select one)">
    A. Opening every web citation.
    B. Recomputing totals and comparing the summary to the body for internal consistency.
    C. Asking a second model.
    D. Regenerating the answer.

    **Answer: B.** Internal-consistency checks need no external source and are fast, so they come first. A is valuable but costlier and comes later. C and D are weak checks.
  </AccordionItem>

  <AccordionItem title="Q4 · A citation has a plausible title, real-looking authors and a real journal name. How do you validate it? (Select one)">
    A. Accept it; the journal is real.
    B. Open the cited source and confirm it exists and supports the claim.
    C. Ask for the DOI in the same chat and accept it.
    D. Regenerate to see if it recurs.

    **Answer: B.** Fabricated citations reuse real-sounding elements, so only opening the source validates it. A trusts appearance. C can also be fabricated. D tests recurrence, not existence.
  </AccordionItem>

  <AccordionItem title="Q5 · An internal-only memo contains a revenue figure that will feed next year's hiring plan. A colleague says 'it's internal, skip the check'. What is the BEST response? (Select one)">
    A. Agree; internal content needs no checking.
    B. Verify the figure, because its downstream use in a hiring decision makes it consequential.
    C. Only check the tone.
    D. Ask the model if it is confident.

    **Answer: B.** Stakes follow the downstream use, not the label; a figure feeding hiring must be verified. A ignores downstream stakes. C checks the wrong thing. D is a weak same-chat check.
  </AccordionItem>

  <AccordionItem title="Q6 · A research summary has eight citations; the analyst opens two and one does not exist. What should they do? (Select one)">
    A. Use it; half the sample was fine.
    B. Discard only the paragraph with the bad citation.
    C. Treat the whole summary as unverified: check every citation and re-verify claims resting on failed ones.
    D. Ask the model to replace the missing citation.

    **Answer: C.** One fabrication invalidates the sample, so verification must become exhaustive. A ignores the failed sample. B leaves other risks. D papers over the problem.
  </AccordionItem>

  <AccordionItem title="Q7 · Which action correctly matches the provenance of a number the model COMPUTED? (Select one)">
    A. Open a web citation.
    B. Recompute it with a tool or by hand.
    C. Ask a second model.
    D. Lower the temperature.

    **Answer: B.** A computed number is verified by recomputation, not by web search or cross-model agreement. A is for cited claims. C and D are weak checks.
  </AccordionItem>

  <AccordionItem title="Q8 · Which output MOST clearly requires a mandatory human review gate before acting? (Select one)">
    A. A brainstorm list you will rewrite.
    B. A press release naming a third party, going to media today.
    C. A draft outline for your own notes.
    D. A reformatting of your own paragraph.

    **Answer: B.** External, irreversible, third-party-naming publication is high-stakes and requires review. A, C and D are low-stakes internal or reversible tasks.
  </AccordionItem>

  <AccordionItem title="Q9 · Two colleagues verify a high-stakes figure by asking two different AI tools; both agree, so they accept it. What is the FLAW? (Select one)">
    A. None; agreement confirms accuracy.
    B. Both tools can share the same wrong prior, so agreement is weak; verify against an authoritative source.
    C. They should have used one tool twice.
    D. They should have asked a third AI.

    **Answer: B.** Cross-model agreement is not independent verification. A is false. C and D do not make the check authoritative.
  </AccordionItem>

  <AccordionItem title="Q10 · Which TWO checks require NO external source? (Select two)">
    A. Confirming a table's rows sum to its stated total.
    B. Checking that the summary's recommendation matches the analysis body.
    C. Confirming a cited paper exists.
    D. Verifying a market statistic from model memory.
    E. Confirming a quoted price against a public list.

    **Answer: A and B.** Both are internal-consistency checks against the output itself. C, D and E all require an external or authoritative source.
  </AccordionItem>

  <AccordionItem title="Q11 · A throwaway internal brainstorm list is produced for your own use. What verification is appropriate? (Select one)">
    A. Line-by-line verification with expert sign-off.
    B. A quick sanity read, fixing obvious errors.
    C. Web-search every item.
    D. Route it through legal review.

    **Answer: B.** Low-stakes, reversible, internal output warrants proportional (light) verification. A, C and D over-check a throwaway list.
  </AccordionItem>

  <AccordionItem title="Q12 · Which TWO are weak 'checks' that do not verify accuracy? (Select two)">
    A. Asking 'are you sure?' in the same chat.
    B. Regenerating the answer and seeing the same result.
    C. Opening a cited source to confirm it.
    D. Recomputing a total with a spreadsheet.
    E. Comparing the paraphrase against the supplied document.

    **Answer: A and B.** Same-chat self-review inherits the bias and regeneration tests stability, not truth. C, D and E are genuine verification methods.
  </AccordionItem>

  <AccordionItem title="Q13 · A board pack must be signed off in an hour; it reads fluently but rows don't match the headline total and one citation is unfindable. What should you do FIRST? (Select one)">
    A. Send it; it reads well and time is short.
    B. Recompute the totals from the rows, since that free internal check shows the table can't be trusted as-is.
    C. Ask the chat if the numbers are right.
    D. Lower the temperature and regenerate the pack.

    **Answer: B.** The cheapest, highest-value first move is the internal recompute, needing no external source. A is the fluency trap. C reuses biased reasoning. D changes variance, not grounding.
  </AccordionItem>

  <AccordionItem title="Q14 · A dashboard shows 92% overall accuracy for an AI-assisted tagging task and a manager wants to expand it. Which TWO checks matter MOST first? (Select two)">
    A. Break accuracy down per category to find a failing segment.
    B. Confirm failures are not concentrated in the highest-stakes categories.
    C. Report only the 92% headline.
    D. Assume 92% holds everywhere.
    E. Switch to a cheaper model to save cost.

    **Answer: A and B.** Aggregates hide a badly failing segment, so per-category breakdown and checking high-stakes concentration are essential. C and D mask the risk. E is unrelated to validation.
  </AccordionItem>
</Accordions>

## Key takeaways

- You own the accuracy of anything you act on; verify proportional to stakes.
- Do the free internal-consistency checks first – recompute totals, compare summary to body.
- Use the provenance test: supplied text → passage; citation → open it; memory → independent source; computed → recompute.
- Open citations; a real-sounding name proves nothing, and one fabrication invalidates the sample.
- Irreversible, regulated, external or personal-data outputs require a human review gate.
- Same-chat self-review, two AIs agreeing, regenerating and lowering temperature are not verification.
- Stakes follow downstream use, not the 'internal' label; over-checking throwaways is also wrong.
