# D5 · Reviewing and Verifying Agent Work

Verifying output you did not watch being produced — evidence trails, artefacts, spot-checking, reproducing a claim, and reviewing a long agent run efficiently.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

Worth **16%** — about **8 of 50 items**. This domain tests the skill that makes delegation safe: verifying work you did not watch being produced. When you prompt, you see every step; when you delegate, you get a result and a story about how it was reached. This domain is about trusting the *evidence*, not the story — reading an evidence trail, checking artefacts, spot-checking against the definition of done, reproducing a claim, and doing all of that efficiently on a long run without re-doing the work yourself.

## What you need to know

You cannot verify an unwatched agent by re-reading its confident summary — that is the story, not the evidence. Instead you verify against the **definition of done** (from D2), you inspect the **artefacts** the agent produced (files, drafts, the actual changes), and you follow the **evidence trail** (what it did, which sources it used, where each claim came from). You **spot-check** representative and high-stakes items rather than everything, and you **reproduce** a critical claim independently to confirm it. On a long run, you review the **plan, the checkpoints and the artefacts** rather than the full transcript. The governing rule: a fluent completion summary is not evidence, and "it said it did X" is not the same as "X is true".

## Learning objectives

By the end of this page you should be able to:

1. **Verify** agent output against the definition of done rather than against its own summary.
2. **Distinguish** the agent's claim about what it did from evidence that it actually did it.
3. **Follow** an evidence trail: sources used, steps taken, provenance of each claim.
4. **Spot-check** representatively and reproduce a high-stakes claim independently.
5. **Review a long run efficiently** by reading the plan, checkpoints and artefacts, not the whole transcript.
6. **Detect** silent partial completion and unsupported claims.

---

## 5.1 The story is not the evidence

An agent ends a run with a summary: "I reviewed all twelve accounts, drafted responses, and everything is ready." That sentence is a *claim*. It may be true, partly true, or confidently wrong. Verification means checking the claim against something independent of the claim.

| The agent says… | Verify by… |
| --- | --- |
| "I covered all twelve accounts" | Counting the twelve artefacts against the list |
| "All figures are from the Q1 report" | Opening the report and matching a sample |
| "I fixed every broken link" | Testing a sample of links, including ones it didn't mention |
| "The summary is complete" | Checking it against the definition of done, item by item |

```text
   agent's summary  ──►  "I did X, Y, Z"   (the STORY — take as a claim)
                              │
                              ▼
   verification     ──►  artefacts + evidence trail + definition of done
                         (the EVIDENCE — this is what you trust)
```

:::tip[Assessment signal]
When a stem says the agent "reported success", "summarised what it did", or the output "looks complete", the correct answer verifies against artefacts or the definition of done — not "trust the summary" and not "ask the agent if it's sure" (a same-session self-check carries the same reasoning).
:::

## 5.2 Verify against the definition of done

The definition of done you wrote in D2 is your verification checklist. This is why the two domains are linked: a task with no definition of done cannot be verified, because there is no acceptance test to check against. With one, review becomes mechanical — tick each criterion against the actual output.

| Definition-of-done criterion | Verification check |
| --- | --- |
| "Covers all four competitors" | Count: are all four present? |
| "One sourced price per tier" | Sample: does each price have a working source link? |
| "Flags anything unverifiable" | Look for the flags; their absence is itself suspicious |
| "Fits one slide, ≤ 6 bullets" | Measure it |

If you find yourself unable to verify a result, the gap is usually upstream: the definition of done was too vague. Note it for the next iteration (D6).

## 5.3 Artefacts over assertions

An **artefact** is the concrete thing the agent produced or changed — the file, the draft, the actual edited document, the list of records touched. Artefacts are verifiable; assertions are not. Prefer to review the artefact directly rather than the agent's description of it.

| Assertion (weaker) | Artefact (stronger) |
| --- | --- |
| "I updated the pricing table" | The updated table itself |
| "I emailed the five clients" | The five drafts / the sent-items entries |
| "I removed the outdated sections" | A diff showing what was removed |
| "I found three issues" | The three issues, each with its location |

The habit to build: ask for the work product, not the report of the work product. When an agent can show the artefact, you can check it; when it can only describe it, you are back to trusting the story.

## 5.4 The evidence trail

A good agent run leaves a trail: the plan it followed, the tools and sources it used, and — ideally — the provenance of each claim (which source each fact came from). Following the trail is how you verify without redoing the work.

```text
   PLAN            what it intended to do
     │
   STEPS/TOOLS     what it actually did, in order
     │
   SOURCES         which documents/systems it read
     │
   CLAIMS          each output claim, linked to its source
     │
   ARTEFACTS       the concrete outputs produced
```

The provenance test from foundational verification applies directly: for any claim that matters, ask *where did this come from?* A claim traceable to a supplied document is cheap to confirm; a claim with no source in the trail is a lead to verify, not a fact to trust.

## 5.5 Spot-checking and reproducing a claim

You rarely have time to verify everything, and you don't need to. **Spot-checking** means checking a representative and risk-weighted sample; **reproducing** means independently re-deriving one high-stakes claim to confirm it.

| Technique | When to use | How |
| --- | --- | --- |
| **Representative spot-check** | Bulk output, uniform items | Check a random sample; if any fail, widen the check |
| **Risk-weighted spot-check** | Mixed stakes | Check the highest-stakes items in full |
| **Reproduce a claim** | A critical number or conclusion | Re-derive it yourself (recompute, re-open the source) |
| **Adversarial check** | Suspicious fluency | Check the thing it *didn't* mention, or the edge case |

The key spot-check discipline: **a failed sample invalidates the sample.** If you check five of fifty items and one is wrong, you cannot assume the other forty-five are fine — the failure means you widen the check or reject the batch, exactly as with a fabricated citation invalidating a whole output.

## 5.6 Reviewing a long run efficiently

A twenty-minute run can produce a transcript longer than the work would have taken you. Reading all of it defeats the purpose of delegating. Efficient review is structured, not linear.

```text
   Efficient long-run review (in order, stop early if a step fails):
   1. Read the PLAN         ── did it understand the objective?
   2. Check the CHECKPOINTS ── what did it pause on / decide?
   3. Verify against DONE   ── tick each acceptance criterion
   4. Inspect the ARTEFACTS ── the actual outputs, not the narration
   5. Spot-check + reproduce── a sample + one high-stakes claim
   (only read the full transcript if something above fails)
```

The order matters: a misread objective at step 1 means you can stop and re-brief without wading through the rest. Reading the transcript top-to-bottom is the slow, last-resort move — reserve it for when a structured check surfaces a problem you need to trace.

:::tip[Assessment signal]
When a stem describes a long run and asks how to review it "efficiently", the answer is the structured path — plan, checkpoints, definition of done, artefacts, spot-check — not "read the entire transcript" and not "trust the final summary".
:::

## Decision framework

Use the **TRACE** review method on any delegated result: **T**est against done, **R**eproduce a key claim, **A**rtefacts not assertions, **C**heck a representative sample, **E**vidence trail for provenance.

| Step | Question | What you do |
| --- | --- | --- |
| **T — Test against done** | Does it meet every acceptance criterion? | Tick the definition of done item by item |
| **R — Reproduce** | Is the most critical claim actually true? | Independently re-derive one high-stakes claim |
| **A — Artefacts** | Am I checking the work or the report of it? | Inspect the concrete output/diff, not the summary |
| **C — Check a sample** | Do representative items hold up? | Spot-check risk-weighted; a failure widens the check |
| **E — Evidence trail** | Where did each claim come from? | Follow sources/provenance; unsourced claims are leads |

Apply it tomorrow: on your next delegated result, run TRACE before you act on it. If any step can't be completed, the gap is often an upstream brief defect to fix in D6.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Trusting the completion summary | It's fluent and confident | Verify against artefacts and the definition of done |
| "Ask the agent if it's sure" | Feels like a check | Same-session self-review carries the same reasoning; check independently |
| Verifying nothing because it "looks complete" | Polish reads as correctness | Tick the acceptance criteria; look for what's missing |
| Checking assertions, not artefacts | The report is easier to read | Inspect the actual output/diff/records |
| Assuming a passed sample means all passed | Small check felt sufficient | A failed sample invalidates it; a passed one only covers what you checked |
| Reading the whole transcript to be thorough | Feels rigorous | Use the structured path; transcript is last resort |
| Not reproducing a high-stakes number | The agent showed working | Re-derive critical claims independently |
| No way to verify at all | The definition of done was vague | Note it and sharpen the brief next iteration (D6) |

## Scenario challenge

**Scenario.** Lena delegated an agent to audit her team's 60-page knowledge base for outdated links and stale product references. Forty minutes later the agent returns a tidy report: "Reviewed all 60 pages. Found and fixed 14 broken links and updated 9 outdated product references. The knowledge base is now current." The summary is clear and confident, and Lena is tempted to close the task. But she has an hour before a stakeholder relies on the knowledge base, and the stakes are real — customers use these pages.

**Expert reasoning trace.**

1. **Treat the summary as a claim.** "Reviewed all 60 pages… now current" is the story. It might be true, but Lena has no evidence yet — she didn't watch the run. Closing the task on the summary is the fluency trap.
2. **Verify against the definition of done first.** Her brief asked for outdated links *and* stale references fixed across all pages. She checks the artefacts: does a diff or change list show edits across the pages, or only a subset? If the agent touched 40 of 60 pages, "reviewed all 60" is already suspect — a possible silent partial completion.
3. **Artefacts over assertions.** Rather than trusting "fixed 14 broken links", she looks at the actual changes — the list of links it reports fixing — and tests a sample of them. She also tests a few links it *didn't* mention (the adversarial check), because a link it missed is exactly what the summary won't surface.
4. **Spot-check, and let a failure widen the check.** She samples five of the 14 "fixed" links; four resolve, one still 404s. That failed sample means she can't trust the remaining nine by assumption — she widens to check all 14, and now distrusts the "9 references updated" claim too.
5. **Reproduce a high-stakes claim.** The most consequential reference — the current pricing page link customers hit — she opens herself to confirm it points to the live page, rather than trusting the report.
6. **Follow the evidence trail for the "all 60 pages" claim.** She checks the run's step log: which pages did it actually open? If the trail shows only 44 pages read, the "all 60" claim is false and there's an unreviewed remainder — a silent partial completion she'd never have caught from the summary.
7. **Decide and iterate.** She corrects the missed links, re-scopes the unreviewed pages, and notes for D6 that the brief needed a definition of done requiring a per-page checklist and a flag for any page it couldn't process — so the next run can't quietly stop short.

**The point.** The confident summary was the least reliable part of the deliverable. Verifying against the definition of done, inspecting artefacts, spot-checking (and widening on a failure), reproducing the highest-stakes claim, and following the evidence trail caught a silent partial completion that "trust the summary" or "ask the agent if it's sure" would have missed — and did it in far less time than re-reading 60 pages.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "The summary says it's done, so it's done" | Fluent summaries read as truth | The summary is a claim; verify against evidence |
| "Ask the agent to confirm it did the work" | Feels like a verification step | Same-session self-check reuses the same reasoning; check independently |
| "It looks complete, no need to check" | Polish signals competence | Silent partial completion looks complete; verify coverage |
| "Read the whole transcript to be sure" | Thoroughness feels safe | Structured review (plan→done→artefacts→sample) is faster and catches more |
| "Five samples passed, so all passed" | Small check felt enough | A passed sample only covers itself; a failed one invalidates the batch |
| "It showed its working, so the number's right" | Working looks like proof | Reproduce high-stakes claims independently |

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · An agent ends a run with 'I completed all tasks successfully.' What is the BEST next step? (Select one)">
    A. Close the task; the summary confirms success.
    B. Verify the result against the definition of done and inspect the artefacts it produced.
    C. Ask the agent in the same session whether it is sure.
    D. Re-run the whole task to compare.

    **Answer: B.** The summary is a claim; verification means checking artefacts and the acceptance criteria independently. Trusting the summary (A) is the fluency trap. A same-session self-check (C) reuses the agent's own reasoning. Re-running everything (D) wastes the delegation and doesn't verify the original.
  </AccordionItem>

  <AccordionItem title="Q2 · Why is 'ask the agent if it completed the work correctly' a weak verification method? (Select one)">
    A. It costs too many tokens.
    B. A same-session self-check relies on the same reasoning that produced the result, so it can confirm its own error.
    C. Agents can't answer questions about their work.
    D. It's actually the strongest method.

    **Answer: B.** Asking the agent to self-assess in the same session carries the original reasoning and bias, so it may confirm a mistake. Token cost (A) isn't the issue. Agents can describe their work (C) — that's exactly the unreliable claim. It is not the strongest method (D).
  </AccordionItem>

  <AccordionItem title="Q3 · What is an 'artefact' in the context of verifying agent work? (Select one)">
    A. The agent's summary of what it did.
    B. The concrete output or change it produced — the file, draft, diff or records.
    C. The model's temperature setting.
    D. The run's start time.

    **Answer: B.** An artefact is the actual work product you can inspect and check, unlike the agent's description of it. The summary (A) is an assertion, not an artefact. Temperature (C) and start time (D) are run metadata, not the work product.
  </AccordionItem>

  <AccordionItem title="Q4 · You spot-check 5 of 50 bulk items and one fails. What should you conclude? (Select one)">
    A. The other 45 are fine; one failure is acceptable.
    B. The failed sample invalidates the assumption of uniform quality; widen the check or reject the batch.
    C. Re-run only the failed item and ship the rest.
    D. Nothing; spot-checks are unreliable.

    **Answer: B.** A failure in the sample means you can no longer assume the unchecked items are correct, so you widen the check or reject. Assuming the rest are fine (A) ignores the signal. Fixing only the one (C) leaves 44 unverified. Spot-checks are useful, not unreliable (D).
  </AccordionItem>

  <AccordionItem title="Q5 · What is the MOST efficient way to review a long, mostly-unwatched agent run? (Select one)">
    A. Read the entire transcript top to bottom.
    B. Read the plan and checkpoints, verify against the definition of done, inspect artefacts, then spot-check — reading the full transcript only if something fails.
    C. Trust the final summary.
    D. Ask the agent to shorten its transcript.

    **Answer: B.** Structured review — plan, checkpoints, done-criteria, artefacts, spot-check — is faster and catches more than a linear read. The full transcript (A) is the slow last resort. Trusting the summary (C) verifies nothing. Shortening the transcript (D) doesn't verify the work.
  </AccordionItem>

  <AccordionItem title="Q6 · An agent claims 'reviewed all 60 pages' but the evidence trail shows it read only 44. What has occurred? (Select one)">
    A. Nothing unusual.
    B. A silent partial completion — it stopped short but reported full coverage; the unreviewed pages must be handled.
    C. A hallucinated model.
    D. A temperature error.

    **Answer: B.** Reporting full coverage while the trail shows partial work is a silent partial completion, caught only by checking the evidence trail against the claim. It is not normal (A). It isn't a 'hallucinated model' (C) or a temperature issue (D) — it's a coverage-vs-claim mismatch.
  </AccordionItem>

  <AccordionItem title="Q7 · For a single high-stakes number in an agent's report, what is the strongest check? (Select one)">
    A. Trust it because the agent showed its working.
    B. Reproduce it independently — recompute or re-open the source yourself.
    C. Ask the agent to recompute it in the same session.
    D. Round it to hide any error.

    **Answer: B.** Independently reproducing the critical claim confirms it without relying on the agent's own reasoning. Shown working (A) can still be wrong. A same-session recompute (C) reuses the original reasoning. Rounding (D) conceals rather than checks.
  </AccordionItem>

  <AccordionItem title="Q8 · Which TWO checks best verify an agent's claim that it 'fixed all broken links'? (Select two)">
    A. Test a sample of the links it says it fixed.
    B. Test some links it did not mention, in case it missed them.
    C. Read its summary again more carefully.
    D. Ask it to confirm in the same chat.
    E. Assume completion because the report is detailed.

    **Answer: A and B.** Testing a sample of the claimed fixes (A) checks the work it reported, and testing links it didn't mention (B) is the adversarial check that catches what the summary hides. Re-reading the summary (C) and asking in-session (D) verify nothing independent. Detail (E) isn't evidence of completion.
  </AccordionItem>

  <AccordionItem title="Q9 · Why is the definition of done central to verifying agent work? (Select one)">
    A. It sets the model temperature.
    B. It is the acceptance checklist you tick against the actual output; without it there's nothing objective to verify against.
    C. It makes the agent run faster.
    D. It replaces the need to inspect artefacts.

    **Answer: B.** The definition of done is the objective checklist that turns review into ticking criteria; a vague or missing one leaves nothing to verify against. It doesn't set temperature (A) or speed (C), and it complements — not replaces — artefact inspection (D).
  </AccordionItem>

  <AccordionItem title="Q10 · An agent's report reads well but you find you cannot verify one of its claims at all. What does this most likely indicate? (Select one)">
    A. The claim is definitely correct.
    B. The brief's definition of done was too vague to check — note it and sharpen the brief for the next iteration.
    C. The model is broken.
    D. Verification is impossible for agents.

    **Answer: B.** An unverifiable claim usually reflects an upstream brief defect — no checkable criterion — which you fix in the next iteration (D6). Inability to verify doesn't make the claim correct (A). It's not a broken model (C), and agent work is verifiable when the brief supports it (D).
  </AccordionItem>

  <AccordionItem title="Q11 · Which is the difference between an assertion and an artefact you should prefer to review? (Select one)">
    A. Prefer the assertion — the summary is easier to read.
    B. Prefer the artefact — the actual updated table or diff, over the claim 'I updated the table'.
    C. They are identical.
    D. Prefer whichever is shorter.

    **Answer: B.** The artefact is the checkable work product; the assertion is only a claim about it, so you review the artefact. Preferring the summary (A) trusts the story. They are not identical (C), and length (D) is irrelevant to which is verifiable.
  </AccordionItem>

  <AccordionItem title="Q12 · A stakeholder will rely on an agent-produced deliverable in an hour. Which TWO actions give the most verification value in limited time? (Select two)">
    A. Verify against the definition of done and reproduce the single highest-stakes claim.
    B. Risk-weighted spot-check of the most consequential items.
    C. Re-read the agent's summary twice.
    D. Re-run the entire task from scratch.
    E. Ask the agent to grade its own work.

    **Answer: A and B.** Ticking the acceptance criteria plus reproducing the top-stakes claim (A) and risk-weighting the spot-check to the most consequential items (B) concentrate limited time where errors matter most. Re-reading the summary (C) and self-grading (E) verify nothing independent, and re-running everything (D) wastes the hour.
  </AccordionItem>
</Accordions>

## Key takeaways

- The agent's completion **summary is a claim**, not evidence; verify against something independent of it.
- Verify against the **definition of done** — it is your acceptance checklist, and a vague one leaves nothing to check.
- Prefer **artefacts over assertions**: inspect the actual output, diff or records, not the report of them.
- Follow the **evidence trail** — plan, steps, sources, provenance — to verify without redoing the work.
- **Spot-check** representatively and by risk; a failed sample invalidates the batch, a passed one only covers what you checked.
- **Reproduce** the single highest-stakes claim independently; shown working is not proof.
- Review a long run **structurally** (plan → checkpoints → done → artefacts → sample); read the full transcript only as a last resort.
- Watch for **silent partial completion** — a claim of full coverage the evidence trail contradicts.
