# D2 · Output Evaluation and Validation

Evaluating accuracy and completeness, detecting hallucinations, inconsistency and bias, fact-checking, deciding when human review is required, and adapting outputs for audience and format.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This is the heaviest domain on the Associate exam – roughly **13 of 60 items**. It tests whether you treat Claude's output as a draft to be verified rather than a fact to be forwarded. Nearly every item hinges on one question: *what could be wrong here, and how would I know?*

## Learning objectives

By the end of this page you should be able to:

1. Evaluate an output for **accuracy, completeness, relevance and consistency**.
2. Recognise the signatures of **hallucination**, **internal inconsistency** and **bias**.
3. Choose an appropriate **fact-checking** strategy proportional to the stakes.
4. Decide **when human review is mandatory** versus optional.
5. Adapt an output for a specific **audience** and choose the right **output format** (inline text, artifact, structured data).

---

## 2.1 The four evaluation lenses

Every output can be audited on four axes. Exam items usually target one axis and use the others as distractors.

| Lens | Question to ask | Typical failure | Detection tactic |
| --- | --- | --- | --- |
| **Accuracy** | Are the facts, figures, names, dates and quotes true? | Hallucinated statistic; misattributed quote; invented URL | Spot-check anything checkable; ask for sources; cross-reference |
| **Completeness** | Does it cover everything the task required? | Answers 3 of 4 asked questions; silently drops an edge case | Re-read the original brief; use a checklist |
| **Relevance** | Does it answer *this* question for *this* audience? | Generic essay when a decision memo was needed | Compare against the stated purpose |
| **Consistency** | Do parts of the output agree with each other and with inputs? | Table totals do not match; conclusion contradicts body | Recompute; compare sections; compare to source doc |

:::tip[Exam signal]
When a stem says an output "looks polished", "reads confidently" or "was produced quickly", the item is testing whether you will still verify it. Fluency is not evidence of accuracy.
:::

## 2.2 Hallucination – what it is and how it looks

A **hallucination** is content that is fluent and plausible but not grounded in the input or in reality. LLMs generate the most probable continuation, not the most verified one. Hallucinations cluster in predictable places:

- **Specific numerics** – percentages, dates, prices, counts, especially for niche topics.
- **Citations** – paper titles, authors, URLs, case law, ISBNs. Often a real-sounding blend of real elements.
- **Proper nouns** – product names, people, internal project names the model has never seen.
- **Long-tail facts** – anything rare in training data.
- **Filling gaps** – when the source document does not contain the answer, the model may still produce one.

### The provenance test

For any claim that matters, ask: **where did this come from?**

| Source of the claim | Trust posture |
| --- | --- |
| Quoted or paraphrased from a document *you supplied* | Verify the paraphrase against the document; usually reliable |
| Retrieved with a citation in research mode | Open the citation; confirm it says what Claude says it says |
| From model knowledge with no source | Treat as a lead, not a fact; verify independently before use |
| A number the model computed | Recompute yourself or ask Claude to show working, then check |

### Practical detection heuristics

1. **Ask for sources, then check them.** A fabricated citation is a red flag for the surrounding paragraph too.
2. **Ask the same question two ways.** Divergent answers indicate low grounding.
3. **Ask "what in the provided material supports this?"** A grounded model points to text; an ungrounded one restates the claim.
4. **Look for excessive specificity on obscure topics.** "Adoption rose 37.4% in Q3 2019" about a small company is suspicious.
5. **Check dates against the model's knowledge cutoff.** Claims about very recent events without web research are suspect.

:::caution[Do not use Claude to verify Claude in the same session]
Asking the same conversation "are you sure?" retains the reasoning that produced the error. Use an **independent** check: a fresh session, a different source, or a human. This is the same principle the Architect exam calls "same-session self-review bias".
:::

## 2.3 Internal inconsistency

Inconsistency is cheaper to detect than hallucination because you do not need external sources – you only need to compare the output with itself and with its inputs.

Common patterns:

- A summary table whose totals do not equal the sum of rows.
- An executive summary recommending option A while the analysis body favours option B.
- A translated document where a defined term is rendered three different ways.
- A timeline where an effect precedes its cause.
- A "three risks" list with four items or two.

**Tactic:** ask Claude to produce a **self-check table** (claims vs. evidence location) *and then verify that table yourself*. The table makes inconsistencies visible; your verification closes the loop.

## 2.4 Bias

Bias in outputs shows up as **skewed framing, unrepresentative examples, stereotyped assumptions or asymmetric treatment** of groups, vendors or options.

| Bias type | Example | Mitigation |
| --- | --- | --- |
| Framing bias | Pros listed for one vendor, cons for another | Ask for a symmetric pros/cons table with equal depth |
| Demographic bias | Personas default to one gender or culture | Specify diverse, representative attributes; review before publishing |
| Confirmation bias (yours) | You prompted "explain why plan A is best" | Prompt neutrally: "compare A and B against criteria X, Y, Z" |
| Recency/availability bias | Over-weights whatever was in the pasted document | Supply balanced sources; ask what is missing |
| Sycophancy | Model agrees with your stated view | Ask for the strongest counter-argument; use a fresh session |

The exam expects you to know that **bias mitigation is a shared responsibility**: prompt design reduces it, human review catches what remains, and organisational review processes exist for external-facing content.

## 2.5 Fact-checking proportional to stakes

Not every output needs the same rigour. Match effort to consequences.

```text
Stakes ────────────────────────────────────────────────────►
Low                     Medium                        High
brainstorm, first       internal memo, slide          external publication,
draft, internal chat    deck, customer email draft    legal/medical/financial,
                                                      regulatory filing, pricing
Spot-check obvious      Verify all numbers and        Line-by-line verification,
errors; sanity read     names; open 1–2 sources       independent SME review,
                                                      documented sign-off
```

Signals in a stem that push toward **high rigour**: "will be published", "regulator", "client-facing", "contract", "compliance", "medical", "financial figures", "press release", "irreversible".

## 2.6 When human review is required

Human review is **mandatory** (not merely advisable) when any of these hold:

1. **Irreversibility** – the action cannot be undone (sending, publishing, paying, deleting).
2. **Regulated domain** – legal, medical, financial advice, HR decisions, safety-critical instructions.
3. **External audience** – customers, press, regulators, the public.
4. **Personal data or protected characteristics** – anything touching PII or fairness.
5. **Organisational policy says so** – many AI use policies name review gates explicitly.
6. **Novel or high-uncertainty content** – the model is operating outside well-trodden knowledge.

Human review is **optional** for: internal brainstorming, first drafts you will rewrite, formatting or restructuring of your own content, exploratory analysis you will re-run.

:::note[The exam's phrasing]
Look for "human-in-the-loop", "review gate", "sign-off" or "subject-matter expert" in correct options. Options that skip review for irreversible or regulated actions are wrong regardless of how confident the output is.
:::

## 2.7 Adapting output for an audience

The same analysis needs different shapes for different readers. Evaluation includes checking **fit**.

| Audience | What they need | What to strip | Prompt lever |
| --- | --- | --- | --- |
| Executives | Decision, options, risk, ask; one page | Method detail, hedging | "Write a one-page decision memo with a recommendation up front" |
| Engineers | Precise specification, edge cases, constraints | Marketing tone | "Use precise terminology; list assumptions and edge cases" |
| Customers | Plain language, empathy, next steps | Internal jargon, blame | "Plain language, 8th-grade reading level, warm but direct" |
| Regulators / legal | Exact terms, citations, no overclaiming | Superlatives | "Neutral tone; cite the clause; avoid absolute claims" |
| New hires | Context, definitions, why | Assumed knowledge | "Define every acronym on first use" |

Evaluation questions to ask: Is the reading level right? Is the tone right? Is the length right? Is the **call to action** clear? Does it overclaim?

## 2.8 Choosing the output format

<Tabs>
  <TabItem label="Inline text">
    **Use when:** the answer is short, conversational, or will be copied into another tool. Fast to iterate; no persistence.

    **Avoid when:** you will revise the deliverable several times, or it has structure (code, long document, table) that benefits from a dedicated view.
  </TabItem>
  <TabItem label="Artifact">
    **Use when:** the output is a standalone deliverable you will iterate on – a document, a slide outline, a chart, a mock-up, a script, a spreadsheet-like table. Artifacts persist alongside the conversation, can be versioned by iteration, and are easy to copy or download.

    **Avoid when:** you just need a quick fact or a one-line rewrite.
  </TabItem>
  <TabItem label="Structured data">
    **Use when:** the result will be consumed by another system or pasted into a spreadsheet – lists of records, JSON, CSV, Markdown tables with fixed columns. Ask for the exact schema and column order you need.

    **Avoid when:** nuance matters more than machine-readability; forcing structure can drop qualifiers.
  </TabItem>
</Tabs>

**Rule of thumb for the exam:** iterative deliverable → artifact; machine-consumed → structured data; conversational → inline.

## 2.9 A repeatable evaluation checklist

Use this on anything you will act on:

1. **Re-read the brief.** Did the output address every part? (Completeness)
2. **Circle every number, name, date, quote and citation.** Verify each against a source proportional to stakes. (Accuracy)
3. **Check internal agreement** – totals, summary vs. body, defined terms. (Consistency)
4. **Check fit** – audience, tone, length, format, call to action. (Relevance)
5. **Check fairness** – symmetric treatment, representative examples. (Bias)
6. **Decide the review gate** – is this irreversible, regulated, external or personal-data-bearing? If yes, route to human review before use.
7. **Record what you verified** if the output feeds a decision others will rely on.

---

## 2.10 A verification decision tree

Faced with any output, an expert routes it through a small number of questions. The tree below turns the four lenses into a repeatable triage.

```text
Will I ACT on this output (send, publish, decide, pay)?
│
├─ No  ─► Low stakes. Sanity-read; fix obvious errors; proceed.
│
└─ Yes ─► Is the action irreversible, regulated, external, or
          touching personal data / protected characteristics?
          │
          ├─ Yes ─► HIGH stakes: verify every claim; route to a
          │         human/SME review gate; document sign-off.
          │
          └─ No  ─► MEDIUM stakes: verify all numbers, names,
                    dates, citations; check internal consistency;
                    a second reader for tone.
```

Then, for each *checkable claim*, apply the provenance test:

| Where did the claim come from? | Cheapest sufficient check |
| --- | --- |
| A document you supplied | Locate the supporting sentence; confirm the paraphrase |
| A research-mode citation | Open the source; confirm it says what Claude claims |
| The model's own knowledge | Treat as a lead; verify against an independent source |
| A number the model computed | Recompute independently (calculator/spreadsheet/analysis tool) |
| Internal (totals, cross-references) | Recompute or compare sections – no external source needed |

:::tip[Exam signal]
The correct answer picks the *cheapest check that is sufficient for the stakes*. Internal-consistency failures (totals, contradictions) never require external sources – options that reach for the web to check arithmetic are over-engineered distractors.
:::

## 2.11 Completeness and relevance in depth

Accuracy gets the attention, but **completeness** and **relevance** failures are just as common and easier to miss because the text reads well.

| Failure | How it hides | Detection |
| --- | --- | --- |
| Answered 3 of 4 sub-questions | The three answered ones are thorough | Re-read the brief as a checklist; tick each part |
| Dropped an edge case or exclusion | The main case is handled confidently | Ask "what did the task list that is not here?" |
| Right facts, wrong audience | Fluent but pitched at the wrong reader | Compare tone/level against the stated audience |
| Answered a related but different question | Plausible and on-topic | Restate the exact question and compare |
| Buried the recommendation | Everything is present but unusable | Check the ask: decision-first for executives? |

**Before/after – an executive brief that was accurate but irrelevant:**

```text
Before:  A 3-page, method-heavy analysis with the recommendation on page 3.
         Every fact is correct, but the CFO cannot find the decision.
Check:   Relevance + audience fit fail – accuracy passed but usefulness did not.
After:   One page, recommendation and the ask in the first two lines, three
         supporting bullets, method in a short appendix. Same facts, usable.
```

:::tip[Exam signal]
An output can be fully accurate and still wrong for the task if it is incomplete or mis-pitched. Watch for options that only check facts when the stem's problem is a missing sub-answer or the wrong audience.
:::

## 2.12 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "Confident, fluent, well-formatted output is accurate." | Fluency is a property of language models, not a truth signal. | The core D2 trap: forwarding polished-but-unverified output. |
| "If I gave Claude the document, its summary must be right." | Extraction can still misread or fabricate; verify against the text. | Provenance test on supplied documents. |
| "Asking 'are you sure?' in the same chat verifies it." | Same-session self-review carries the original bias. | Self-report / same-session distractor. |
| "Two models agreeing means the answer is correct." | Both can share the same wrong prior. | Weak-verification distractor. |
| "A real-looking journal name proves the citation is valid." | Fabricated citations blend real elements; open the source. | Citation-verification item. |
| "Lower temperature removes hallucination." | It reduces variance, not ungroundedness. | Setting-as-fix distractor. |
| "Internal-only content need not be fact-checked." | Internal figures feed external decisions. | Stakes-follow-downstream-use item. |
| "High overall accuracy means the specific claim is fine." | Aggregates hide the specific error that matters. | Aggregate-metric distractor. |

## 2.13 Scenario walkthrough – a board pack under time pressure

**Scenario.** Marcus has 90 minutes before a board pack is due. Claude produced a polished market analysis: a summary table of segment sizes, five cited sources, and a recommendation to enter Segment C. The prose is fluent and the totals "look right". Two of the five citations are to reputable-sounding reports. A colleague says "it reads great, just send it". Marcus notices the segment table's rows sum to a different total than the headline figure, and one citation is a report he cannot find anywhere.

**Expert reasoning trace.**

1. **Set the stakes.** This is external-facing, decision-driving, and hard to walk back once the board acts. That forces HIGH-rigour verification and a review gate – "it reads great, send it" (the fluency trap) is rejected outright.
2. **Attack internal consistency first – it is free.** The rows not summing to the headline is a consistency failure checkable without any source: recompute. This alone means the table cannot be trusted as-is.
3. **Apply the provenance test to citations.** The unfindable report is a probable fabrication; a fabricated citation invalidates sample-based trust, so *every* citation and the claims resting on them must be checked, not just the bad paragraph (rejecting "remove only the bad paragraph").
4. **Reject same-session self-check.** Asking the chat "are you sure about these numbers?" reuses the reasoning that produced them; verification must be independent (recompute, open sources).
5. **Reject "lower temperature and regenerate".** That changes variance, not grounding, and wastes the 90 minutes.
6. **Route the gate.** Because it is board-facing, the corrected pack still needs a human/SME sign-off before release.

**Exam-correct decision:** recompute the totals independently, verify every citation (treating the whole output as unverified once one fails), correct or remove unsupported claims, and pass a human review gate before sending. **Not** trust the fluency, **not** a same-session self-check, **not** a temperature change, **not** a partial fix of only the flagged paragraph.

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| "The output was detailed and well-formatted, so it can be sent" | Fluency ≠ accuracy |
| "Ask Claude in the same chat if it is sure" | Same-session self-review carries the same bias |
| "The citation had a real journal name, so it is valid" | Fabricated citations blend real elements; open the source |
| "Overall accuracy looked high" | Aggregate impressions hide specific errors; verify the specific claims that matter |
| "It is only an internal memo, no need to check the figures" | Internal figures feed external decisions; verify numbers proportional to downstream use |
| "Publish, then correct if someone complains" | Publication is irreversible; review gate is required |
| "Use temperature 0 / ask for 'only facts' to eliminate hallucination" | Reduces variance, does not guarantee grounding |
| "Two AI tools agree, so the answer is verified" | Both can share the same wrong prior; use an independent, authoritative check |
| "Remove only the paragraph with the bad citation" | One fabricated citation invalidates the sample; verify the whole output |
| "The facts are all correct, so the output is fine" | It can still be incomplete or mis-pitched for the audience; check relevance and completeness |
| "Recompute the table by searching the web" | Internal-consistency errors are checked by recomputation, not external sources |

---

## Practice questions

Each item states how many responses to select. Attempt before revealing.

<Accordions>
  <AccordionItem title="Q1 · A marketing analyst asks Claude for market-size statistics for a niche B2B segment. Claude returns five precise percentages with year-over-year growth figures but no sources. What is the BEST next step? (Select one)">
    A. Use the figures; they are consistent with each other.
    B. Ask Claude, in the same conversation, whether it is confident in the figures.
    C. Treat the figures as leads, ask for sources, and verify each against an independent source before use.
    D. Lower the model's temperature and regenerate to get more accurate figures.

    **Answer: C.** Precise, unsourced numerics on a niche topic are the classic hallucination signature. Internal consistency (A) is not evidence of truth. Same-session self-check (B) shares the original reasoning bias. Temperature (D) affects variance, not grounding.
  </AccordionItem>

  <AccordionItem title="Q2 · An operations manager uses Claude to summarise a 40-page supplier contract. The summary states the termination notice period is 60 days. Which action MOST appropriately validates this claim? (Select one)">
    A. Accept it; the contract was supplied in the prompt so Claude had the text.
    B. Ask Claude to quote the exact clause and page, then open the contract and confirm.
    C. Ask a second AI tool the same question and accept the answer if it matches.
    D. Ask Claude to summarise the contract again and compare.

    **Answer: B.** Provenance test: the claim comes from a supplied document, so the correct check is to locate the supporting text and confirm it. Supplying the document (A) does not guarantee correct extraction. Two models agreeing (C) is weak evidence. Regenerating (D) checks stability, not accuracy.
  </AccordionItem>

  <AccordionItem title="Q3 · Claude drafts a customer-facing email explaining a service outage and includes a sentence attributing the cause to a named third-party vendor. Which TWO actions are required before sending? (Select two)">
    A. Verify the vendor attribution with the incident owner.
    B. Route the draft through the organisation's external-communications review.
    C. Ask Claude to make the tone more apologetic.
    D. Send immediately to minimise customer confusion.
    E. Regenerate the email at a lower temperature.

    **Answer: A and B.** External, irreversible communication that names a third party is high-stakes: verify the factual claim and pass the organisational review gate. Tone (C) is an optional improvement. Sending immediately (D) skips mandatory review. Temperature (E) is irrelevant to the risk.
  </AccordionItem>

  <AccordionItem title="Q4 · A project manager asks Claude to compare two software vendors. The output lists six strengths for Vendor A and six weaknesses for Vendor B. What is the MOST likely issue and the best remedy? (Select one)">
    A. Hallucination – ask for sources for each point.
    B. Framing bias – ask for a symmetric comparison with equal-depth pros and cons for both vendors against the same criteria.
    C. Incompleteness – ask for more strengths for Vendor A.
    D. Inconsistency – ask Claude to recompute the totals.

    **Answer: B.** Asymmetric treatment of options is framing bias, often triggered by how the question was asked. The remedy is a neutral, criteria-based, symmetric structure. Sources (A) may also help but do not fix the asymmetry.
  </AccordionItem>

  <AccordionItem title="Q5 · A team lead wants a one-page brief for the executive committee based on a 15-page analysis Claude produced. Which output choice and check are MOST appropriate? (Select one)">
    A. Paste the 15 pages inline and let executives skim.
    B. Ask for a one-page artifact with the recommendation first, then verify that the recommendation matches the analysis body and that every figure traces to the source.
    C. Ask for a JSON summary so it can be loaded into a dashboard.
    D. Ask Claude to make the analysis more detailed.

    **Answer: B.** Audience fit (executives need the decision up front, one page), format fit (iterative deliverable → artifact), and a consistency check between summary and body plus accuracy check on figures.
  </AccordionItem>

  <AccordionItem title="Q6 · An HR coordinator drafts interview feedback summaries with Claude for six candidates. Before the summaries are used in hiring decisions, what is REQUIRED? (Select one)">
    A. Nothing additional; the summaries are internal.
    B. Human review by the hiring panel for accuracy and for language that could reflect bias on protected characteristics.
    C. Ask Claude to rate each candidate numerically to remove subjectivity.
    D. Regenerate each summary twice and keep the longest version.

    **Answer: B.** Hiring decisions are regulated, high-stakes and touch protected characteristics – a mandatory human-review gate. Numerical ratings (C) do not remove bias and may launder it. "Internal" (A) does not exempt regulated decisions.
  </AccordionItem>

  <AccordionItem title="Q7 · Claude produces a research summary in research mode with eight citations. The analyst opens two at random; one supports the claim, one does not exist. What should the analyst do? (Select one)">
    A. Use the summary; 50% of the sample was fine.
    B. Discard only the paragraph with the bad citation.
    C. Treat the entire summary as unverified: check every citation, and independently re-verify any claim whose citation fails.
    D. Ask Claude to replace the missing citation.

    **Answer: C.** A fabricated citation is a red flag for the whole output. Random sampling that finds a failure invalidates the sample-based trust; move to exhaustive verification. Replacing one citation (D) does not address the others.
  </AccordionItem>

  <AccordionItem title="Q8 · A financial analyst asks Claude to compute quarterly totals from a pasted table. Claude returns a total that looks reasonable. What is the MOST appropriate verification? (Select one)">
    A. Trust it; arithmetic is deterministic.
    B. Ask Claude to show its working and recompute the total independently (spreadsheet or calculator).
    C. Ask Claude "are you sure?"
    D. Round the total to reduce error.

    **Answer: B.** LLMs can make arithmetic errors, especially on long tables. Financial figures are medium-to-high stakes; recompute independently. Arithmetic is not reliably deterministic for a language model (A).
  </AccordionItem>

  <AccordionItem title="Q9 · Which output format is MOST appropriate for a list of 200 customer records with fixed fields that will be imported into a CRM? (Select one)">
    A. Inline prose
    B. An artifact containing a narrative summary
    C. Structured data (CSV or JSON) with a specified schema and column order
    D. A slide deck

    **Answer: C.** Machine-consumed, fixed-field data → structured format with the exact schema requested.
  </AccordionItem>

  <AccordionItem title="Q10 · A communications team member asks Claude to translate a 20-page policy into Spanish. On review, the key term for 'data controller' appears in three different Spanish renderings. What lens caught this, and what is the fix? (Select one)">
    A. Accuracy – ask for sources.
    B. Consistency – supply a glossary of defined terms and ask for a terminology check table before final review.
    C. Relevance – shorten the document.
    D. Bias – ask for a neutral tone.

    **Answer: B.** Inconsistent rendering of a defined term is an internal-consistency failure. Supplying a glossary and requesting a terminology table gives a checkable artefact.
  </AccordionItem>

  <AccordionItem title="Q11 · Which of the following are reliable indicators that an output should receive line-by-line human verification? (Select two)">
    A. It will be filed with a regulator.
    B. It was generated in under ten seconds.
    C. It contains dosage or legal-liability language.
    D. It uses bullet points.
    E. It is longer than one page.

    **Answer: A and C.** Regulated filing and medical/legal content are high-stakes triggers. Speed, formatting and length are irrelevant to stakes.
  </AccordionItem>

  <AccordionItem title="Q12 · A user notices Claude tends to agree with whatever position they state in the prompt. What is this behaviour and the BEST mitigation when evaluating a plan? (Select one)">
    A. Hallucination; ask for citations.
    B. Sycophancy; ask neutrally for the strongest arguments on both sides, or request a critique in a fresh session before revealing your preference.
    C. Inconsistency; ask for a table.
    D. Framing; make the prompt longer.

    **Answer: B.** Agreement with the user's stated position is sycophancy. Neutral framing and independent critique counter it.
  </AccordionItem>

  <AccordionItem title="Q13 · An analyst must sign off a board pack in an hour. It reads fluently, but the segment table's rows do not sum to the headline total and one of five citations cannot be found anywhere. What should the analyst do FIRST? (Select one)">
    A. Send it; it reads well and time is short.
    B. Recompute the totals from the rows (a free internal check), since a consistency failure means the table cannot be trusted as-is.
    C. Ask Claude in the same chat whether the numbers are correct.
    D. Lower the temperature and regenerate the whole pack.

    **Answer: B.** The cheapest, highest-value first move is the internal-consistency recompute, which needs no external source. Sending on fluency (A) is the fluency trap; a same-session self-check (C) reuses the biased reasoning; regenerating at lower temperature (D) changes variance, not grounding.
  </AccordionItem>

  <AccordionItem title="Q14 · Claude answers three of four questions in a stakeholder brief thoroughly and omits the fourth, but the brief reads as complete. Which evaluation lens catches this and how? (Select one)">
    A. Accuracy; open external sources.
    B. Completeness; re-read the original brief as a checklist and tick each required part.
    C. Bias; balance the examples.
    D. Relevance; change the audience.

    **Answer: B.** A silently dropped sub-question is a completeness failure, caught by checking the output against the brief as a checklist. Accuracy sourcing (A), bias balancing (C) and audience changes (D) do not detect a missing part.
  </AccordionItem>

  <AccordionItem title="Q15 · A fully accurate 3-page analysis is handed to a CFO who complains it is useless because she cannot find the decision. What TWO fixes address this? (Select two)">
    A. Restructure to put the recommendation and the ask in the first two lines.
    B. Cut method detail to a short appendix and lead with three supporting bullets.
    C. Add more supporting data throughout.
    D. Switch to a more capable model.
    E. Ask Claude if it is sure the facts are right.

    **Answer: A and B.** The problem is relevance/audience fit, not accuracy: an executive needs the decision first and the method stripped back. Adding data (C) worsens it; a bigger model (D) does not change fit; re-checking facts (E) addresses the wrong lens.
  </AccordionItem>

  <AccordionItem title="Q16 · Claude cites a study with a plausible title, real-sounding authors and a journal name, supporting a key claim. What is the MOST appropriate validation before relying on it? (Select one)">
    A. Accept it; the journal name is real.
    B. Open the cited source and confirm it exists and actually says what Claude claims.
    C. Ask Claude to provide the DOI in the same chat and accept it.
    D. Regenerate to see if the same citation appears.

    **Answer: B.** Fabricated citations blend real-sounding elements; the only valid check is to open the source and confirm it supports the claim. A real journal name (A) proves nothing; a same-chat DOI (C) can also be fabricated; regeneration (D) tests stability, not existence.
  </AccordionItem>

  <AccordionItem title="Q17 · Two colleagues verify a high-stakes figure by asking two different AI tools; both give the same number, so they treat it as confirmed. What is the FLAW? (Select one)">
    A. None; agreement confirms accuracy.
    B. Both tools can share the same wrong prior, so agreement is weak evidence; verify against an authoritative independent source.
    C. They should have used the same tool twice.
    D. They should have asked a third AI.

    **Answer: B.** Two models agreeing is not independent verification because they can be wrong the same way; an authoritative source or human is needed. Agreement is not confirmation (A); repeating one tool (C) or adding a third AI (D) does not make the check authoritative.
  </AccordionItem>

  <AccordionItem title="Q18 · A dashboard reports 92% overall accuracy for a Claude-assisted tagging task, and a manager wants to expand it. Which check is MOST important before expanding? (Select two)">
    A. Break accuracy down per category to find a segment that is failing.
    B. Confirm the failing categories are not the highest-stakes ones.
    C. Report only the 92% headline to leadership.
    D. Assume 92% is fine everywhere.
    E. Switch to a cheaper model to save cost.

    **Answer: A and B.** An aggregate can hide a category that fails badly, so per-segment breakdown and checking whether failures concentrate in high-stakes categories are essential. Reporting only the headline (C) and assuming uniformity (D) mask the risk; cost (E) is unrelated to the validation question.
  </AccordionItem>

  <AccordionItem title="Q19 · An internal-only memo contains a revenue figure that will feed next year's hiring plan. A colleague says 'it's internal, no need to check the number'. What is the BEST response? (Select one)">
    A. Agree; internal content does not need verification.
    B. Verify the figure, because its downstream use in a hiring decision makes it consequential regardless of the 'internal' label.
    C. Only check the tone.
    D. Ask Claude if it is confident.

    **Answer: B.** Stakes follow the downstream use, not the 'internal' label; a figure feeding hiring must be verified. 'Internal' does not exempt it (A); tone (C) is the wrong lens; a same-chat confidence check (D) is not verification.
  </AccordionItem>
</Accordions>

## Key takeaways

- Fluency is not accuracy. Verify anything you will act on, proportional to stakes.
- Hallucinations cluster in numbers, citations, proper nouns and long-tail facts.
- Use the provenance test: supplied document → check the text; research citation → open it; model knowledge → verify independently.
- Check consistency against inputs and within the output; it is the cheapest check.
- Bias is mitigated by neutral prompting, symmetric structures and human review.
- Irreversible, regulated, external or personal-data outputs **require** human review.
- Format follows use: artifact for iterative deliverables, structured data for machines, inline for conversation.
- Pick the cheapest sufficient check; internal-consistency errors never need external sources.
- Accurate output can still fail on completeness or relevance – check against the brief and the audience.
- Stakes follow downstream use; two AI tools agreeing is not independent verification.
