# D4 · Evaluation, Testing & Optimization

Evaluation metrics, building eval datasets, harness design, rubric grading and LLM-as-judge, A/B testing and significance, offline vs online evals, regression suites, failure diagnosis, and the cost/latency optimisation playbook.

import { Accordions, AccordionItem, Tabs, TabItem, Steps } from '@prosefly/astro-components';

This domain is worth **16% – roughly 10 of 63 items**. It tests whether you can prove a system works and keep it working: choosing the right metrics, building trustworthy eval datasets, running calibrated LLM-as-judge and A/B tests, catching regressions in CI, diagnosing *which layer* failed, and optimising cost and latency without degrading quality. The dominant judgement: **measure per-segment with independent evaluation**, never aggregate-only or self-report.

## Learning objectives

By the end of this page you should be able to:

1. Choose evaluation **metrics** (accuracy, latency p50/p95, cost per task, safety, security) per use case.
2. Build **eval datasets**: golden sets, synthetic data, edge cases, per-segment coverage.
3. Design a **test harness** and **rubric grading**; run **LLM-as-judge** in a separate session/model and **calibrate** it.
4. Run **pairwise A/B tests** and reason about **statistical significance**.
5. Distinguish **offline vs online** evals and run **regression suites in CI**.
6. **Diagnose** prompt failure vs hallucination vs model mismatch vs retrieval failure.
7. Apply the **cost/latency/token optimisation playbook** and roll out changes via **canary**.

---

## 4.1 Evaluation metrics

Different use cases have different definitions of "good". An architect picks the metrics that map to the success criteria (D1).

| Metric | Definition | When it dominates |
| --- | --- | --- |
| Accuracy / task success | Correct outcome vs ground truth or rubric | Most quality-critical tasks |
| Latency p50 / p95 | Median / tail response time | Interactive/SLA-bound systems (watch **p95**, not just p50) |
| Cost per task | \$ per completed unit of work | High-volume, cost-sensitive systems |
| Safety | Rate of harmful/policy-violating outputs | User-facing, regulated |
| Security | Prompt-injection/exfiltration resistance | Tool-using agents, untrusted inputs |

:::tip[Exam signal]
"Average latency is fine but some users wait 12 s" → you are being tested on **p95/tail latency**, not the mean. "Overall accuracy is 92%" while one segment fails → the answer is **per-segment metrics**, not the aggregate (anti-pattern #10).
:::

---

## 4.2 Building eval datasets

An eval is only as trustworthy as its dataset. Coverage matters more than size.

| Dataset element | Purpose | Sourcing |
| --- | --- | --- |
| Golden set | Curated inputs with verified expected outputs | Hand-labelled by SMEs; the source of truth |
| Synthetic data | Scale coverage cheaply; explore rare inputs | Generate with a model, then human-review a sample |
| Edge cases | Exercise the failure boundary | Mined from production errors, adversarial inputs |
| Per-segment coverage | Ensure every document type / customer tier / language is represented | Stratify the set so no segment is invisible |

:::caution[Aggregate blindness]
A golden set that is 90% one document type will report high accuracy while the 10% type silently fails. **Stratify** the dataset and report metrics per segment.
:::

---

## 4.3 Harness design and rubric grading

A **test harness** runs each eval input through the system and scores the output. Scoring methods, from cheapest to richest:

| Method | Use when | Caveat |
| --- | --- | --- |
| Exact / string / regex match | Deterministic outputs (IDs, classifications, JSON fields) | Brittle for free text |
| Rubric grading | Structured criteria (e.g. correctness, completeness, tone each 0–2) | Needs a clear rubric; use LLM-as-judge or humans to apply it |
| LLM-as-judge | Free-text quality at scale | Must use a **separate model/session** and be **calibrated** |
| Human review | Highest-stakes, ambiguous, or calibration reference | Slow/expensive; use as the gold standard |

A rubric turns fuzzy quality into checkable dimensions:

```text
Criterion        Weight  Scale
Correctness       0.5    0 wrong · 1 partial · 2 fully correct
Grounding         0.3    0 unsupported · 1 partial · 2 fully cited
Tone/format       0.2    0 off · 1 minor · 2 on-spec
Score = Σ (weight × normalised criterion)
```

---

## 4.4 LLM-as-judge and calibration

LLM-as-judge scales evaluation, but only if it is trustworthy.

- Use a **different model or at least a fresh session** than the one under test — same-session self-review carries the original reasoning bias (anti-pattern #9).
- **Calibrate** the judge against a set of **human labels**: measure agreement (e.g. accuracy vs human, correlation). If the judge disagrees with humans, fix the rubric/prompt before trusting it at scale.
- Watch for judge biases: position bias (favours the first option), length bias (favours longer answers), self-preference (favours its own family's style).

```text
 human-labelled sample ──▶ run judge ──▶ compare
        │                                   │
        └────── agreement ≥ threshold? ──── yes ─▶ trust judge at scale
                                            no  ─▶ revise rubric/judge prompt, recalibrate
```

:::caution[Never let the model grade its own homework in-session]
Asking the producing session "is this correct?" retains the bias that produced the answer. Evaluation must be independent.
:::

---

## 4.5 Pairwise A/B testing and statistical significance

To compare prompt A vs prompt B (or model A vs B), **pairwise** comparison on the same inputs is more sensitive than comparing separate averages.

<Steps>

1. Run both variants on the **same** eval inputs (paired design).

2. Have an independent judge/human pick the winner per item (or score each).

3. Compute the win rate and test whether the difference is **statistically significant** — a 52% win on 30 items is noise; the same on 2,000 items may be real. Use a significance test and report the confidence interval.

4. Only promote the winner if the improvement is significant **and** holds per segment (no regression on any segment).

</Steps>

:::tip[Exam signal]
"Prompt B looks a bit better on a handful of examples" → the correct answer invokes **statistical significance / larger sample**, not "ship B because it looks better". Small-sample eyeballing is a trap.
:::

---

## 4.6 Offline vs online evaluation

| | Offline eval | Online eval |
| --- | --- | --- |
| When | Pre-release, in CI | In production, live traffic |
| Data | Golden/synthetic sets | Real user interactions |
| Measures | Correctness vs known answers, regressions | Real outcomes: deflection, thumbs, task completion, cost |
| Risk | May not reflect production distribution | Exposes users to changes (mitigate with canary) |

You need **both**: offline gates releases; online catches distribution shift the golden set missed.

### Regression suites in CI

Every prompt, model or retrieval change must run the **regression suite** before promotion. A regression suite is the accumulated golden set plus every past production failure turned into a test. This is how you prevent a fix for one segment from breaking another.

```text
PR opened ─▶ CI runs regression suite (offline evals, per-segment)
   ├─ any segment regresses ─▶ block merge
   └─ all pass ─▶ allow ─▶ canary online ─▶ ramp
```

---

## 4.7 Diagnosing failures: which layer broke?

When output is wrong, isolate the layer before "fixing" anything. Fixing the wrong layer is the most expensive mistake in this domain.

| Symptom | Likely cause | Confirm by | Fix at |
| --- | --- | --- | --- |
| Answer uses facts not in provided context | Hallucination / weak grounding | Check faithfulness against retrieved chunks | Grounding instruction, citations |
| Right chunks retrieved, wrong answer shape/format | Prompt failure | Inspect prompt vs output; test prompt in isolation | Prompt/template |
| Confident-but-wrong after a data change | Retrieval/indexing (stale) | Log retrieved chunk IDs | Re-index / freshness pipeline (D3) |
| Fails only on hardest cases, fine elsewhere | Model mismatch (too small) | Re-run failures on a stronger model | Route/escalate model tier |
| Fails only on one segment | Coverage / segment-specific issue | Per-segment metrics | Targeted data/prompt/retrieval fix |

```text
Wrong output ─▶ Was the right context retrieved?
   ├─ No  ─▶ retrieval/indexing failure (D3)
   └─ Yes ─▶ Is the answer unsupported by that context?
             ├─ Yes ─▶ hallucination / grounding
             └─ No  ─▶ Is only the hardest tier failing?
                       ├─ Yes ─▶ model mismatch (escalate tier)
                       └─ No  ─▶ prompt/format failure
```

---

## 4.8 The cost / latency / token optimisation playbook

Optimise only after you can measure (per-segment) and without regressing the quality bar.

| Lever | Cuts | Trade-off / caveat |
| --- | --- | --- |
| Prompt caching (stable-prefix-first) | Input cost, latency | Needs stable prefix; ~1024-token minimum |
| Batching (Message Batches API) | 50% cost | Up to 24 h latency — only latency-tolerant work |
| Routing / cascades | Cost | Adds a classifier/validation step |
| Output trimming / structured outputs | Output cost, latency | Don't drop needed content |
| Effort tuning (lower where adequate) | Thinking tokens, latency | Too low hurts hard tasks |
| Model right-sizing | Cost, latency | Re-validate quality on the smaller model |
| Fewer tool round-trips / progressive discovery | Latency, tokens | Requires tool-set discipline |

```text
Cost too high?  ─▶ cache stable prefix ─▶ route cheap-first ─▶ batch tolerant work ─▶ trim output
Latency too high? ─▶ smaller/faster model ─▶ fast mode ─▶ fewer round-trips ─▶ lower effort ─▶ stream
(after each change: re-run the regression suite per segment)
```

:::tip[Exam signal]
Optimisation answers that skip re-validation are wrong. Every cost/latency change must be **re-evaluated per segment** — a cheaper model or lower effort that quietly regresses one segment is a failure, not a win.
:::

---

## 4.9 Canary rollouts of prompts and models

Never flip a prompt or model change to 100% at once.

<Steps>

1. Pass the offline regression suite (per segment).

2. Release to a **small canary** slice of live traffic with online metrics and automatic rollback guardrails.

3. Compare canary vs control on real outcomes and cost; ramp by percentage if healthy.

4. Keep the previous version deployable for instant rollback.

</Steps>

---

## 4.10 A/B significance reasoning with worked numbers

The exam does not require you to run a t-test by hand, but it does require you to know **when a difference is real**. The key intuition: a small sample can produce a large apparent win by chance.

**Worked example — small sample.** Prompt B beats A on **7 of 12** paired items (58% win rate).

```text
Under the null (no difference), each item is a coin flip (p = 0.5).
Getting ≥ 7 of 12 heads by pure chance is very common (~39% two-sided).
→ 7/12 is well within noise. Do NOT promote B.
```

**Worked example — large sample.** On **2,000** paired items B wins **1,080** (54%).

```text
Expected under null = 1,000; std dev ≈ sqrt(n·p·(1-p)) = sqrt(2000·0.25) ≈ 22.4
Observed excess = 1,080 − 1,000 = 80 wins  →  z ≈ 80 / 22.4 ≈ 3.6
z ≈ 3.6 ⇒ p < 0.001  → the 54% win is statistically significant.
```

Same-direction result (B better), but only the large sample supports promotion — and only if it also holds **per segment** (no segment regresses).

| Situation | Correct action |
| --- | --- |
| Big win, tiny sample | Collect more data; do not promote |
| Small win, huge sample, significant, no segment regression | Promote via canary |
| Significant overall but one segment regresses | Do **not** promote; the aggregate hides a regression |
| Judge is uncalibrated | Fix calibration before trusting any A/B result |

:::tip[Exam signal]
Any option that says "ship B because it won N of a dozen" is the trap. The correct answer invokes **larger sample + statistical significance + no per-segment regression**. "MOST" and "FIRST" qualifiers usually point at "gather a significant sample" over "ship now".
:::

---

## 4.11 An observability and eval-record schema

To evaluate and debug at scale you need a consistent record per request. A minimal schema:

```json
{
  "correlation_id": "9f3a-...",
  "request_id": "req_01H...",
  "timestamp": "2026-09-15T11:02:33Z",
  "segment": { "tenant": "42", "language": "es", "ticket_type": "refund" },
  "model": "claude-sonnet-5",
  "prompt_version": "support-agent@3.2.1",
  "retrieval": { "k": 50, "kept": 6, "chunk_ids": ["c_101","c_233"], "recall_at_k": 0.94 },
  "usage": { "input_tokens": 6120, "output_tokens": 1440, "thinking_tokens": 300 },
  "cost_usd": 0.0271,
  "latency_ms": { "ttft": 610, "total": 5230 },
  "stop_reason": "end_turn",
  "eval": { "faithfulness": 0.9, "rubric_score": 1.7, "judge_model": "claude-opus-5" },
  "outcome": { "human_override": false, "user_thumb": "up" }
}
```

| Field group | Enables |
| --- | --- |
| `segment` | Per-segment metrics (the anti-aggregate control) |
| `prompt_version` / `model` | Attribute regressions to a specific change; A/B attribution |
| `retrieval` | Diagnose retrieval vs grounding; recall@k tracking |
| `usage` / `cost_usd` | Cost-per-task dashboards; routing decisions |
| `latency_ms` | p50/p95 SLA tracking (track `total`, watch the tail) |
| `eval` / `outcome` | Offline judge scores vs online human signals |

:::caution[If it isn't logged, you can't evaluate it]
Segment tags, prompt/model version, retrieval detail and per-request cost are the fields most often missing — and their absence is why teams can only report a misleading aggregate. Design the schema before launch.
:::

---

## 4.12 Scenario walkthrough: promoting a prompt change safely

**Scenario.** A team hand-tests a new support prompt on 15 favourite tickets, sees "clearly better" answers, and wants to ship to 100% tomorrow. The system serves English and Spanish tickets and three tenant tiers. The judge is the same Sonnet 5 session that generated the answers.

**Expert reasoning trace.**

<Steps>

1. **Reject the sample and the judge.** 15 cherry-picked items prove nothing, and grading in the producing session carries its bias (anti-pattern #9). First fix the method.

2. **Build the eval set.** Use a **stratified** golden set covering both languages and all three tiers, plus edge cases mined from production failures.

3. **Independent, calibrated judge.** Grade with a **separate model/session** (or humans), calibrated against human labels to catch length/position bias.

4. **Paired A/B with significance.** Run A vs B on the same inputs; require a **statistically significant** win on a large enough sample, and confirm **no per-segment regression** (e.g. Spanish must not drop).

5. **Gate in CI, then canary.** The per-segment regression suite must pass; then canary on a small live slice with rollback before ramping.

6. **Keep the prior version hot** for instant rollback.

</Steps>

**Why the tempting alternatives are wrong:** "ship because it looked better on 15" is small-sample eyeballing; "trust the same-session self-grade" is same-session bias; "flip to 100% to gather signal faster" removes the safety net; "test only English because that's most traffic" reintroduces the aggregate-blindness that hid the Spanish failure.

---

## 4.13 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "A big win on a few examples means B is better." | Small samples are dominated by noise; test for significance. | Small-sample eyeballing is the classic A/B trap. |
| "Overall accuracy is the metric that matters." | Per-segment metrics catch failures the aggregate hides. | Anti-pattern #10 appears in most D4 items. |
| "Mean latency reflects user experience." | Interactive SLAs live on p95/p99 tail. | Mean-only distractors are wrong. |
| "The model can grade its own answer." | Same-session self-review carries the producing bias. | Independent, calibrated judge is required. |
| "LLM-as-judge is objective." | Judges have length/position/self-preference bias; calibrate. | Length-bias stems test this. |
| "Optimise cost, then measure later." | Every cost change must be re-validated per segment. | Skipping re-eval is a wrong answer. |
| "Fix the prompt when output is wrong." | Diagnose the failing layer first (retrieval/grounding/prompt/model). | Fixing the wrong layer is the expensive mistake. |

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Reporting only aggregate accuracy | Masks per-segment failures (anti-pattern #10) |
| Using mean latency as the SLA metric | Hides tail; use p95 |
| LLM-as-judge in the same session/model as the system | Same-session self-review bias (anti-pattern #9) |
| Trusting a judge without calibrating against humans | Judge biases (position/length/self-preference) go undetected |
| "B looks better on 10 examples, ship it" | No statistical significance; small-sample noise |
| Optimising cost without re-running evals | May silently regress a segment |
| Fixing the prompt when retrieval was stale | Wrong layer; diagnose before fixing |
| Flipping a model change to 100% at once | No canary; no safe rollback |
| Golden set skewed to one segment | Aggregate looks fine while a segment fails |
| Batching latency-sensitive interactive traffic | 24 h latency violates the SLA |
| Promoting a variant that wins overall but regresses one segment | Aggregate hides the regression; require per-segment no-regression |
| Trusting an A/B result from an uncalibrated judge | Judge bias contaminates the comparison; calibrate first |
| Omitting segment/version/cost fields from the eval record | Forces misleading aggregate-only reporting |
| Confusing a big win rate on a tiny sample with significance | Small samples are noise; compute/estimate significance |
| Tracking only `total` latency without TTFT for streaming UX | TTFT drives perceived latency in interactive apps |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · A support bot reports 92% overall accuracy, but complaints come only from refund-related tickets. What does the architect need? (Select one)">
    A. Nothing; 92% exceeds the target.
    B. Per-segment metrics that break accuracy down by ticket type, revealing the refund segment's failure the aggregate hides.
    C. A bigger golden set of the same mix.
    D. A higher-effort model for all tickets.

    **Answer: B.** Aggregate accuracy masks a segment failure (anti-pattern #10). Stratified, per-segment metrics expose the refund problem so it can be fixed. The aggregate (A) is misleading; a bigger same-mix set (C) still hides it; escalating all tickets (D) overpays without diagnosing.
  </AccordionItem>

  <AccordionItem title="Q2 · An eval uses the same model and conversation that produced the answer to grade whether the answer is correct. Why is this unsound? (Select one)">
    A. It costs too much.
    B. Same-session self-review retains the reasoning bias that produced the answer; the judge must be a separate model/session and be calibrated against humans.
    C. The judge should always be a bigger model.
    D. LLM-as-judge is never valid.

    **Answer: B.** Grading in the producing session carries the original bias (anti-pattern #9). Independence plus human calibration is required. Cost (A) isn't the core issue; the judge need not be bigger (C); LLM-as-judge is valid when independent and calibrated (D).
  </AccordionItem>

  <AccordionItem title="Q3 · Prompt B beats prompt A on 6 of 10 hand-picked examples. What is the correct conclusion? (Select one)">
    A. Ship B; it won the comparison.
    B. The sample is far too small to be significant; run a paired A/B on a large, per-segment eval set and test for statistical significance before promoting.
    C. Ship A; it is the incumbent.
    D. Average the two prompts.

    **Answer: B.** A 6/10 result is noise; promotion requires a paired test on a representative sample with statistical significance and no per-segment regression. Shipping on eyeballed small samples (A), defaulting to the incumbent (C), or "averaging" prompts (D) are all unsound.
  </AccordionItem>

  <AccordionItem title="Q4 · Average latency is 2 s but users complain of slow responses. Metrics show p95 = 11 s. What should the SLA track? (Select one)">
    A. Mean latency only.
    B. Tail latency (p95/p99), since the mean hides the slow tail users actually experience.
    C. Total request count.
    D. Token count only.

    **Answer: B.** A healthy mean with a bad p95 means the tail is the real problem; SLAs for interactive systems track p95/p99. The mean (A) hides it; request count (C) and tokens (D) are not latency SLAs.
  </AccordionItem>

  <AccordionItem title="Q5 · An LLM judge consistently rates the longer of two answers higher regardless of correctness. What is happening and the fix? (Select two)">
    A. Length bias in the judge.
    B. Calibrate against human labels and revise the rubric to score correctness/grounding explicitly, not length.
    C. The judge is perfectly reliable.
    D. Always pick the longer answer.
    E. Delete the eval entirely.

    **Answer: A and B.** The judge exhibits length bias; the remedy is calibration against humans and a rubric that scores the dimensions that matter. The judge is not reliable (C), rewarding length (D) is the bug, and deleting evals (E) abandons measurement.
  </AccordionItem>

  <AccordionItem title="Q6 · A change fixed accuracy for enterprise-tier tickets but no one checked other tiers. How should this be prevented in future? (Select one)">
    A. Manual spot-checks after release.
    B. A per-segment regression suite in CI that blocks merge if any segment regresses.
    C. Trust the author's judgement.
    D. Only test the segment that changed.

    **Answer: B.** A per-segment regression suite in CI is exactly the guard against fixing one segment while breaking another. Manual spot-checks (A) and author trust (C) are unreliable; testing only the changed segment (D) is what caused the risk.
  </AccordionItem>

  <AccordionItem title="Q7 · A RAG answer includes a fact not present in the retrieved chunks, though recall@k is high. Which layer failed and what is the fix? (Select one)">
    A. Retrieval; re-index.
    B. Generation/grounding (hallucination); tighten the answer-only-from-context instruction, add citations, and verify faithfulness — retrieval is fine.
    C. Model mismatch; use Haiku.
    D. Infrastructure; add retries.

    **Answer: B.** High recall means the right context was present, so an unsupported fact is a grounding/hallucination failure at generation. Re-indexing (A) targets a healthy layer; a smaller model (C) won't help; retries (D) are unrelated.
  </AccordionItem>

  <AccordionItem title="Q8 · A team wants to cut cost 50% on a nightly bulk-summarisation job that has no latency requirement. Which lever is BEST and what must follow? (Select one)">
    A. Lower effort blindly and ship.
    B. Move the job to the Message Batches API (50% discount, within 24 h) and re-run the per-segment regression suite to confirm no quality regression.
    C. Switch to Haiku for all interactive traffic too.
    D. Disable evaluation to save compute.

    **Answer: B.** A latency-tolerant bulk job is the Batch API's use case; every cost change is followed by per-segment re-validation. Blind effort cuts (A) risk quality; changing interactive traffic (C) is out of scope; disabling evals (D) removes the safety net.
  </AccordionItem>

  <AccordionItem title="Q9 · A system fails only on the hardest 5% of cases and is fine elsewhere. What is the most likely cause and fix? (Select one)">
    A. Retrieval failure; re-chunk everything.
    B. Model mismatch — the tier is too small for the hardest cases; route/escalate those to a stronger model (cascade) while keeping the cheap model for the rest.
    C. Prompt failure; rewrite the whole prompt.
    D. Infrastructure; add more replicas.

    **Answer: B.** A clean 'fails only on the hardest cases' signature points to model capability; a cascade escalates just those cases, preserving cost elsewhere. Re-chunking everything (A) and rewriting the prompt (C) target layers that work on the other 95%; replicas (D) don't affect correctness.
  </AccordionItem>

  <AccordionItem title="Q10 · How should a new prompt version reach production safely? (Select two)">
    A. Pass the offline per-segment regression suite first.
    B. Canary on a small live slice with online metrics and automatic rollback, then ramp.
    C. Flip to 100% immediately to gather signal faster.
    D. Let each team edit the prompt in place.
    E. Skip offline evals if the change is small.

    **Answer: A and B.** Safe rollout is offline regression gate → canary with rollback → ramp. Flipping to 100% (C) removes the safety net; in-place per-team edits (D) destroy governance; skipping offline evals for 'small' changes (E) is how regressions slip through.
  </AccordionItem>

  <AccordionItem title="Q11 · Which pair are the RIGHT metrics for a high-volume, cost-sensitive classification service with an interactive SLA? (Select two)">
    A. Cost per task.
    B. p95 latency.
    C. Mean latency only.
    D. Total tokens generated across the fleet.
    E. Number of prompt versions.

    **Answer: A and B.** Cost-sensitive + interactive → cost per task and p95 (tail) latency are the binding metrics. Mean latency (C) hides the tail; total tokens (D) and prompt-version count (E) are not user-facing quality/SLA metrics.
  </AccordionItem>

  <AccordionItem title="Q12 · An eval dataset is 85% English support tickets; the system later fails badly on Spanish tickets in production. What was the flaw and the fix? (Select one)">
    A. The dataset was too small.
    B. The dataset lacked per-segment coverage (language); stratify the golden set so every language/segment is represented and reported separately.
    C. The model is broken.
    D. Online evals are unnecessary.

    **Answer: B.** An unstratified set makes a whole segment invisible to offline evals. Stratifying by language (and other segments) with per-segment reporting surfaces the gap. Size (A) isn't the issue; the model isn't broken (C); online evals (D) are still needed but the root cause here is coverage.
  </AccordionItem>

  <AccordionItem title="Q13 · On 2,000 paired items prompt B wins 1,080 (54%); on a separate 12-item hand test B won 7. Which result should drive promotion, and why? (Select one)">
    A. The 12-item test, because the answers looked clearly better.
    B. The 2,000-item result: a 54% win at n=2,000 is ~3.6 standard deviations from chance (p < 0.001), so it is statistically significant — provided no segment regresses.
    C. Neither; A/B testing is unreliable.
    D. Average both win rates.

    **Answer: B.** At n=2,000 the excess of 80 wins over the 1,000 expected is ≈3.6σ (std dev ≈ 22.4), which is significant; 7/12 is within coin-flip noise. The small test (A) is noise; A/B is reliable at scale (C); averaging win rates (D) is meaningless.
  </AccordionItem>

  <AccordionItem title="Q14 · A prompt change wins significantly overall but Spanish-tier accuracy drops 6 points. What is the correct decision? (Select one)">
    A. Promote it; the overall win is significant.
    B. Do not promote as-is; a per-segment regression blocks promotion even when the aggregate improves — fix the Spanish regression first.
    C. Promote and monitor complaints.
    D. Drop Spanish from the eval set.

    **Answer: B.** Per-segment no-regression is a hard gate; an aggregate win that hides a segment regression is anti-pattern #10. Promoting anyway (A, C) ships a known regression; dropping Spanish (D) re-hides it.
  </AccordionItem>

  <AccordionItem title="Q15 · A team can only report a single overall accuracy number and cannot tell which segment fails. Which observability gap is the root cause? (Select one)">
    A. Missing GPU metrics.
    B. Eval records lack a `segment` tag (and prompt/model version), so metrics cannot be stratified or attributed.
    C. The judge model is too small.
    D. Latency is not logged.

    **Answer: B.** Without segment tags and version fields, only an aggregate is computable and regressions can't be attributed. GPU metrics (A) are irrelevant; judge size (C) doesn't create the reporting gap; latency (D) is a different signal.
  </AccordionItem>

  <AccordionItem title="Q16 · Which TWO fields are MOST essential in a per-request eval record to support per-segment evaluation and A/B attribution? (Select two)">
    A. A `segment` tag (tenant/language/type).
    B. The `prompt_version` and `model` used.
    C. The server's CPU temperature.
    D. A random UUID with no linkage.
    E. The marketing campaign name.

    **Answer: A and B.** Segment tags enable stratified metrics; prompt/model version enables attributing regressions and A/B results to a specific change. CPU temperature (C), an unlinked UUID (D), and a campaign name (E) don't support evaluation.
  </AccordionItem>

  <AccordionItem title="Q17 · Before promoting a new prompt validated only by the same session that produced the answers, what must change FIRST? (Select one)">
    A. Increase the model's effort.
    B. Grade with an independent, calibrated judge (separate model/session) on a stratified set, because same-session self-review carries the producing bias.
    C. Ship to canary immediately.
    D. Delete the old prompt version.

    **Answer: B.** Same-session grading is anti-pattern #9; the method must be fixed with an independent, human-calibrated judge before any promotion decision. Effort (A) doesn't fix bias; canary (C) is premature; deleting the prior version (D) removes rollback.
  </AccordionItem>

  <AccordionItem title="Q18 · An interactive streaming UI feels slow even though total latency is acceptable. Which metric should be added? (Select one)">
    A. Total tokens generated.
    B. Time-to-first-token (TTFT), because perceived latency in streaming UIs is driven by how fast output starts.
    C. Daily request count.
    D. Number of prompt versions.

    **Answer: B.** In streaming UIs users perceive responsiveness by when tokens start (TTFT), not just total time. Total tokens (A), request count (C), and version count (D) are not latency-experience metrics.
  </AccordionItem>
</Accordions>

## Key takeaways

- Pick metrics that map to success criteria: accuracy, **p95** (not mean) latency, cost per task, safety, security — and report them **per segment**.
- Build **stratified** eval datasets (golden + synthetic + edge cases) so no segment is invisible.
- Use **rubric grading** and **LLM-as-judge in a separate model/session**, and **calibrate** the judge against human labels.
- Compare variants with **paired A/B tests** and require **statistical significance** plus no per-segment regression before promoting.
- Run **offline** evals in CI as a regression gate and **online** evals to catch distribution shift; you need both.
- **Diagnose the failing layer** (retrieval vs grounding vs prompt vs model vs segment) before fixing anything.
- Optimise with caching, batching, routing, trimming and effort tuning — then **re-validate per segment**.
- Roll out via **canary** with automatic rollback; never flip to 100% at once.
- Judge **significance, not vibes**: a big win on 12 items is noise; a small win on thousands can be real (z ≈ excess ÷ √(n·p·(1−p))) — and it must hold **per segment**.
- Design the **eval record schema** (segment, prompt/model version, retrieval, usage, cost, latency, judge score, outcome) before launch — you cannot evaluate what you did not log.
- Track **TTFT** for streaming UX alongside p95 total latency; perceived speed is driven by first-token time.
