# D3 · Evaluating AI Applications

Designing evals with representative datasets and graders, offline versus online evaluation, preventing regressions, failure analysis and the prompt optimizer for AI applications on the OpenAI API.

import { Accordions, AccordionItem } from '@prosefly/astro-components';

This domain is about **16%** of the OAI-API mock – roughly **10 of 60 items** – and mirrors the Academy *Evaluate AI Applications* course (70 min). It tests whether you can prove an AI application works and keep it working: building a representative dataset, choosing a grader, running offline before online, catching regressions, and analysing failures instead of guessing. One nuance to know: **the Evals API is now grouped under "Legacy APIs" in the docs**, but the *practice* of evaluating is as central as ever – say "the Evals API (now grouped with legacy APIs)" rather than presenting it as the newest surface.

## What you need to know

An eval is a dataset plus a grader plus a metric. The dataset must be **representative** of real inputs, including the hard and rare cases, not a handful of cherry-picked examples. The grader decides pass/fail: exact match or code for deterministic checks, or an **LLM-as-judge** with a rubric for open-ended quality. You run **offline** evals before shipping (on a fixed dataset) and **online** evals after (sampling live traffic), and you keep a baseline so any prompt or model change that lowers the score is caught as a **regression**. When something fails, you do **failure analysis** – cluster the failures, find the pattern, fix the cause – rather than tweaking the prompt at random. The **prompt optimizer** can improve a prompt against your eval.

## Learning objectives

By the end of this page you should be able to:

1. **Build a representative eval dataset**, including edge and adversarial cases.
2. **Choose a grader** – exact match, code-based, or LLM-as-judge with a rubric.
3. Distinguish **offline** and **online** evaluation and use each correctly.
4. **Prevent regressions** by baselining and re-running evals on every change.
5. **Analyse failures** by clustering rather than tweaking at random.
6. Use the **prompt optimizer** and know where evals sit in the current docs.

---

## 3.1 What an eval is

```text
        ┌──────────────┐    ┌────────────┐    ┌──────────┐
INPUT ─►│  your app     │─► │  OUTPUT     │─►│  GRADER   │─► pass / score
        │ (prompt+model)│    │             │    │(rule or  │
        └──────────────┘    └────────────┘    │  judge)  │
                ▲                              └──────────┘
                │                                   │
          DATASET (representative inputs)     METRIC (accuracy, rubric score)
```

An eval turns "it looks better" into a number you can compare across versions. Without one, every prompt change is a guess and every regression ships silently.

:::tip[Assessment signal]
Stems with "how do you know it works", "prove the change is an improvement", "before you ship", or "the new prompt seems better" are eval items. The correct answer measures on a dataset; "it looks good in a few tries" is the distractor.
:::

## 3.2 Representative datasets

The dataset is the eval. A biased dataset gives a confident, wrong answer.

| Property | Why it matters | How to get it |
| --- | --- | --- |
| **Representative** | Score must reflect real traffic | Sample real inputs, not invented easy ones |
| **Covers edge cases** | Rare inputs are where failures hide | Add the hard, ambiguous, adversarial cases deliberately |
| **Labelled / has references** | Grading needs ground truth or a rubric | Have SMEs label; or define rubric criteria |
| **Right size** | Too small is noisy; too big is slow/costly | Enough that a 1–2% change is meaningful (often 100–500) |
| **Held out** | Prevents tuning to the test | Keep a set the prompt was never tuned on |

## 3.3 Choosing a grader

| Grader | Use for | Example |
| --- | --- | --- |
| **Exact / string match** | Deterministic, single correct answer | Classification label, extracted field |
| **Code-based** | Checkable rules, structure, ranges | "Is valid JSON and `priority` in 1–4" |
| **LLM-as-judge** | Open-ended quality against a rubric | "Is the summary faithful and complete?" |
| **Human** | Highest-stakes or rubric calibration | Sampling to validate the judge |

For an LLM judge, write an explicit rubric and validate it against human labels on a sample – an unchecked judge just moves the trust problem.

```python
grader = {
    "type": "score_model",
    "model": "gpt-5.6-terra",
    "input": [{"role": "system", "content": (
        "Score the ANSWER 1-5 for faithfulness to CONTEXT. "
        "5 = every claim is supported; 1 = unsupported claims. "
        "Return only the integer.")}],
}
```

## 3.4 Offline versus online

```text
OFFLINE  ── fixed dataset, before you ship ──►  gate the release
   │                                             (regression check)
   ▼
ONLINE   ── sample of live traffic, after ──►   catch real-world drift
                                                 (new inputs, edge cases)
```

- **Offline** evals run on a curated dataset in CI-style gates; they answer "is version B at least as good as version A?"
- **Online** evals sample production traffic and grade it (often with an LLM judge or thumbs-up signals); they catch inputs your dataset never had.

Both matter: offline stops known regressions before release; online surfaces unknown failures after.

## 3.5 Preventing regressions

The value of an eval is comparison over time. Keep a **baseline** score and re-run the eval on every prompt edit, model change, or dependency bump.

| Change | Regression risk | Guard |
| --- | --- | --- |
| Prompt edit | Fixes one case, breaks two others | Re-run the full eval, compare to baseline |
| Model upgrade | Behaviour shifts subtly | Re-run before switching production traffic |
| Reasoning-effort change | Quality/cost trade shifts | Re-measure both quality and cost |
| New retrieval source | Grounding changes | Re-run grounded-answer eval |

:::caution[The single-example trap]
"It handled my one hard example, ship it" is the most common regression cause: a change that fixes the visible case silently breaks unseen ones. Only the full dataset comparison protects you.
:::

## 3.6 Failure analysis

When the score drops, do not tweak randomly. **Cluster the failures and find the pattern.**

```text
1. Pull every failing item from the eval run.
2. Read them; group by shared cause
   (e.g., "long inputs", "numbers", "one category", "ambiguous asks").
3. Form a hypothesis about the cause.
4. Change ONE thing that addresses the largest cluster.
5. Re-run the eval; confirm the cluster shrank without new regressions.
```

Fixing the largest cluster first is the highest-leverage move; chasing individual failures is slow and often introduces regressions.

## 3.7 The prompt optimizer and where evals live

- The **prompt optimizer** takes a prompt and an eval and proposes improved wording measured against your metric – it automates the tune-and-measure loop, but it needs a good dataset and grader to be meaningful.
- In the current docs, **Evals, fine-tuning, Agent Builder and the Assistants API are grouped under "Legacy APIs".** Evals still matter conceptually – getting started, working with evals, the prompt optimizer, external models, best practices, graders – so describe it as "the Evals API (now grouped with legacy APIs)" and keep using the practice; do not present it as the newest surface or drop it.

## Decision framework

Use the **PROVE** framework to make any quality claim defensible.

| Letter | Step | Question |
| --- | --- | --- |
| **P** | Pick the metric | What number defines "good enough"? |
| **R** | Representative data | Does the dataset match real traffic and edge cases? |
| **O** | One grader | Exact match, code, or a validated LLM judge? |
| **V** | Versus baseline | Is the new version at least as good as the last? |
| **E** | Examine failures | What is the largest failure cluster, and its cause? |

The step teams skip most is **R** – they evaluate on easy invented examples and are surprised in production.

## Common mistakes

| Mistake | Why it happens | What to do instead |
| --- | --- | --- |
| Shipping on "it looks better" | Vibes feel like evidence | Compare a metric on a held-out dataset |
| Evaluating on easy invented examples | Faster than sampling real inputs | Build a representative set with edge cases |
| No baseline to compare against | Nobody recorded version A's score | Store the baseline; every change is compared to it |
| Trusting an unvalidated LLM judge | The judge seems authoritative | Validate the judge against human labels on a sample |
| Tweaking the prompt at random after a failure | Faster than analysis | Cluster failures, fix the largest cause, re-measure |
| Only running offline evals | CI feels sufficient | Add online evals to catch real-world drift |
| Testing one example and generalising | The visible case is compelling | The single example is where regressions hide |
| Treating evals as obsolete because they are 'legacy' | The docs regrouped them | The practice is central; use the Evals API knowingly |

## Scenario challenge

**Scenario.** Your RAG-based policy assistant answers HR questions. A colleague edits the system prompt to fix a complaint about verbose answers, tries three questions that now look tighter, and wants to deploy. You have an eval of 240 real questions with reference answers and a faithfulness judge, plus a baseline of 91% faithful / 88% complete.

**Expert reasoning trace.**

1. **Reject "three questions looked good".** Three cherry-picked cases are not evidence; the change might have improved brevity at the cost of completeness on the questions no one checked.
2. **Run the full eval against the baseline.** Suppose faithfulness holds at 91% but completeness drops to 79%. That is a regression: the shorter prompt is now omitting required detail. Without the dataset comparison, this ships silently.
3. **Do failure analysis, not another prompt tweak.** Pull the newly failing items. If they cluster on multi-part questions ("what is the policy and how do I appeal?"), the cause is the prompt trimming secondary parts, not verbosity in general.
4. **Change one thing.** Adjust the prompt to keep completeness on multi-part questions while trimming filler, then re-run. Confirm completeness recovers toward baseline without hurting faithfulness.
5. **Consider online too.** Even a passing offline eval will not have every real phrasing; add an online faithfulness sample so drift surfaces after launch.
6. **Use the optimizer with care.** The prompt optimizer could propose wording, but only because you now have a representative dataset and a validated grader for it to optimise against.

**Exam-correct decision:** run the full eval against the baseline, catch the completeness regression, cluster the failures, fix the largest cause, re-measure, and add an online sample. **Not** deploy on three good-looking answers, **not** keep tweaking wording without measuring.

## Assessment traps

| Trap | Why it is tempting | The discriminator |
| --- | --- | --- |
| "It handled my hard example, ship it" | The visible case is convincing | One example hides regressions; compare the full dataset |
| "The new answers read better, deploy" | Fluency feels like quality | A metric on a representative set is the evidence |
| "The LLM judge said 5/5, so it's right" | Judges look authoritative | Validate the judge against human labels first |
| "Offline eval passed, we're done" | CI gating feels complete | Online evals catch inputs the dataset lacked |
| "Evals are legacy, skip them" | The docs regrouped them | The practice is central; use it deliberately |
| "Tweak the prompt until failures drop" | Feels like progress | Cluster failures and fix the cause, then re-measure |

## Practice questions

Each item states how many responses to select. Commit before revealing.

<Accordions>
  <AccordionItem title="Q1 · A developer changes a prompt, tries two examples that look better, and wants to deploy. What should they do FIRST? (Select one)">
    A. Deploy; two good examples are enough.
    B. Run the full eval on a representative dataset and compare the score to the baseline.
    C. Ask the model if the new prompt is better.
    D. Increase the reasoning effort.

    **Answer: B.** Only a dataset comparison against the baseline shows whether the change helps overall. Two examples (A) hide regressions, asking the model (C) is not evidence, and raising effort (D) is unrelated to validating the change.
  </AccordionItem>

  <AccordionItem title="Q2 · Which dataset is MOST appropriate for evaluating a support-triage classifier? (Select one)">
    A. Ten hand-written easy examples.
    B. A representative sample of real tickets including rare categories and ambiguous cases, labelled by SMEs.
    C. Only the tickets the model already gets right.
    D. Synthetic tickets generated by the same model.

    **Answer: B.** A representative, edge-case-covering, SME-labelled set reflects real performance. Easy invented examples (A) and model-generated tickets (D) are unrepresentative, and evaluating only on wins (C) is circular.
  </AccordionItem>

  <AccordionItem title="Q3 · You need to grade whether extracted invoice fields exactly match ground truth. Which grader is BEST? (Select one)">
    A. LLM-as-judge with a subjective rubric.
    B. Exact / code-based match against the labelled field values.
    C. Human review of every invoice.
    D. Thumbs-up signals from users.

    **Answer: B.** Deterministic field extraction has ground truth, so exact/code matching is precise and cheap. An LLM judge (A) adds noise for a deterministic check, per-invoice human review (C) does not scale, and thumbs-up (D) is not field-level ground truth.
  </AccordionItem>

  <AccordionItem title="Q4 · What is the difference between offline and online evaluation? (Select one)">
    A. Offline uses a fixed curated dataset before shipping; online samples live production traffic after.
    B. Offline is for small teams and online is for large teams.
    C. They are the same thing with different names.
    D. Online evals replace the need for offline evals.

    **Answer: A.** Offline evals gate releases on a fixed dataset; online evals grade real traffic to catch drift. They are complementary (contradicting C and D), and the distinction is about timing and data, not team size (B).
  </AccordionItem>

  <AccordionItem title="Q5 · An LLM-as-judge grader is scoring summaries for faithfulness. What must you do before trusting its scores? (Select one)">
    A. Nothing; the model is authoritative.
    B. Validate the judge against human labels on a sample and refine the rubric until they agree.
    C. Use the same model that produced the summaries as the judge, with no rubric.
    D. Trust it only if it uses `gpt-6-astra`.

    **Answer: B.** An LLM judge must be calibrated against human labels with an explicit rubric, or it just relocates the trust problem. It is not automatically authoritative (A), self-grading with no rubric (C) is weak, and model size alone (D) does not validate it.
  </AccordionItem>

  <AccordionItem title="Q6 · After a model upgrade, the eval score drops on a cluster of long-input questions. What is the correct response? (Select two)">
    A. Read the failing long-input items and form a hypothesis about the cause.
    B. Change one thing that addresses the cluster, then re-run the eval.
    C. Roll back and never upgrade any model.
    D. Randomly reword the prompt until the score recovers.
    E. Ignore it because the average is still acceptable.

    **Answer: A and B.** Failure analysis means clustering, hypothesising, changing one thing, and re-measuring. Never upgrading (C) is an overreaction, random rewording (D) risks new regressions, and ignoring a cluster (E) leaves a real failure mode in production.
  </AccordionItem>

  <AccordionItem title="Q7 · Why keep a baseline score for an application's eval? (Select one)">
    A. To show leadership a nice chart only.
    B. So every prompt, model or dependency change can be compared to it and regressions are caught before shipping.
    C. Because the API requires it.
    D. To avoid having to write graders.

    **Answer: B.** A baseline is the reference every change is measured against, which is how regressions are detected. Charts (A) are a side effect, the API does not require it (C), and it does not replace graders (D).
  </AccordionItem>

  <AccordionItem title="Q8 · How should you describe the Evals API given the current OpenAI docs structure? (Select one)">
    A. As the newest, recommended surface for all new work.
    B. As the Evals API now grouped under Legacy APIs, while the practice of evaluating remains central.
    C. As fully removed and unavailable.
    D. As identical to the Responses API.

    **Answer: B.** The docs regrouped Evals (with fine-tuning, Agent Builder, Assistants) under Legacy APIs, yet evaluation is still essential practice. It is neither the newest surface (A), nor removed (C), nor the same as Responses (D).
  </AccordionItem>

  <AccordionItem title="Q9 · A team runs only offline evals and is surprised by production failures on phrasings never in their dataset. What is missing? (Select one)">
    A. A larger model.
    B. Online evaluation sampling live traffic to catch inputs the offline dataset lacked.
    C. Higher reasoning effort.
    D. More prompt engineering.

    **Answer: B.** Offline evals cannot cover inputs they never contained; online evals sample real traffic to surface them. A bigger model (A), more effort (C) or more prompt work (D) do not address the missing feedback loop.
  </AccordionItem>

  <AccordionItem title="Q10 · What does the prompt optimizer require to be useful? (Select one)">
    A. Nothing; it improves prompts blindly.
    B. A representative eval dataset and a grader, because it optimises the prompt against your metric.
    C. A `gpt-6-astra` subscription only.
    D. That you disable all other evals.

    **Answer: B.** The optimizer tunes a prompt against a metric, so it needs a dataset and grader to measure against. It is not blind (A), does not depend on a specific model tier (C), and does not require disabling evals (D) – it uses them.
  </AccordionItem>

  <AccordionItem title="Q11 · A prompt change fixes the one complaint you received but you have no dataset. What is the SAFEST next step before shipping? (Select one)">
    A. Ship immediately; the complaint is resolved.
    B. Build a small representative dataset with references, set a baseline with the old prompt, then compare the new prompt.
    C. Ask three colleagues if it looks better.
    D. Increase max output tokens.

    **Answer: B.** Without a dataset you cannot detect regressions, so build one and baseline before comparing. Shipping on one fix (A) risks silent breakage, opinions (C) are not measurement, and output tokens (D) are irrelevant.
  </AccordionItem>

  <AccordionItem title="Q12 · Which TWO are legitimate reasons to include human review in an eval process? (Select two)">
    A. To calibrate and validate an LLM-as-judge grader on a sample.
    B. To grade the highest-stakes outputs where errors are costly.
    C. To replace the dataset entirely with opinions.
    D. To make the eval slower on purpose.
    E. Because automated graders can never be trusted for anything.

    **Answer: A and B.** Humans calibrate judges and grade high-stakes items where the cost of error justifies the effort. Replacing the dataset with opinions (C) removes rigour, slowing on purpose (D) is pointless, and automated graders are trustworthy for many deterministic checks (E).
  </AccordionItem>

  <AccordionItem title="Q13 · Your eval dataset is 12 items and results swing wildly between runs. What is the MOST likely problem? (Select one)">
    A. The model is broken.
    B. The dataset is too small to give a stable, meaningful signal; enlarge it to a representative size.
    C. Reasoning effort is too low.
    D. The grader should be human-only.

    **Answer: B.** A tiny dataset is noisy, so small changes swing the score; a representative set (often 100–500) stabilises it. The model is likely fine (A), effort (C) does not fix sample noise, and switching to humans (D) does not address size.
  </AccordionItem>

  <AccordionItem title="Q14 · A grounded-answer eval measures whether answers are supported by retrieved context. When should it be re-run? (Select one)">
    A. Only once, at the very start of the project.
    B. Whenever the prompt, model, or retrieval source changes, and periodically online.
    C. Never; retrieval is deterministic.
    D. Only when a user complains.

    **Answer: B.** Any change to prompt, model or retrieval can alter grounding, so re-run the eval on each and sample online. Running once (A) ignores drift, retrieval is not fully deterministic in grounding effect (C), and waiting for complaints (D) is reactive and late.
  </AccordionItem>
</Accordions>

## Key takeaways

- An eval is a representative dataset plus a grader plus a metric – it turns "looks better" into a comparable number.
- Build the dataset from real inputs and deliberately include edge and adversarial cases.
- Match the grader to the task: exact/code for deterministic checks, a validated LLM judge for open-ended quality.
- Run offline evals to gate releases and online evals to catch real-world drift.
- Keep a baseline and re-run on every prompt, model, effort or retrieval change to catch regressions.
- Do failure analysis by clustering, fixing the largest cause, and re-measuring – not by random tweaks.
- The Evals API is now grouped with Legacy APIs, but evaluation is still central practice; the prompt optimizer needs a good dataset and grader.
