AI Cert Prep
Type to search documentation.

Appendix · Claude

Evals Cookbook

Golden-set design, graders, LLM-as-judge calibration, A/B testing with significance, CI regression YAML, per-segment reporting and cost/latency dashboards.

Evals turn “it seems better” into evidence. The exam-correct posture: a golden set, the cheapest sufficient grader, judges in a separate session, per-segment reporting, and a CI regression gate on prompt/model changes.

Exam signal

“Overall accuracy is 95%” is a trap — demand per-segment metrics. “Are you sure?” in the same chat is a trap — use a fresh session/model. Comparing two prompts without significance is a trap — use pairwise A/B with a significance test.

Golden-set design

  1. Cover the distribution — sample real inputs across segments (document type, language, difficulty, edge cases), not just easy ones.
  2. Include hard/adversarial cases — injection attempts, ambiguous inputs, known past failures.
  3. Label with expected outputs or rubrics — exact answers where possible; rubrics for subjective tasks.
  4. Size — enough per segment to detect regressions (aim ≥ 30–50 per segment for stable rates).
  5. Version and freeze — a golden set is a controlled asset; changing it invalidates comparisons.
  6. Separate holdout — keep a set you never tune against to catch overfitting to the eval.

Graders — cheapest sufficient first

GraderUse whenCostReliability
Exact / normalised matchDeterministic answers (classification, extraction fields)FreeHigh
Regex / structuralFormat, presence of fields, schema validityFreeHigh
Heuristic (numeric tolerance, set overlap)Numbers, listsFreeMedium
Rubric gradingSubjective quality with clear criteriaLowMedium
LLM-as-judgeOpen-ended quality, pairwise preferenceHigherMedium (needs calibration)
HumanGround truth, calibration referenceHighestHighest

Escalate only as far as the task requires — exact match before rubric before LLM-judge.

LLM-as-judge calibration

The judge must run in a separate session, ideally a different model, and be calibrated against human labels before you trust it.

  1. Draft the rubric with explicit, non-overlapping criteria and a scale.
  2. Have humans label a calibration set (e.g., 100 items).
  3. Run the judge on the same set; measure agreement (e.g., Cohen’s κ, or % agreement).
  4. If agreement is low, tighten the rubric, add few-shot exemplars of each grade, or reduce the scale (binary is easier to calibrate than 1–10).
  5. Re-measure; only deploy the judge once agreement meets your bar (e.g., κ ≥ 0.6).
  6. Re-calibrate when the model, rubric or task shifts.
Calibration symptomFix
Judge too lenientAdd negative exemplars; sharpen “fail” criteria
Judge inconsistent run-to-runLower temperature; binary rubric; majority vote of 3
Position bias in pairwiseRandomise A/B order; average both orderings
Verbosity biasInstruct to ignore length; penalise unsupported claims

A/B testing with significance

Pairwise comparison of prompt/model B against A. Do not eyeball a few outputs — test.

Worked example: B wins 118 of 200 head-to-head comparisons (ties split). Is B really better?

text
n = 200, wins = 118, p̂ = 0.59
Two-sided test of p = 0.5:
z = (118 − 100) / sqrt(200 × 0.5 × 0.5) = 18 / 7.07 ≈ 2.55
p-value ≈ 0.011 → significant at α = 0.05
95% CI on win rate: 0.59 ± 1.96 × sqrt(0.59×0.41/200) = 0.59 ± 0.068 → [0.52, 0.66]

B is significantly better. Had B won 108/200: z ≈ 1.13, p ≈ 0.26 — not significant; do not ship on that alone.

PitfallFix
Too few samplesPower the test; more items for small effects
Peeking / stopping earlyFix n in advance or use sequential-test corrections
Ignoring tiesDefine handling up front
Aggregate onlyAlso test per segment

CI regression gate

Run the golden set on every prompt/model change; block the deploy if a segment regresses beyond tolerance.

.github/workflows/evals.yml
name: evals
on: [pull_request]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- run: pip install -r eval/requirements.txt
- name: Run golden set
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: python eval/run.py --golden eval/golden.jsonl --out eval/report.json
- name: Gate
run: |
python - <<'PY'
import json, sys
r = json.load(open("eval/report.json"))
# per-segment gate: no segment below its baseline minus 2 points
fails = [s for s, m in r["by_segment"].items() if m["accuracy"] < r["baseline"][s] - 0.02]
if fails:
print("Regressed segments:", fails); sys.exit(1)
print("All segments within tolerance.")
PY

Gate on per-segment thresholds, not just the aggregate — that is what catches a single failing document type.

Per-segment reporting

text
Segment n Accuracy Faithfulness Δ vs baseline
-------------- --- -------- ------------ -------------
invoices 120 0.94 0.95 +0.01
contracts 90 0.88 0.91 −0.00
receipts (scan) 60 0.61 ◄ 0.72 ◄ −0.18 ◄
-------------- --- -------- ------------ -------------
OVERALL 270 0.86 0.90 −0.03

Aggregate 0.86 looks fine; the scanned-receipts segment at 0.61 is the real story. Always report the breakdown.

Cost / latency dashboards

MetricWhyWatch for
Cost per request / per 1kBudgetSpike after a model or prompt change
Tokens in/out (p50/p95)Cost + context pressureGrowing prompts → cache or trim
Latency p50 / p95 / p99UX / SLOp95 breach even when p50 is fine
Cache hit rateCost efficiencyDrop → prefix churned
Error rate by type (429/5xx/refusal)Reliability429 → rate-limit tier; refusals → prompt/policy
Fallback rateCapacityHigh → capacity or model issue
Quality (eval score) over timeRegressionSilent drift after upstream changes

Report p95/p99, not just averages — an average hides the tail that violates the SLO.

Common misconceptions

MisconceptionRealityWhy it matters on the exam
“95% overall means we’re good”A segment may fail; report per-segmentAggregate-metric anti-pattern
“The model can grade its own answer”Use a separate session/model, calibratedSame-session-review anti-pattern
“B looked better in a few outputs”Test significance with enough samplesEyeballing distractor
“LLM-judge is objective”Needs calibration vs humans; has biasesUncalibrated-judge distractor
“Average latency meets the SLO”Watch p95/p99 tailsAggregate-latency distractor
“Evals are a one-time thing”CI regression + online monitoring, continuousPoint-in-time distractor
“Change the golden set to pass”Freeze it; changing invalidates comparisonOverfitting distractor

Scenario walkthrough

A team wants to switch the extraction pipeline from Sonnet 5 to Haiku 4.5 to cut cost. Aggregate accuracy on a quick test looks “about the same”. How do you decide correctly?

  1. Golden set, per-segment — run both models on the frozen set, break results down by document type.
  2. Right grader — exact/structural match on the extracted fields (deterministic), not an LLM judge.
  3. Significance — where subjective, pairwise A/B with a significance test, not eyeballing.
  4. Find the hidden regression — Haiku matches on clean PDFs but drops 18 points on scanned receipts (aggregate hid it).
  5. Decide by constraint — if scanned receipts matter, keep Sonnet 5 for that segment (cascade), Haiku for the rest.
  6. Gate it — add the per-segment thresholds to CI so a future switch cannot silently regress.
  7. Watch cost + latency — confirm the saving is real on the dashboard (p95, cost per 1k).

Rejected alternatives: switching on aggregate accuracy (aggregate-metric), judging with the same model in the same session (same-session review), and eyeballing “about the same” without significance (eyeballing distractor).

Key takeaways

  • Golden set first: cover the distribution, include hard cases, version and freeze it.
  • Use the cheapest sufficient grader; LLM-as-judge must be separate-session and calibrated.
  • Prove improvements with pairwise A/B + significance, not a few outputs.
  • Gate deploys with a CI regression suite on per-segment thresholds.
  • Monitor cost and latency at p95/p99, and quality continuously — evals are ongoing, not one-time.

Last updated Sep 18, 2026