Domains
D8 · Eval, Testing and Debugging
Identifying integration-layer vs model-output errors, recovery, trace analysis, request-ID logging, reproducibility, eval basics, regression tests in CI and per-segment metrics.
This domain is roughly 1 of 53 items, but it ties the others together. It tests whether you can tell an integration-layer failure from a model-output failure, reproduce and trace problems, and evaluate quality without fooling yourself. The theme: measure per-segment, reproduce deterministically, and judge output in a separate session.
Learning objectives
By the end of this page you should be able to:
- Distinguish integration-layer from model-output errors and choose recovery.
- Do trace analysis and log request IDs.
- Reproduce issues with a pinned model and
temperature: 0where appropriate. - Apply eval basics: golden set, rubric, LLM-as-judge in a separate session.
- Add regression tests in CI and use per-segment metrics.
8.1 Integration-layer vs model-output errors
The first question in any incident: which layer failed? Integration-layer errors live in the plumbing (auth, transport, request shape, rate limits, parsing). Model-output errors live in what Claude produced (hallucination, format drift, refusal, truncation). Retries and timeouts fix the former; prompt/context/eval changes fix the latter. Applying the wrong fix to the wrong layer is the classic distractor.
Error taxonomy
| Layer | Error | Diagnostic signal | Recovery |
|---|---|---|---|
| Integration | Auth | 401 authentication / 403 permission | Check API key, org, scopes; do not retry blindly |
| Integration | Schema / bad request | 400 invalid_request (e.g. tool_choice:"any" on Fable 5.1) | Fix request shape; use auto + strict:true / structured outputs |
| Integration | Timeout | Client timeout, no response | Set sane timeouts; retry with backoff; consider streaming |
| Integration | Rate limit | 429 rate_limit, retry-after header | Exponential backoff + jitter; respect retry-after; batch/queue |
| Integration | Parsing | JSON decode error on the response body | Validate/repair; request structured outputs; validation-retry |
| Model output | Hallucination | Fluent claim with no grounding in inputs | Grounding, citations, RAG, validation; verify claims |
| Model output | Format drift | Output nearly-valid but off-schema | strict:true tool schema / structured outputs; validation-retry |
| Model output | Refusal | stop_reason: "refusal" | Reframe legitimate request; adjust system prompt; escalate |
| Model output | Truncation | stop_reason: "max_tokens" | Raise max_tokens; chunk the task; stream and continue |
Exam signal
First ask which layer. A 429, a timeout, or a JSON decode error is integration — fix with backoff, timeouts, or request/schema changes. A hallucination, off-schema output, or refusal is model output — fix with prompt/context/model changes and validation. stop_reason and the HTTP status code tell you which layer.
8.2 Trace analysis walkthrough
Log the request-id and a correlation ID for every call, plus model, stop_reason, usage and latency. For agents, log the tool sequence. Traces let you pinpoint where a multi-step run went wrong and give Anthropic support a handle.
resp = client.messages.create(model=MODEL, max_tokens=512, messages=msgs)log.info("claude_call", extra={ "request_id": resp._request_id, "model": resp.model, "stop_reason": resp.stop_reason, "in": resp.usage.input_tokens, "out": resp.usage.output_tokens})Reading a trace of a stuck agent:
corr_id=abc-123 step 1 request_id=req_01A stop_reason=tool_use tools=[search_orders] in=1,240 out=180 step 2 request_id=req_01B stop_reason=tool_use tools=[get_order(id=9981)] in=1,910 out=95 step 3 request_id=req_01C stop_reason=tool_use tools=[get_order(id=9981)] in=2,600 out=95 ← repeat step 4 request_id=req_01D stop_reason=tool_use tools=[get_order(id=9981)] in=3,290 out=95 ← loop ... step N request_id=req_01Z stop_reason=max_tokens in=9,900 out=512 ← truncatedDiagnosis: the loop repeats the same tool call and input tokens climb each turn — the harness is not feeding the tool result back, or termination is driven by something other than stop_reason. This is anti-patterns #1/#2 (natural-language termination and iteration caps) rather than a model defect. The max_tokens at the end is a downstream symptom, not the cause. Fix the loop: after each tool_use, append the tool result and continue until stop_reason is end_turn.
8.3 Reproducibility
To reproduce a model-output issue, control every variable you can:
- Pin the exact model snapshot — not a floating alias.
- Set
temperature: 0where appropriate to minimise variance. - Freeze the system prompt and tools and replay the same messages in order.
resp = client.messages.create( model="claude-sonnet-5", # pinned snapshot temperature=0, # minimise variance max_tokens=512, system=FROZEN_SYSTEM_PROMPT, tools=FROZEN_TOOLS, messages=RECORDED_MESSAGES, # exact replay)Non-determinism means the output still is not byte-identical, but pinning + temperature: 0 + a frozen prompt makes issues far more repeatable and comparisons fair. Note: Fable 5.x / Opus 5 / Sonnet 5 harnesses must be append-only — editing or reordering earlier turns invalidates later thinking blocks, so replay by appending, never by rewriting history.
8.4 Eval basics and an eval harness
| Element | What it is |
|---|---|
| Golden set | Curated inputs with known-good outputs |
| Exact match | For deterministic outputs (labels, extractions) |
| Rubric | Explicit scoring criteria for open-ended output |
| LLM-as-judge | A different model/session scores output against the rubric |
Same-session self-review
Never have the model grade its own output in the same session — it retains the reasoning bias that produced it (anti-pattern #9). Use a separate session and ideally a different model as judge.
A minimal harness combining exact-match and an LLM-as-judge in a separate call:
import jsonfrom anthropic import Anthropic
client = Anthropic()GEN_MODEL = "claude-sonnet-5" # system under testJUDGE_MODEL = "claude-opus-5" # different model, separate call
def generate(case): r = client.messages.create( model=GEN_MODEL, temperature=0, max_tokens=512, system="Extract the invoice total as JSON: {'total_cents': int}.", messages=[{"role": "user", "content": case["input"]}], ) return r.content[0].text
def exact_match(pred, expected): try: return json.loads(pred).get("total_cents") == expected["total_cents"] except json.JSONDecodeError: return False # parsing failure counts as a miss, never silently passed
RUBRIC = ( "Score 1 if the answer is faithful to the source and correctly formatted, " "else 0. Return JSON: {'score': 0 or 1, 'reason': str}.")
def llm_judge(case, pred): # Separate session/model — no shared context with the generator. r = client.messages.create( model=JUDGE_MODEL, temperature=0, max_tokens=256, system=RUBRIC, messages=[{"role": "user", "content": f"SOURCE:\n{case['input']}\n\nANSWER:\n{pred}"}], ) return json.loads(r.content[0].text)
def run(golden): results = [] for case in golden: pred = generate(case) results.append({ "id": case["id"], "segment": case["segment"], "exact": exact_match(pred, case["expected"]), "judge": llm_judge(case, pred)["score"], }) return results8.5 Regression tests in CI
Run the golden set in CI on every prompt/model/config change and fail the build on regression. Pin the model and temperature: 0 for stable comparisons.
name: eval-regressionon: pull_request: paths: ["prompts/**", "src/**", "evals/**"]jobs: eval: runs-on: ubuntu-latest steps: - uses: actions/checkout@v4 - uses: actions/setup-python@v5 with: python-version: "3.12" - run: pip install anthropic - name: Run golden-set eval env: ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }} run: python evals/run.py --golden evals/golden.jsonl --min-score 0.95 # run.py exits non-zero if aggregate OR any per-segment score drops below threshold8.6 Per-segment metrics
Report metrics per segment (per document type, per language, per intent). Aggregate accuracy masks per-segment failures (anti-pattern #10) — the single most common eval trap on the exam.
| Segment | Cases | Accuracy | Verdict |
|---|---|---|---|
| Overall (aggregate) | 1,000 | 92% | Looks fine — misleading |
| invoices | 600 | 98% | Healthy |
| receipts | 300 | 95% | Healthy |
| handwritten | 100 | 61% | The real problem, hidden by the aggregate |
The 92% aggregate passes a naive gate, yet handwritten forms fail 4 in 10. Gate CI on the worst segment, not the mean, and always break metrics down before shipping.
8.7 Choosing the right eval method for the output type
Not every output is scored the same way. Picking the wrong scoring method is a subtle exam trap.
| Output type | Best scoring method | Why |
|---|---|---|
| Labels / classifications | Exact match (accuracy, F1 per class) | Deterministic ground truth |
| Structured extraction | Field-level match against expected JSON | Partial credit and per-field diagnosis |
| Open-ended prose | Rubric + LLM-as-judge (separate session) | No single correct string |
| Comparative quality | Pairwise A/B (which of two is better) | Humans/judges compare more reliably than absolute-score |
| Retrieval/grounding | Citation/faithfulness checks | Catches ungrounded claims |
# Pairwise A/B: ask a separate judge which answer is better, randomising order to avoid position bias.import random
def pairwise(judge_client, question, ans_a, ans_b): first, second = ("A", ans_a, "B", ans_b) if random.random() < 0.5 else ("B", ans_b, "A", ans_a) label1, text1, label2, text2 = first r = judge_client.messages.create( model="claude-opus-5", max_tokens=128, temperature=0, system="Reply with the label of the better answer: 'A' or 'B'. JSON: {'winner': 'A'|'B'}.", messages=[{"role": "user", "content": f"Q:{question}\n\n{label1}:\n{text1}\n\n{label2}:\n{text2}"}]) import json return json.loads("".join(b.text for b in r.content if b.type == "text"))["winner"]Exam signal
‘Deterministic labels’ → exact match. ‘Open-ended quality’ → rubric + LLM-as-judge in a separate session. ‘Which version is better’ → pairwise A/B. ‘Grounding’ → citation/faithfulness. Match the method to the output type.
8.8 Debugging decision tree: which layer, which fix
A fast triage flow turns a vague ‘it’s broken’ into a targeted fix. The HTTP status and stop_reason are your first two signals.
Failure observed │ ├─ HTTP 4xx/5xx? ──► INTEGRATION layer │ ├─ 429 / 500 / 529 ──► transient: backoff + jitter, honour retry-after │ ├─ 400 invalid_request ──► fix request shape (e.g. forced tool on Fable 5.1) │ ├─ 401 / 403 ──► auth/permissions: fix key/org/scopes (do NOT retry) │ └─ timeout / JSON decode ──► timeouts/streaming; schema + validation-retry │ └─ HTTP 200 but bad answer? ──► MODEL-OUTPUT layer ├─ stop_reason == max_tokens ──► truncation: continue / raise cap / chunk ├─ stop_reason == refusal ──► safety decline: reframe / policy path ├─ off-schema-but-close ──► strict/structured outputs + validation-retry └─ confident but wrong ──► grounding: docs-first, citations, RAG, claim validation| Symptom | Layer | Wrong fix (distractor) | Right fix |
|---|---|---|---|
429 under load | Integration | Rewrite the prompt | Backoff + retry-after; batch/queue |
400 on forced tool (Fable 5.1) | Integration | Retry with backoff | auto/structured outputs |
Truncated JSON, max_tokens | Model output | Rotate API key | Raise cap / continue / chunk |
| Fabricated value | Model output | Raise temperature | Grounding + claim validation |
Exam signal
Read the status code and stop_reason first. A 429/timeout/JSON error is integration; a hallucination/off-schema/refusal is model output. Applying an integration fix to a model-output problem (or vice-versa) is the classic distractor.
8.9 Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| Any failure can be retried | Only transient integration errors (429/5xx/529) are retryable | Wrong-layer fix |
A 400 might be transient | It is deterministic; fix the request | Distinguishes layers |
| Aggregate accuracy is enough | It masks per-segment failures (anti-pattern #10) | The top eval trap |
| Same-session self-review is efficient | The judge inherits the generator’s bias (anti-pattern #9) | Judge-design trap |
temperature: 0 gives reproducible bytes | It reduces variance; pin the snapshot too | Reproducibility nuance |
A max_tokens end is the root cause of a loop | It is a downstream symptom; fix loop/stop_reason handling | Trace-reading trap |
| Empty output means ‘no results’ | It may be silently suppressed error (anti-pattern #7) | Silent-failure trap |
| Exact match fits every output | Open-ended output needs a rubric/LLM-judge; comparisons need pairwise | Eval-method selection |
8.10 Scenario walkthrough: an eval that lied
Scenario. A document-classification service reports 93% accuracy on its golden set and passes CI, yet production incidents keep involving handwritten forms and non-English documents. Investigation reveals: (1) CI gates only on aggregate accuracy; (2) the eval had the generator model grade its own open-ended rationales in the same session; (3) an intermittent 400 invalid_request in production is being retried with exponential backoff and never succeeding; and (4) a helper returns an empty list when a call throws, which the caller treats as ‘no matches’. Fix the eval and the debugging.
Expert reasoning trace.
- Break the aggregate. 93% overall hides that handwritten and non-English segments fail badly (anti-pattern #10). Report per-segment metrics (by document type and language) and gate CI on the worst segment, not the mean.
- Fix the judge. Same-session self-grading (anti-pattern #9) rubber-stamps the generator. Move to an LLM-as-judge in a separate session, ideally a different model, against an explicit rubric; for the open-ended rationales, consider pairwise A/B with randomised order.
- Fix the
400triage. A400 invalid_requestis an integration/schema error and deterministic — retrying with backoff loops forever. Read the error: likely a forbidden parameter (e.g. forcedtool_choiceon Fable 5.1). Fix the request shape (auto/structured outputs); do not retry. - Fix the silent suppression. Returning an empty list on exception (anti-pattern #7) reports failure as success. Propagate the error with diagnostic context, or fail loudly; never let empty-as-success flow downstream.
- Make it reproducible. Pin the model snapshot and set
temperature: 0for the eval; logrequest_id,stop_reason,usageand latency so incidents are traceable. - Reject the tempting alternatives. ‘Raise the aggregate threshold to 95%’ — still masks per-segment failures. ‘Retry the 400 more aggressively’ — deterministic error. ‘Trust the empty result as no matches’ — silent suppression. ‘Have the same model re-grade to save cost’ — reintroduces bias.
Correct decision. Per-segment metrics with worst-segment CI gating; a separate-session/different-model LLM-judge (pairwise A/B for open-ended); fix the 400 at the request layer (no backoff); surface the suppressed error loudly; pin snapshot + temperature: 0 and log request IDs for traceability.
Exam traps in this domain
| Trap | Why it is wrong |
|---|---|
| Retrying a model-output error like a network error | Wrong layer; fix the prompt/eval, not backoff |
Treating a 400 schema error as transient | It is deterministic; fix the request shape, do not retry |
| Same-session self-review | Anti-pattern #9; use a separate judge session/model |
| Reporting only aggregate accuracy | Anti-pattern #10; masks per-segment failures |
| Not logging request IDs | No trace handle for debugging/support |
| Reproducing with random temperature and unpinned model | Not reproducible; pin snapshot + temperature 0 |
| Treating empty output as success | Silent suppression (anti-pattern #7) |
| No CI regression suite | Prompt/model changes silently regress |
| Judging with the same model in the same context | Bias; use a different session/model |
Reading max_tokens truncation as the root cause of a loop | It is a downstream symptom; fix the loop/stop_reason handling |
| Raising the aggregate threshold instead of gating per-segment | Still masks the failing segment; gate on the worst segment |
| Using exact match for open-ended output | Use a rubric + LLM-as-judge (separate session); pairwise A/B for comparisons |
Not reading the status code and stop_reason before choosing a fix | They tell you the layer; the wrong-layer fix is the classic distractor |
| Ignoring position bias in pairwise judging | Randomise A/B order so the judge is not swayed by position |
Treating a deterministic 400 as flaky and retrying | Fix the request shape; backoff will loop forever |
Practice questions
Q1 · A service intermittently fails with JSON decode errors and 429s, and separately sometimes extracts the wrong invoice total. How should the team triage? (Select one)
A. Treat both as model problems and rewrite the prompt.
B. Separate layers: fix 429s with backoff and JSON errors with validation-retry (integration); fix wrong extraction with prompt/context engineering and validation (model output).
C. Treat both as network problems and add retries.
D. Increase max_tokens for both.
Answer: B. The failures are in different layers and need different fixes. Blanket prompt rewrites (A) ignore the integration errors; retries (C) do nothing for wrong extraction; raising max_tokens (D) addresses neither a 429 nor a wrong value.
Q2 · A model scores 91% overall on the eval, but a production incident involves handwritten forms. What eval practice would have surfaced this, and how should output be judged? (Select two)
A. Report per-segment metrics (by document type) instead of only aggregate. B. Have the same session grade its own output. C. Use an LLM-as-judge in a separate session/model against a rubric. D. Only track overall accuracy. E. Skip evals in CI.
Answer: A and C. Per-segment metrics expose the handwritten-forms failure (anti-pattern #10), and a separate-session judge avoids self-review bias (anti-pattern #9). Aggregate-only (D), same-session grading (B) and no CI (E) are the anti-patterns.
Q3 · Every request to Claude Fable 5.1 returns `400 invalid_request`; the code sets `tool_choice: {'type':'any'}`. Which layer is this and what is the fix? (Select one)
A. Model-output error; rewrite the prompt to be clearer.
B. Integration/schema error; Fable 5.1 rejects forced any/{type:tool}, so use auto with an instruction, strict:true, or structured outputs.
C. Rate-limit error; add exponential backoff.
D. Transient error; retry with jitter.
Answer: B. A 400 on a forbidden parameter is a deterministic integration/schema error specific to Fable 5.1’s tool-choice restriction; the fix is to change the request. It is not about prompt quality (A); it is not a 429 (C); and retrying a deterministic 400 (D) will always fail again.
Q4 · A trace shows an agent calling the same tool with identical input on every step, with input tokens climbing each turn, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)
A. The model ran out of output tokens; raise max_tokens.
B. The harness is not feeding the tool result back and/or termination is not driven by stop_reason; fix the loop to append results and stop on end_turn.
C. Rate limiting; add backoff.
D. A hallucination; add citations.
Answer: B. Repeating the same call with growing context is a loop-control defect (anti-patterns #1/#2). The final max_tokens is a downstream symptom, not the cause (A). It is not a transport (C) or grounding (D) issue.
Q5 · A response comes back with `stop_reason: 'max_tokens'` and the JSON is cut off mid-object. Which layer, and what are TWO valid recoveries? (Select two)
A. Model-output truncation; raise max_tokens for the call.
B. Chunk the task or stream and continue generation.
C. It is an auth error; rotate the API key.
D. Silently return the partial JSON as success.
E. Retry unchanged with backoff.
Answer: A and B. max_tokens means the output was truncated; raising the limit or splitting/streaming the work recovers it. It is not auth (C); returning partial output as success is silent suppression, anti-pattern #7 (D); retrying unchanged (E) truncates again.
Q6 · To reproduce a model-output bug reliably, which combination should the team use? (Select one)
A. Latest floating model alias, temperature 1, paraphrased prompt.
B. Pinned model snapshot, temperature: 0, frozen system prompt and tools, exact message replay.
C. Any model, as long as the prompt is similar.
D. A different model each run to average out noise.
Answer: B. Controlling the model snapshot, temperature, prompt, tools, and message history maximises repeatability. Floating aliases and non-zero temperature (A), loose prompts (C), and varying the model (D) all inject variance that defeats reproduction.
Q7 · An eval pipeline has the generator model grade its own answers in the same conversation. What is wrong, and what is the fix? (Select one)
A. Nothing; self-grading is efficient. B. Same-session self-review retains the reasoning bias that produced the answer; run an LLM-as-judge in a separate session, ideally a different model. C. Use exact match for everything instead. D. Grade only the aggregate.
Answer: B. Same-session self-review (anti-pattern #9) is biased because the judge shares the generator’s context. A separate-session, different-model judge removes that bias. Exact match (C) does not fit open-ended output; aggregate-only grading (D) is a different anti-pattern.
Q8 · Which artefacts should be logged for every Claude call to enable trace analysis? (Select two)
A. request_id and a correlation ID.
B. The user’s password.
C. model, stop_reason, usage tokens, latency, and (for agents) the tool sequence.
D. Nothing, to save storage.
E. Only the final answer text.
Answer: A and C. Request/correlation IDs plus model, stop reason, token usage, latency, and tool sequence give a full trace handle for debugging and support. Logging secrets (B) is a security violation; logging nothing (D) or only the answer (E) leaves you blind during incidents.
Q9 · A CI job runs the golden set but only fails the build if aggregate accuracy drops. A per-segment regression on 'legal documents' ships to production. What should CI do instead? (Select one)
A. Keep gating on aggregate; it is simpler. B. Gate on the worst-performing segment as well as the aggregate, failing if any segment drops below its threshold. C. Remove the eval from CI. D. Only run evals monthly.
Answer: B. Gating on per-segment thresholds catches the failures an aggregate hides (anti-pattern #10). Aggregate-only gating (A) is exactly what let the regression through; removing (C) or slowing (D) evals makes it worse.
Q10 · A function returns an empty list when the Claude call raises an exception, and the caller treats empty as 'no results found'. Which anti-pattern is this and how is it fixed? (Select one)
A. Aggregate metrics; add per-segment reporting.
B. Silent error suppression (anti-pattern #7); surface the error with diagnostic context instead of returning empty-as-success.
C. Same-session self-review; use a separate judge.
D. Iteration cap; drive from stop_reason.
Answer: B. Returning empty on failure hides errors and reports failure as success — silent suppression (anti-pattern #7). The fix is to propagate the error with context. The other options name unrelated anti-patterns.
Q11 · Output is almost valid JSON but occasionally emits an extra trailing field not in the schema. Which layer, and what is the most robust fix? (Select one)
A. Integration timeout; add retries.
B. Model-output format drift; enforce the schema with strict: true on the tool or structured outputs, plus a validation-retry.
C. Rate limit; add backoff.
D. Auth error; rotate the key.
Answer: B. Off-schema-but-close output is format drift, a model-output problem; strict:true/structured outputs plus validation-retry enforce the schema. It is not a timeout (A), rate limit (C), or auth (D) issue.
Q12 · A team wants to prevent prompt and model changes from silently degrading quality. Which practice is essential? (Select one)
A. Manual spot-checks whenever someone remembers. B. A regression suite over a golden set run automatically in CI on every prompt/model/config change, with per-segment thresholds that fail the build. C. Trusting the model’s self-reported confidence. D. Only measuring latency.
Answer: B. Automated CI regression over a golden set with per-segment gating is the discipline that catches silent quality drops. Ad-hoc manual checks (A) are unreliable; self-reported confidence (C) is an anti-pattern; latency (D) does not measure quality.
Q13 · An eval scores open-ended rationales. Which scoring method is appropriate, and which is not? (Select one)
A. Exact string match against a single reference answer. B. A rubric with an LLM-as-judge in a separate session (and pairwise A/B for comparisons), not exact match. C. Aggregate accuracy only. D. The generator grading itself for speed.
Answer: B. Open-ended output has no single correct string, so use a rubric + separate-session judge, with pairwise A/B for comparisons. Exact match (A) fits deterministic labels, not prose; aggregate-only (C) masks segments; self-grading (D) is anti-pattern #9.
Q14 · A production call intermittently returns `400 invalid_request` and is retried with exponential backoff, never succeeding. What is the correct triage? (Select one)
A. Add more retries with longer backoff.
B. Recognise a 400 is a deterministic integration/schema error (e.g. a forbidden parameter such as forced tool_choice on Fable 5.1); fix the request shape rather than retrying.
C. Treat it as a model hallucination and rewrite the prompt.
D. Rotate the API key.
Answer: B. A 400 will never succeed on retry; read it and fix the request. More backoff (A) loops forever; it is not a model-output problem (C); auth rotation (D) addresses 401/403, not a 400.
Q15 · A trace shows an agent repeating the same tool call with growing input tokens, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)
A. The model ran out of output tokens; raise max_tokens.
B. A loop-control defect: the harness is not feeding the tool result back and/or termination is not driven by stop_reason; fix the loop. The final max_tokens is a downstream symptom.
C. Rate limiting; add backoff.
D. A hallucination; add citations.
Answer: B. Repeating the same call with growing context is a loop-control problem (anti-patterns #1/#2); the max_tokens end is a symptom, not the cause (A). It is not transport (C) or grounding (D).
Q16 · A CI eval gates only on aggregate accuracy (93%), and a handwritten-forms regression ships. What TWO changes fix the eval process? (Select two)
A. Report and gate on per-segment metrics (by document type), failing if any segment drops below its threshold. B. Use an LLM-as-judge in a separate session/model against a rubric for the open-ended parts. C. Raise the aggregate threshold to 96%. D. Have the same session grade its own output. E. Run evals only monthly.
Answer: A and B. Per-segment gating surfaces the hidden failure (anti-pattern #10) and a separate-session judge avoids self-review bias (anti-pattern #9). A higher aggregate (C) still masks segments; self-grading (D) is #9; monthly evals (E) slow detection.
Q17 · A helper returns an empty list when the Claude call throws, and the caller treats empty as 'no matches'. Which anti-pattern is this and the fix? (Select one)
A. Aggregate metrics; add per-segment reporting.
B. Silent error suppression (anti-pattern #7); propagate the error with diagnostic context instead of returning empty-as-success.
C. Same-session self-review; use a separate judge.
D. Iteration cap; drive from stop_reason.
Answer: B. Returning empty on failure reports failure as success — silent suppression (anti-pattern #7). Surface the error with context. The others name unrelated anti-patterns.
Q18 · When comparing two prompt versions with an LLM judge, results flip depending on which answer is shown first. What is the issue and the fix? (Select one)
A. The judge model is broken; switch models. B. Position bias in pairwise judging; randomise the A/B order (and optionally average both orders) so position does not decide the winner. C. Temperature is too low; raise it. D. Use exact match instead.
Answer: B. Pairwise judges can be swayed by answer position; randomising order (or scoring both orders) removes the bias. It is not a broken model (A); temperature (C) is not the cause; exact match (D) does not fit open-ended comparison.
Key takeaways
- Triage by layer: integration errors (429/5xx/timeouts/JSON/roles) get retries and request fixes; model-output errors get prompt/context/model changes and validation.
- Log request IDs, model,
stop_reason,usageand latency for trace analysis. - Reproduce with a pinned snapshot and
temperature: 0, accepting residual non-determinism. - Evaluate with a golden set and rubric; use LLM-as-judge in a separate session (never same-session self-review).
- Run regression tests in CI and report per-segment metrics – aggregates hide the failures that matter.
- Match the eval method to the output: exact match for labels, field-level match for extraction, rubric + separate-session LLM-judge for prose, pairwise A/B (with randomised order) for comparisons.
- Triage with the status code and
stop_reasonfirst:429/timeout/JSON = integration; hallucination/off-schema/refusal = model output — the wrong-layer fix is the classic distractor. - Gate CI on the worst-performing segment, not the aggregate; raising the aggregate threshold still masks a failing segment.
- A deterministic
400never succeeds on retry — fix the request shape (e.g. forced tool choice on Fable 5.1) rather than adding backoff.
Last updated Sep 18, 2026