AI Cert Prep
Type to search documentation.

Domains

D8 · Eval, Testing and Debugging

Identifying integration-layer vs model-output errors, recovery, trace analysis, request-ID logging, reproducibility, eval basics, regression tests in CI and per-segment metrics.

This domain is roughly 1 of 53 items, but it ties the others together. It tests whether you can tell an integration-layer failure from a model-output failure, reproduce and trace problems, and evaluate quality without fooling yourself. The theme: measure per-segment, reproduce deterministically, and judge output in a separate session.

Learning objectives

By the end of this page you should be able to:

  1. Distinguish integration-layer from model-output errors and choose recovery.
  2. Do trace analysis and log request IDs.
  3. Reproduce issues with a pinned model and temperature: 0 where appropriate.
  4. Apply eval basics: golden set, rubric, LLM-as-judge in a separate session.
  5. Add regression tests in CI and use per-segment metrics.

8.1 Integration-layer vs model-output errors

The first question in any incident: which layer failed? Integration-layer errors live in the plumbing (auth, transport, request shape, rate limits, parsing). Model-output errors live in what Claude produced (hallucination, format drift, refusal, truncation). Retries and timeouts fix the former; prompt/context/eval changes fix the latter. Applying the wrong fix to the wrong layer is the classic distractor.

Error taxonomy

LayerErrorDiagnostic signalRecovery
IntegrationAuth401 authentication / 403 permissionCheck API key, org, scopes; do not retry blindly
IntegrationSchema / bad request400 invalid_request (e.g. tool_choice:"any" on Fable 5.1)Fix request shape; use auto + strict:true / structured outputs
IntegrationTimeoutClient timeout, no responseSet sane timeouts; retry with backoff; consider streaming
IntegrationRate limit429 rate_limit, retry-after headerExponential backoff + jitter; respect retry-after; batch/queue
IntegrationParsingJSON decode error on the response bodyValidate/repair; request structured outputs; validation-retry
Model outputHallucinationFluent claim with no grounding in inputsGrounding, citations, RAG, validation; verify claims
Model outputFormat driftOutput nearly-valid but off-schemastrict:true tool schema / structured outputs; validation-retry
Model outputRefusalstop_reason: "refusal"Reframe legitimate request; adjust system prompt; escalate
Model outputTruncationstop_reason: "max_tokens"Raise max_tokens; chunk the task; stream and continue

Exam signal

First ask which layer. A 429, a timeout, or a JSON decode error is integration — fix with backoff, timeouts, or request/schema changes. A hallucination, off-schema output, or refusal is model output — fix with prompt/context/model changes and validation. stop_reason and the HTTP status code tell you which layer.


8.2 Trace analysis walkthrough

Log the request-id and a correlation ID for every call, plus model, stop_reason, usage and latency. For agents, log the tool sequence. Traces let you pinpoint where a multi-step run went wrong and give Anthropic support a handle.

python
resp = client.messages.create(model=MODEL, max_tokens=512, messages=msgs)
log.info("claude_call", extra={
"request_id": resp._request_id, "model": resp.model,
"stop_reason": resp.stop_reason, "in": resp.usage.input_tokens,
"out": resp.usage.output_tokens})

Reading a trace of a stuck agent:

text
corr_id=abc-123
step 1 request_id=req_01A stop_reason=tool_use tools=[search_orders] in=1,240 out=180
step 2 request_id=req_01B stop_reason=tool_use tools=[get_order(id=9981)] in=1,910 out=95
step 3 request_id=req_01C stop_reason=tool_use tools=[get_order(id=9981)] in=2,600 out=95 ← repeat
step 4 request_id=req_01D stop_reason=tool_use tools=[get_order(id=9981)] in=3,290 out=95 ← loop
...
step N request_id=req_01Z stop_reason=max_tokens in=9,900 out=512 ← truncated

Diagnosis: the loop repeats the same tool call and input tokens climb each turn — the harness is not feeding the tool result back, or termination is driven by something other than stop_reason. This is anti-patterns #1/#2 (natural-language termination and iteration caps) rather than a model defect. The max_tokens at the end is a downstream symptom, not the cause. Fix the loop: after each tool_use, append the tool result and continue until stop_reason is end_turn.


8.3 Reproducibility

To reproduce a model-output issue, control every variable you can:

  1. Pin the exact model snapshot — not a floating alias.
  2. Set temperature: 0 where appropriate to minimise variance.
  3. Freeze the system prompt and tools and replay the same messages in order.
python
resp = client.messages.create(
model="claude-sonnet-5", # pinned snapshot
temperature=0, # minimise variance
max_tokens=512,
system=FROZEN_SYSTEM_PROMPT,
tools=FROZEN_TOOLS,
messages=RECORDED_MESSAGES, # exact replay
)

Non-determinism means the output still is not byte-identical, but pinning + temperature: 0 + a frozen prompt makes issues far more repeatable and comparisons fair. Note: Fable 5.x / Opus 5 / Sonnet 5 harnesses must be append-only — editing or reordering earlier turns invalidates later thinking blocks, so replay by appending, never by rewriting history.


8.4 Eval basics and an eval harness

ElementWhat it is
Golden setCurated inputs with known-good outputs
Exact matchFor deterministic outputs (labels, extractions)
RubricExplicit scoring criteria for open-ended output
LLM-as-judgeA different model/session scores output against the rubric

Same-session self-review

Never have the model grade its own output in the same session — it retains the reasoning bias that produced it (anti-pattern #9). Use a separate session and ideally a different model as judge.

A minimal harness combining exact-match and an LLM-as-judge in a separate call:

python
import json
from anthropic import Anthropic
client = Anthropic()
GEN_MODEL = "claude-sonnet-5" # system under test
JUDGE_MODEL = "claude-opus-5" # different model, separate call
def generate(case):
r = client.messages.create(
model=GEN_MODEL, temperature=0, max_tokens=512,
system="Extract the invoice total as JSON: {'total_cents': int}.",
messages=[{"role": "user", "content": case["input"]}],
)
return r.content[0].text
def exact_match(pred, expected):
try:
return json.loads(pred).get("total_cents") == expected["total_cents"]
except json.JSONDecodeError:
return False # parsing failure counts as a miss, never silently passed
RUBRIC = (
"Score 1 if the answer is faithful to the source and correctly formatted, "
"else 0. Return JSON: {'score': 0 or 1, 'reason': str}."
)
def llm_judge(case, pred):
# Separate session/model — no shared context with the generator.
r = client.messages.create(
model=JUDGE_MODEL, temperature=0, max_tokens=256,
system=RUBRIC,
messages=[{"role": "user",
"content": f"SOURCE:\n{case['input']}\n\nANSWER:\n{pred}"}],
)
return json.loads(r.content[0].text)
def run(golden):
results = []
for case in golden:
pred = generate(case)
results.append({
"id": case["id"], "segment": case["segment"],
"exact": exact_match(pred, case["expected"]),
"judge": llm_judge(case, pred)["score"],
})
return results

8.5 Regression tests in CI

Run the golden set in CI on every prompt/model/config change and fail the build on regression. Pin the model and temperature: 0 for stable comparisons.

yaml
name: eval-regression
on:
pull_request:
paths: ["prompts/**", "src/**", "evals/**"]
jobs:
eval:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install anthropic
- name: Run golden-set eval
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: python evals/run.py --golden evals/golden.jsonl --min-score 0.95
# run.py exits non-zero if aggregate OR any per-segment score drops below threshold

8.6 Per-segment metrics

Report metrics per segment (per document type, per language, per intent). Aggregate accuracy masks per-segment failures (anti-pattern #10) — the single most common eval trap on the exam.

SegmentCasesAccuracyVerdict
Overall (aggregate)1,00092%Looks fine — misleading
invoices60098%Healthy
receipts30095%Healthy
handwritten10061%The real problem, hidden by the aggregate

The 92% aggregate passes a naive gate, yet handwritten forms fail 4 in 10. Gate CI on the worst segment, not the mean, and always break metrics down before shipping.


8.7 Choosing the right eval method for the output type

Not every output is scored the same way. Picking the wrong scoring method is a subtle exam trap.

Output typeBest scoring methodWhy
Labels / classificationsExact match (accuracy, F1 per class)Deterministic ground truth
Structured extractionField-level match against expected JSONPartial credit and per-field diagnosis
Open-ended proseRubric + LLM-as-judge (separate session)No single correct string
Comparative qualityPairwise A/B (which of two is better)Humans/judges compare more reliably than absolute-score
Retrieval/groundingCitation/faithfulness checksCatches ungrounded claims
python
# Pairwise A/B: ask a separate judge which answer is better, randomising order to avoid position bias.
import random
def pairwise(judge_client, question, ans_a, ans_b):
first, second = ("A", ans_a, "B", ans_b) if random.random() < 0.5 else ("B", ans_b, "A", ans_a)
label1, text1, label2, text2 = first
r = judge_client.messages.create(
model="claude-opus-5", max_tokens=128, temperature=0,
system="Reply with the label of the better answer: 'A' or 'B'. JSON: {'winner': 'A'|'B'}.",
messages=[{"role": "user",
"content": f"Q:{question}\n\n{label1}:\n{text1}\n\n{label2}:\n{text2}"}])
import json
return json.loads("".join(b.text for b in r.content if b.type == "text"))["winner"]

Exam signal

‘Deterministic labels’ → exact match. ‘Open-ended quality’ → rubric + LLM-as-judge in a separate session. ‘Which version is better’ → pairwise A/B. ‘Grounding’ → citation/faithfulness. Match the method to the output type.


8.8 Debugging decision tree: which layer, which fix

A fast triage flow turns a vague ‘it’s broken’ into a targeted fix. The HTTP status and stop_reason are your first two signals.

text
Failure observed
│
├─ HTTP 4xx/5xx? ──► INTEGRATION layer
│ ├─ 429 / 500 / 529 ──► transient: backoff + jitter, honour retry-after
│ ├─ 400 invalid_request ──► fix request shape (e.g. forced tool on Fable 5.1)
│ ├─ 401 / 403 ──► auth/permissions: fix key/org/scopes (do NOT retry)
│ └─ timeout / JSON decode ──► timeouts/streaming; schema + validation-retry
│
└─ HTTP 200 but bad answer? ──► MODEL-OUTPUT layer
├─ stop_reason == max_tokens ──► truncation: continue / raise cap / chunk
├─ stop_reason == refusal ──► safety decline: reframe / policy path
├─ off-schema-but-close ──► strict/structured outputs + validation-retry
└─ confident but wrong ──► grounding: docs-first, citations, RAG, claim validation
SymptomLayerWrong fix (distractor)Right fix
429 under loadIntegrationRewrite the promptBackoff + retry-after; batch/queue
400 on forced tool (Fable 5.1)IntegrationRetry with backoffauto/structured outputs
Truncated JSON, max_tokensModel outputRotate API keyRaise cap / continue / chunk
Fabricated valueModel outputRaise temperatureGrounding + claim validation

Exam signal

Read the status code and stop_reason first. A 429/timeout/JSON error is integration; a hallucination/off-schema/refusal is model output. Applying an integration fix to a model-output problem (or vice-versa) is the classic distractor.


8.9 Common misconceptions

MisconceptionRealityWhy it matters on the exam
Any failure can be retriedOnly transient integration errors (429/5xx/529) are retryableWrong-layer fix
A 400 might be transientIt is deterministic; fix the requestDistinguishes layers
Aggregate accuracy is enoughIt masks per-segment failures (anti-pattern #10)The top eval trap
Same-session self-review is efficientThe judge inherits the generator’s bias (anti-pattern #9)Judge-design trap
temperature: 0 gives reproducible bytesIt reduces variance; pin the snapshot tooReproducibility nuance
A max_tokens end is the root cause of a loopIt is a downstream symptom; fix loop/stop_reason handlingTrace-reading trap
Empty output means ‘no results’It may be silently suppressed error (anti-pattern #7)Silent-failure trap
Exact match fits every outputOpen-ended output needs a rubric/LLM-judge; comparisons need pairwiseEval-method selection

8.10 Scenario walkthrough: an eval that lied

Scenario. A document-classification service reports 93% accuracy on its golden set and passes CI, yet production incidents keep involving handwritten forms and non-English documents. Investigation reveals: (1) CI gates only on aggregate accuracy; (2) the eval had the generator model grade its own open-ended rationales in the same session; (3) an intermittent 400 invalid_request in production is being retried with exponential backoff and never succeeding; and (4) a helper returns an empty list when a call throws, which the caller treats as ‘no matches’. Fix the eval and the debugging.

Expert reasoning trace.

  1. Break the aggregate. 93% overall hides that handwritten and non-English segments fail badly (anti-pattern #10). Report per-segment metrics (by document type and language) and gate CI on the worst segment, not the mean.
  2. Fix the judge. Same-session self-grading (anti-pattern #9) rubber-stamps the generator. Move to an LLM-as-judge in a separate session, ideally a different model, against an explicit rubric; for the open-ended rationales, consider pairwise A/B with randomised order.
  3. Fix the 400 triage. A 400 invalid_request is an integration/schema error and deterministic — retrying with backoff loops forever. Read the error: likely a forbidden parameter (e.g. forced tool_choice on Fable 5.1). Fix the request shape (auto/structured outputs); do not retry.
  4. Fix the silent suppression. Returning an empty list on exception (anti-pattern #7) reports failure as success. Propagate the error with diagnostic context, or fail loudly; never let empty-as-success flow downstream.
  5. Make it reproducible. Pin the model snapshot and set temperature: 0 for the eval; log request_id, stop_reason, usage and latency so incidents are traceable.
  6. Reject the tempting alternatives. ‘Raise the aggregate threshold to 95%’ — still masks per-segment failures. ‘Retry the 400 more aggressively’ — deterministic error. ‘Trust the empty result as no matches’ — silent suppression. ‘Have the same model re-grade to save cost’ — reintroduces bias.

Correct decision. Per-segment metrics with worst-segment CI gating; a separate-session/different-model LLM-judge (pairwise A/B for open-ended); fix the 400 at the request layer (no backoff); surface the suppressed error loudly; pin snapshot + temperature: 0 and log request IDs for traceability.


Exam traps in this domain

TrapWhy it is wrong
Retrying a model-output error like a network errorWrong layer; fix the prompt/eval, not backoff
Treating a 400 schema error as transientIt is deterministic; fix the request shape, do not retry
Same-session self-reviewAnti-pattern #9; use a separate judge session/model
Reporting only aggregate accuracyAnti-pattern #10; masks per-segment failures
Not logging request IDsNo trace handle for debugging/support
Reproducing with random temperature and unpinned modelNot reproducible; pin snapshot + temperature 0
Treating empty output as successSilent suppression (anti-pattern #7)
No CI regression suitePrompt/model changes silently regress
Judging with the same model in the same contextBias; use a different session/model
Reading max_tokens truncation as the root cause of a loopIt is a downstream symptom; fix the loop/stop_reason handling
Raising the aggregate threshold instead of gating per-segmentStill masks the failing segment; gate on the worst segment
Using exact match for open-ended outputUse a rubric + LLM-as-judge (separate session); pairwise A/B for comparisons
Not reading the status code and stop_reason before choosing a fixThey tell you the layer; the wrong-layer fix is the classic distractor
Ignoring position bias in pairwise judgingRandomise A/B order so the judge is not swayed by position
Treating a deterministic 400 as flaky and retryingFix the request shape; backoff will loop forever

Practice questions

Q1 · A service intermittently fails with JSON decode errors and 429s, and separately sometimes extracts the wrong invoice total. How should the team triage? (Select one)

A. Treat both as model problems and rewrite the prompt. B. Separate layers: fix 429s with backoff and JSON errors with validation-retry (integration); fix wrong extraction with prompt/context engineering and validation (model output). C. Treat both as network problems and add retries. D. Increase max_tokens for both.

Answer: B. The failures are in different layers and need different fixes. Blanket prompt rewrites (A) ignore the integration errors; retries (C) do nothing for wrong extraction; raising max_tokens (D) addresses neither a 429 nor a wrong value.

Q2 · A model scores 91% overall on the eval, but a production incident involves handwritten forms. What eval practice would have surfaced this, and how should output be judged? (Select two)

A. Report per-segment metrics (by document type) instead of only aggregate. B. Have the same session grade its own output. C. Use an LLM-as-judge in a separate session/model against a rubric. D. Only track overall accuracy. E. Skip evals in CI.

Answer: A and C. Per-segment metrics expose the handwritten-forms failure (anti-pattern #10), and a separate-session judge avoids self-review bias (anti-pattern #9). Aggregate-only (D), same-session grading (B) and no CI (E) are the anti-patterns.

Q3 · Every request to Claude Fable 5.1 returns `400 invalid_request`; the code sets `tool_choice: {'type':'any'}`. Which layer is this and what is the fix? (Select one)

A. Model-output error; rewrite the prompt to be clearer. B. Integration/schema error; Fable 5.1 rejects forced any/{type:tool}, so use auto with an instruction, strict:true, or structured outputs. C. Rate-limit error; add exponential backoff. D. Transient error; retry with jitter.

Answer: B. A 400 on a forbidden parameter is a deterministic integration/schema error specific to Fable 5.1’s tool-choice restriction; the fix is to change the request. It is not about prompt quality (A); it is not a 429 (C); and retrying a deterministic 400 (D) will always fail again.

Q4 · A trace shows an agent calling the same tool with identical input on every step, with input tokens climbing each turn, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)

A. The model ran out of output tokens; raise max_tokens. B. The harness is not feeding the tool result back and/or termination is not driven by stop_reason; fix the loop to append results and stop on end_turn. C. Rate limiting; add backoff. D. A hallucination; add citations.

Answer: B. Repeating the same call with growing context is a loop-control defect (anti-patterns #1/#2). The final max_tokens is a downstream symptom, not the cause (A). It is not a transport (C) or grounding (D) issue.

Q5 · A response comes back with `stop_reason: 'max_tokens'` and the JSON is cut off mid-object. Which layer, and what are TWO valid recoveries? (Select two)

A. Model-output truncation; raise max_tokens for the call. B. Chunk the task or stream and continue generation. C. It is an auth error; rotate the API key. D. Silently return the partial JSON as success. E. Retry unchanged with backoff.

Answer: A and B. max_tokens means the output was truncated; raising the limit or splitting/streaming the work recovers it. It is not auth (C); returning partial output as success is silent suppression, anti-pattern #7 (D); retrying unchanged (E) truncates again.

Q6 · To reproduce a model-output bug reliably, which combination should the team use? (Select one)

A. Latest floating model alias, temperature 1, paraphrased prompt. B. Pinned model snapshot, temperature: 0, frozen system prompt and tools, exact message replay. C. Any model, as long as the prompt is similar. D. A different model each run to average out noise.

Answer: B. Controlling the model snapshot, temperature, prompt, tools, and message history maximises repeatability. Floating aliases and non-zero temperature (A), loose prompts (C), and varying the model (D) all inject variance that defeats reproduction.

Q7 · An eval pipeline has the generator model grade its own answers in the same conversation. What is wrong, and what is the fix? (Select one)

A. Nothing; self-grading is efficient. B. Same-session self-review retains the reasoning bias that produced the answer; run an LLM-as-judge in a separate session, ideally a different model. C. Use exact match for everything instead. D. Grade only the aggregate.

Answer: B. Same-session self-review (anti-pattern #9) is biased because the judge shares the generator’s context. A separate-session, different-model judge removes that bias. Exact match (C) does not fit open-ended output; aggregate-only grading (D) is a different anti-pattern.

Q8 · Which artefacts should be logged for every Claude call to enable trace analysis? (Select two)

A. request_id and a correlation ID. B. The user’s password. C. model, stop_reason, usage tokens, latency, and (for agents) the tool sequence. D. Nothing, to save storage. E. Only the final answer text.

Answer: A and C. Request/correlation IDs plus model, stop reason, token usage, latency, and tool sequence give a full trace handle for debugging and support. Logging secrets (B) is a security violation; logging nothing (D) or only the answer (E) leaves you blind during incidents.

Q9 · A CI job runs the golden set but only fails the build if aggregate accuracy drops. A per-segment regression on 'legal documents' ships to production. What should CI do instead? (Select one)

A. Keep gating on aggregate; it is simpler. B. Gate on the worst-performing segment as well as the aggregate, failing if any segment drops below its threshold. C. Remove the eval from CI. D. Only run evals monthly.

Answer: B. Gating on per-segment thresholds catches the failures an aggregate hides (anti-pattern #10). Aggregate-only gating (A) is exactly what let the regression through; removing (C) or slowing (D) evals makes it worse.

Q10 · A function returns an empty list when the Claude call raises an exception, and the caller treats empty as 'no results found'. Which anti-pattern is this and how is it fixed? (Select one)

A. Aggregate metrics; add per-segment reporting. B. Silent error suppression (anti-pattern #7); surface the error with diagnostic context instead of returning empty-as-success. C. Same-session self-review; use a separate judge. D. Iteration cap; drive from stop_reason.

Answer: B. Returning empty on failure hides errors and reports failure as success — silent suppression (anti-pattern #7). The fix is to propagate the error with context. The other options name unrelated anti-patterns.

Q11 · Output is almost valid JSON but occasionally emits an extra trailing field not in the schema. Which layer, and what is the most robust fix? (Select one)

A. Integration timeout; add retries. B. Model-output format drift; enforce the schema with strict: true on the tool or structured outputs, plus a validation-retry. C. Rate limit; add backoff. D. Auth error; rotate the key.

Answer: B. Off-schema-but-close output is format drift, a model-output problem; strict:true/structured outputs plus validation-retry enforce the schema. It is not a timeout (A), rate limit (C), or auth (D) issue.

Q12 · A team wants to prevent prompt and model changes from silently degrading quality. Which practice is essential? (Select one)

A. Manual spot-checks whenever someone remembers. B. A regression suite over a golden set run automatically in CI on every prompt/model/config change, with per-segment thresholds that fail the build. C. Trusting the model’s self-reported confidence. D. Only measuring latency.

Answer: B. Automated CI regression over a golden set with per-segment gating is the discipline that catches silent quality drops. Ad-hoc manual checks (A) are unreliable; self-reported confidence (C) is an anti-pattern; latency (D) does not measure quality.

Q13 · An eval scores open-ended rationales. Which scoring method is appropriate, and which is not? (Select one)

A. Exact string match against a single reference answer. B. A rubric with an LLM-as-judge in a separate session (and pairwise A/B for comparisons), not exact match. C. Aggregate accuracy only. D. The generator grading itself for speed.

Answer: B. Open-ended output has no single correct string, so use a rubric + separate-session judge, with pairwise A/B for comparisons. Exact match (A) fits deterministic labels, not prose; aggregate-only (C) masks segments; self-grading (D) is anti-pattern #9.

Q14 · A production call intermittently returns `400 invalid_request` and is retried with exponential backoff, never succeeding. What is the correct triage? (Select one)

A. Add more retries with longer backoff. B. Recognise a 400 is a deterministic integration/schema error (e.g. a forbidden parameter such as forced tool_choice on Fable 5.1); fix the request shape rather than retrying. C. Treat it as a model hallucination and rewrite the prompt. D. Rotate the API key.

Answer: B. A 400 will never succeed on retry; read it and fix the request. More backoff (A) loops forever; it is not a model-output problem (C); auth rotation (D) addresses 401/403, not a 400.

Q15 · A trace shows an agent repeating the same tool call with growing input tokens, ending in `stop_reason: max_tokens`. What is the root cause? (Select one)

A. The model ran out of output tokens; raise max_tokens. B. A loop-control defect: the harness is not feeding the tool result back and/or termination is not driven by stop_reason; fix the loop. The final max_tokens is a downstream symptom. C. Rate limiting; add backoff. D. A hallucination; add citations.

Answer: B. Repeating the same call with growing context is a loop-control problem (anti-patterns #1/#2); the max_tokens end is a symptom, not the cause (A). It is not transport (C) or grounding (D).

Q16 · A CI eval gates only on aggregate accuracy (93%), and a handwritten-forms regression ships. What TWO changes fix the eval process? (Select two)

A. Report and gate on per-segment metrics (by document type), failing if any segment drops below its threshold. B. Use an LLM-as-judge in a separate session/model against a rubric for the open-ended parts. C. Raise the aggregate threshold to 96%. D. Have the same session grade its own output. E. Run evals only monthly.

Answer: A and B. Per-segment gating surfaces the hidden failure (anti-pattern #10) and a separate-session judge avoids self-review bias (anti-pattern #9). A higher aggregate (C) still masks segments; self-grading (D) is #9; monthly evals (E) slow detection.

Q17 · A helper returns an empty list when the Claude call throws, and the caller treats empty as 'no matches'. Which anti-pattern is this and the fix? (Select one)

A. Aggregate metrics; add per-segment reporting. B. Silent error suppression (anti-pattern #7); propagate the error with diagnostic context instead of returning empty-as-success. C. Same-session self-review; use a separate judge. D. Iteration cap; drive from stop_reason.

Answer: B. Returning empty on failure reports failure as success — silent suppression (anti-pattern #7). Surface the error with context. The others name unrelated anti-patterns.

Q18 · When comparing two prompt versions with an LLM judge, results flip depending on which answer is shown first. What is the issue and the fix? (Select one)

A. The judge model is broken; switch models. B. Position bias in pairwise judging; randomise the A/B order (and optionally average both orders) so position does not decide the winner. C. Temperature is too low; raise it. D. Use exact match instead.

Answer: B. Pairwise judges can be swayed by answer position; randomising order (or scoring both orders) removes the bias. It is not a broken model (A); temperature (C) is not the cause; exact match (D) does not fit open-ended comparison.

Key takeaways

  • Triage by layer: integration errors (429/5xx/timeouts/JSON/roles) get retries and request fixes; model-output errors get prompt/context/model changes and validation.
  • Log request IDs, model, stop_reason, usage and latency for trace analysis.
  • Reproduce with a pinned snapshot and temperature: 0, accepting residual non-determinism.
  • Evaluate with a golden set and rubric; use LLM-as-judge in a separate session (never same-session self-review).
  • Run regression tests in CI and report per-segment metrics – aggregates hide the failures that matter.
  • Match the eval method to the output: exact match for labels, field-level match for extraction, rubric + separate-session LLM-judge for prose, pairwise A/B (with randomised order) for comparisons.
  • Triage with the status code and stop_reason first: 429/timeout/JSON = integration; hallucination/off-schema/refusal = model output — the wrong-layer fix is the classic distractor.
  • Gate CI on the worst-performing segment, not the aggregate; raising the aggregate threshold still masks a failing segment.
  • A deterministic 400 never succeeds on retry — fix the request shape (e.g. forced tool choice on Fable 5.1) rather than adding backoff.

Last updated Sep 18, 2026