AI Cert Prep
Type to search documentation.

Domains

D4 · Prompt and Context Engineering

Prompt engineering principles, structured-output patterns, defensive parsing and validation-retry, and context engineering to prevent drift and bloat.

This domain is roughly 6 of 53 items. It tests whether you can write prompts that reliably produce the shape and quality you need, parse output defensively (never trusting confident text), and engineer context so long-running work does not drift or bloat. The core lesson: the prompt is the interface; the context is the working memory – engineer both.

Learning objectives

By the end of this page you should be able to:

  1. Apply core prompt-engineering principles: clear/direct instructions, XML tags, examples/few-shot, role, chain-of-thought, prefilling, output format.
  2. Produce structured output via output_config.format (JSON schema), tool-use-as-schema, and strict: true.
  3. Parse output defensively with validation-retry and appropriate skepticism.
  4. Prevent context drift and bloat: tool-output pruning, context editing, compaction, summarisation.
  5. Use context isolation via subagents and correct long-context ordering.
  6. Lay out prompts to be caching-aware.

4.1 Prompt-engineering principles

PrincipleWhat it doesExample lever
Clear and directRemoves ambiguityState the task, constraints, and output shape explicitly
XML tagsDelimits sections; separates instructions from data<document>…</document>, <instructions>…</instructions>
Examples / few-shotShows the desired pattern2–5 representative input→output pairs
RoleSets perspective and tonesystem: ‘You are a senior tax analyst.’
Chain-of-thoughtImproves reasoning on hard tasksAsk for step-by-step reasoning, or use thinking
PrefillingConstrains the start of the replyPrefill the assistant turn with { to force JSON
Output formatMakes output parseableSpecify exact schema / structure
python
messages = [
{"role": "user", "content": "Extract fields from <invoice>...</invoice> as JSON."},
{"role": "assistant", "content": "{"}, # prefill forces JSON start
]

Exam signal

“Output is inconsistent / hard to parse” → specify the format, use XML tags, add few-shot examples, and consider prefilling. “Reasoning is shallow on a hard task” → chain-of-thought / thinking.


4.2 Structured-output patterns

Three ways to get machine-readable output, strongest first:

  1. output_config.format with a JSON schema – the model is constrained to emit JSON matching the schema.
  2. Tool-use-as-schema – define a tool whose input_schema is your target shape; the tool_use block’s input is your structured data. Add strict: true to enforce the schema.
  3. Prompt + prefill + validation – ask for JSON, prefill {, then validate (weakest; needs defensive parsing).
json
{
"model": "claude-sonnet-5",
"max_tokens": 512,
"output_config": {
"format": {
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"total": {"type": "number"},
"currency": {"type": "string"},
"line_items": {"type": "array", "items": {"type": "string"}}
},
"required": ["total", "currency"]
}
}
},
"messages": [{"role": "user", "content": "Extract totals from the invoice."}]
}

Fable 5.1

Because Fable 5.1 rejects forced tool_choice, prefer output_config.format or strict: true schemas rather than forcing a tool to obtain structured output.


4.3 Defensive parsing and validation-retry

Never trust that output is valid just because it looks confident. Validate against the schema and retry with the error on failure.

python
import json
from jsonschema import validate, ValidationError
SCHEMA = {"type": "object", "required": ["total", "currency"],
"properties": {"total": {"type": "number"}, "currency": {"type": "string"}}}
def extract(client, messages, attempts=3):
for _ in range(attempts):
resp = client.messages.create(model="claude-sonnet-5", max_tokens=512, messages=messages)
text = "".join(b.text for b in resp.content if b.type == "text")
try:
data = json.loads(text)
validate(data, SCHEMA)
return data
except (json.JSONDecodeError, ValidationError) as e:
messages.append({"role": "assistant", "content": text})
messages.append({"role": "user", "content": f"Invalid output: {e}. Return valid JSON only."})
raise ValueError("could not obtain valid structured output")

Skepticism toward confident output

A fluent, confident answer is not a validated one. Parse defensively, validate against a schema, and surface failures – do not swallow them (anti-pattern #7). Do not rely on the model’s self-reported confidence (anti-pattern #4).


4.4 Context drift and bloat

As a conversation or agent runs, the context fills with tool outputs, intermediate reasoning and stale turns. This bloats cost and latency and causes drift (the model loses the thread or over-weights noise).

TechniqueWhat it does
Tool-output pruningDrop or truncate large/old tool results no longer needed
Context editingProgrammatically clear stale content (e.g., old tool results)
CompactionServer-side summarisation preserving the narrative
SummarisationReplace long history with a running summary
Context isolationPush subtasks to subagents with their own windows

Exam signal

“Long-running agent slows down / loses focus / costs balloon” → context editing, compaction, summarisation, and subagent isolation – not simply a bigger max_tokens.


4.5 Long-context ordering and caching-aware layout

For long inputs, put stable, large content first and the query last:

text
[ system: role + rules ] ← stable, cache here
[ tools ] ← stable, cache here
[ long documents ] ← stable, cache here
[ conversation history ] ← grows
[ the current question / task ] ← variable, last
  • Documents-first, query-last improves grounding and answer quality on long contexts.
  • The same layout is caching-friendly: the stable prefix is cached (0.1× on hits), and only the variable tail is reprocessed.

4.6 Long-context ordering: a worked example

On long inputs, place large stable content first and the query last. Two reasons: grounding quality improves when the model reads the evidence before the question, and the stable prefix becomes cacheable. Below, the documents and instructions come first; the user’s actual question is the final block.

json
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": [
{"type": "text", "text": "You answer strictly from the provided documents. If the answer is not present, say so."}
],
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "<documents>\n<doc id='1'>...50 pages of policy text...</doc>\n<doc id='2'>...contract text...</doc>\n</documents>"},
{"type": "text", "text": "<question>What is the termination notice period, and which clause states it?</question>"}
]}
]
}

The ordering is: system rules → documents → question. If you put the question first, the model reads it without evidence in view, weakening grounding, and the variable question would sit inside what you wanted to cache.

Exam signal

‘Answers to long-document questions are poorly grounded’ or ‘the question is at the top of a huge prompt’ → reorder to documents first, query last. This also sets up caching (next section).


4.7 Caching-aware prompt layout

Prompt caching rewards a stable prefix first. Put the content that repeats byte-for-byte across requests (system prompt, tools, long reference documents) at the front and mark the last stable block with cache_control. Everything variable (the current question, fresh turns) goes after the cached breakpoint.

json
{
"model": "claude-sonnet-5",
"system": [
{"type": "text", "text": "Long stable system instructions and schema...",
"cache_control": {"type": "ephemeral"}}
],
"tools": [
{"name": "search", "description": "...", "input_schema": {"type": "object"}}
],
"messages": [
{"role": "user", "content": [
{"type": "text", "text": "...long reference document reused every request...",
"cache_control": {"type": "ephemeral"}},
{"type": "text", "text": "Current question: ..."}
]}
]
}
  • Cache reads cost ≈ 0.1× base input; writes ≈ 1.25× (5-min TTL) or 2× (1-hour TTL).
  • The minimum cacheable prefix is ~1024 tokens (2048 on Haiku 4.5); shorter prefixes will not cache.
  • Order the layout: system → tools → stable documents → [cache breakpoint] → variable question.
text
[ system + schema ] ┐
[ tools ] │ stable prefix → mark cache_control here (cached at 0.1×)
[ reference documents ] ┘
────────────────────────────── cache breakpoint
[ current question ] variable tail → reprocessed each request

A prefix that changes never caches

If any byte of the prefix differs between requests (a timestamp, a per-user name spliced into the system prompt), the cache misses every time. Keep per-request variability out of the prefix and after the breakpoint.


4.8 Context editing vs compaction

Long agents outgrow the window. Two server-side controls manage it, and the exam expects you to pick the right one.

Context editingCompaction
What it doesProgrammatically clears/trims stale content (e.g., old tool results)Summarises older turns server-side, preserving the narrative
Keeps narrative?No — it removes contentYes — it condenses while retaining the thread
Best whenOld tool outputs are large and no longer neededThe full history matters but is too long to keep verbatim
RiskRemoving something still neededSummary loses a needed detail
json
{
"model": "claude-sonnet-5",
"context_management": {
"edits": [
{"type": "clear_tool_uses", "trigger": {"type": "input_tokens", "value": 100000}}
]
}
}
json
{
"model": "claude-sonnet-5",
"context_management": {
"compaction": {"trigger": {"type": "input_tokens", "value": 150000}}
}
}

Exam signal

‘Old tool results are bloating the window but the conversation thread must stay intact’ → compaction (summarise, preserve narrative). ‘Large stale tool outputs are no longer needed’ → context editing (clear them). Neither is ‘just raise max_tokens’, which only affects output length.


4.9 Few-shot example selection

Few-shot examples steer output shape and quality — but only if they are representative, diverse, and correct. Poor example selection can hurt more than it helps.

PrincipleWhy it matters
RepresentativeExamples should match the real input distribution, including the hard/edge cases
DiverseCover the range of categories/formats so the model does not overfit one pattern
Correct and consistentA wrong or inconsistent example teaches the wrong pattern; keep formatting identical
Ordered deliberatelyGroup by type or escalate difficulty; keep them before the query (stable prefix, cacheable)
Right number (2–5)Enough to show the pattern without bloating context; more is not always better
Dynamic selectionFor varied inputs, retrieve the few examples most similar to the current input

Exam signal

‘The model handles common cases but fails edge cases’ → add representative edge-case examples, not just more common ones. ‘Output overfits one format’ → increase example diversity. Place examples in the stable prefix so they stay cacheable.


4.10 Prefilling, stop sequences and chain-of-thought placement

Three fine-grained levers shape how the model produces output, and the exam tests when each helps.

LeverWhat it doesWhen to usePitfall
PrefillSeed the assistant turn’s start (e.g. {) to constrain formatForce JSON start, force a specific openingCannot prefill when thinking is on (thinking must come first)
stop_sequencesHalt generation at a markerCut off after a delimiter; bound outputThe marker text is not included in output
Chain-of-thought / thinkingReason step-by-step before answeringHard multi-step reasoningDo not ask for reasoning and strict JSON in one block — separate them
python
# Prefill to force a JSON object start (thinking OFF)
messages = [
{"role": "user", "content": "Return the invoice as JSON."},
{"role": "assistant", "content": "{"}, # model continues valid JSON from here
]

Prefill vs thinking

You cannot prefill the assistant turn when extended/adaptive thinking is enabled — the thinking block must be first. If you need both reasoning and a constrained format, use structured outputs (output_config.format) rather than a prefill, or run reasoning in a separate step.


4.11 Grounding and hallucination control

Format reliability (structured outputs) is separate from factual reliability. To keep answers grounded, engineer the context and the request, not just the schema.

TechniqueEffect
Docs-first, query-last orderingModel reads evidence before the question → better grounding
‘Answer only from the provided documents; if absent, say so’Suppresses ungrounded guessing
Document citationsForces the model to point at source spans it used
Retrieval (RAG) of the right spansPuts the actual evidence in context
Validation of claims against sourcesCatches fabricated values before they ship

Exam signal

‘Confident but wrong / fabricated facts’ → grounding controls (docs-first, ‘answer only from context’, citations, RAG, claim validation), not temperature or a bigger max_tokens. Schema conformance does not prevent a wrong-but-well-formatted value.


4.12 Common misconceptions

MisconceptionRealityWhy it matters on the exam
Asking for JSON in the prompt is enoughPrompt-only JSON drifts; use output_config.format / strict + validation-retryStructured-output questions
Confident output is validated outputFluency ≠ validity; parse defensivelyAnti-pattern #7 avoidance
Retrying with the same prompt fixes bad JSONFeed the specific error back so the next attempt targets itValidation-retry design
Schema conformance means the values are correctShape is guaranteed, not semantics; validate values and ground claimsSeparates format from factual reliability
A bigger max_tokens fixes a bloated agentAddress context via editing/compaction/isolationContext-management trap
Order of documents vs question does not matterDocs-first, query-last improves grounding and cachingLong-context ordering
More few-shot examples always helpRepresentative, diverse, correct examples matter more than volumeEdge-case failure questions
You can prefill while thinking is enabledYou cannot; the thinking block must come first — use structured outputs insteadPrefill-vs-thinking trap

4.13 Scenario walkthrough: a flaky invoice extractor

Scenario. An extraction endpoint on Sonnet 5 must return {total_cents: int, currency: str} for each invoice and cite the line it read the total from. In production it fails three ways: (1) about 3% of responses are almost-valid JSON with an extra trailing field; (2) a defensive try/except currently returns {} marked ‘success’ when parsing fails, so bad rows slip through silently; (3) on scanned handwritten invoices it confidently returns a plausible-but-wrong total. The team asks whether to raise temperature, switch to Opus 5, or increase max_tokens.

Expert reasoning trace.

  1. Enforce the shape, do not beg for it. The extra-field drift is fixed by output_config.format with the JSON schema (or a strict: true tool), which constrains output to the schema. Prompt-only JSON is the weakest option.
  2. Stop the silent suppression. Returning {}-as-success is anti-pattern #7. Replace it with a validation-retry loop that validates against the schema and, on failure, feeds the specific error back to the model; if it still fails, fail loudly with diagnostic context. Never report failure as empty success.
  3. Ground the factual reliability separately. The confident-but-wrong total on handwritten invoices is a grounding problem, not a format problem. Enable document citations, instruct ‘answer only from the provided document’, and validate the extracted total against the cited span. Consider representative few-shot examples that include handwritten cases (diversity over volume).
  4. Reject the tempting levers. ‘Raise temperature’ — increases variability, the opposite of what you want. ‘Switch to Opus 5’ — may help marginally but does not fix the missing schema constraint or the silent-suppression bug; it is an over-engineered non-fix. ‘Increase max_tokens’ — addresses truncation, which is not the reported failure.
  5. Order the prompt for grounding and caching. Put the stable instructions/schema and the document first (cacheable prefix), the extraction request last.

Correct decision. Schema-constrained output plus a validation-retry loop that feeds errors back and fails loudly; document citations and ‘answer only from context’ plus value-vs-source validation for factual grounding; representative handwritten few-shot examples; docs-first, query-last layout. Temperature, model swap and max_tokens are rejected as wrong-lever fixes.


Exam traps in this domain

TrapWhy it is wrong
Trusting confident output without validationFluency ≠ validity; validate against a schema
Relying on self-reported confidenceAnti-pattern #4; use programmatic validation
Swallowing parse errors as empty successAnti-pattern #7; surface and retry
Forcing a tool on Fable 5.1 for structured outputReturns 400; use output_config.format / strict
Only raising max_tokens to fix a bloated agentAddress context via editing/compaction/isolation
Putting the query before long documentsHurts grounding and caching; docs first, query last
No examples/format spec when output must be parseableUnder-specified prompts drift in shape
Prompt-based enforcement of critical rulesBelongs in hooks/validation, not the prompt
Splicing per-request values (timestamp, user name) into the cached prefixAny byte change misses the cache; keep variability after the breakpoint
Using compaction to drop large stale tool resultsThat is context editing’s job; compaction summarises to preserve narrative
Adding more common-case few-shot examples to fix edge-case failuresAdd representative edge-case examples; diversity beats volume
Placing few-shot examples after the queryKeep examples in the stable prefix so they stay cacheable and read before the task
Treating schema conformance as factual correctnessShape is guaranteed, not values; ground claims and validate semantics
Raising temperature to fix confident-but-wrong outputFabrication is a grounding problem; use docs-first, citations, RAG and claim validation
Trying to prefill the assistant turn while thinking is enabledThe thinking block must be first; use structured outputs instead
Retrying invalid JSON with the identical promptFeed the specific parse/schema error back so the next attempt targets it
Asking for step-by-step reasoning and strict JSON in one blockSeparate reasoning from the structured answer, or use structured outputs

Practice questions

Q1 · An extraction endpoint must return JSON matching a fixed schema every time. Which approach is most reliable? (Select one)

A. Ask nicely for JSON in the prompt and hope. B. Use output_config.format with a JSON schema (or a strict: true tool schema) plus validation-retry. C. Set temperature: 0 only. D. Increase max_tokens.

Answer: B. Schema-constrained structured output with validation-retry is the reliable pattern. Prompt-only (A) drifts; temperature (C) and max_tokens (D) do not enforce shape.

Q2 · A parser occasionally receives malformed JSON and currently returns an empty object as if successful. What is wrong and the fix? (Select one)

A. Nothing; empty is a safe default. B. It silently suppresses errors (anti-pattern #7); validate against the schema and retry with the error, or fail loudly. C. Lower temperature. D. Switch to Opus 5.

Answer: B. Returning empty success hides failures. Defensive parsing validates and either retries with the error or surfaces it. Temperature (C) and model choice (D) do not fix silent suppression.

Q3 · A long-running agent gets slower and less focused as tool outputs accumulate. Which TWO techniques address this? (Select two)

A. Context editing to clear stale tool results. B. Increase max_tokens. C. Compaction / summarisation of history. D. Add more tools. E. Raise temperature.

Answer: A and C. Context editing and compaction/summarisation reduce bloat and drift. max_tokens (B) caps output; more tools (D) and temperature (E) do not help.

Q4 · A prompt places the user's question before a 300-page document. Answers are poorly grounded. What ordering is better and why? (Select one)

A. Keep it; order does not matter. B. Put the document first and the question last, which improves grounding and enables caching of the stable prefix. C. Put both in the system prompt. D. Split the document across many messages randomly.

Answer: B. Documents-first, query-last improves grounding on long context and makes the stable prefix cacheable. Order does matter (A); random splitting (D) hurts coherence.

Q5 · A team wants to force a specific extraction tool on Fable 5.1 to guarantee structured output. What should they do instead? (Select one)

A. Force the tool anyway; it works on all models. B. Use output_config.format with a JSON schema, since Fable 5.1 rejects forced tool_choice. C. Disable thinking. D. Lower max_tokens.

Answer: B. Fable 5.1 returns 400 for forced tools; use output_config.format (or strict: true / auto + instruction). Forcing (A) fails; thinking (C) and max_tokens (D) are irrelevant.

Q6 · Output shape is inconsistent across runs, making it hard to parse. Which TWO prompt levers most directly improve consistency? (Select two)

A. Provide 2–5 few-shot input→output examples. B. Specify the exact output format (and consider prefilling the assistant turn). C. Increase temperature. D. Add a friendly tone instruction. E. Remove the system prompt.

Answer: A and B. Examples and an explicit format spec (with prefill) constrain the shape. Higher temperature (C) increases variability; tone (D) and removing the system prompt (E) do not help shape.

Q7 · A validation-retry loop rejects the model's JSON but re-sends only the original prompt each time, so the same error recurs. What is the fix? (Select one)

A. Increase the retry count to 10. B. Feed the specific schema/parse error back to the model as a new user turn so it can correct that exact problem. C. Switch to Opus 5. D. Lower max_tokens.

Answer: B. Effective validation-retry appends the failing output and the concrete error (‘Invalid: total must be a number’) so the next attempt targets the actual defect. Blindly retrying (A) repeats the same mistake; model choice (C) and max_tokens (D) do not convey what was wrong.

Q8 · An agent's window is dominated by large, old tool results that are no longer needed, but the conversation narrative must stay coherent for later reasoning. Which control fits, and which does not? (Select one)

A. Compaction, because it deletes the tool results outright. B. Context editing to clear the stale tool results; if the narrative itself were too long, compaction (summarise, preserve narrative) would be the choice. C. Just raise max_tokens. D. Restart the session and lose all state.

Answer: B. Context editing clears stale tool outputs; compaction is for condensing a long narrative while preserving it. max_tokens (C) only caps output; restarting (D) discards needed state. Compaction does not simply delete (A).

Q9 · A prompt reused on every request splices the current timestamp into the system prompt, and cache hit rates are near zero. Why, and what is the fix? (Select one)

A. Caching is disabled on Sonnet 5. B. Any byte change in the prefix (the timestamp) invalidates the cache; move per-request values after the cache breakpoint and keep the prefix stable. C. The prefix is too short; pad it with whitespace. D. Raise the TTL to 1 hour.

Answer: B. A caching prefix must be byte-identical to hit; a per-request timestamp changes it every call. Keep variable content in the tail after the breakpoint. Caching is not disabled (A); padding whitespace (C) does not fix a changing prefix; a longer TTL (D) still requires identical bytes.

Q10 · A summariser answers common documents well but repeatedly mishandles a rare edge-case format. Which TWO changes best improve it? (Select two)

A. Add representative few-shot examples that include the edge-case format. B. Add ten more examples of the common format. C. Increase example diversity to cover the range of formats. D. Raise temperature. E. Remove all examples.

Answer: A and C. Edge-case failures call for representative edge-case examples and greater diversity, not more of the same common case (B). Temperature (D) adds variability; removing examples (E) discards the steering entirely.

Q11 · A long-document QA prompt places the user's question first, then 200 pages of source. Answers are poorly grounded and caching never helps. Which reordering fixes both? (Select one)

A. Keep the order; move the question into the system prompt. B. Put the documents (and stable instructions) first with a cache breakpoint, and the question last. C. Interleave the question between every page. D. Split the documents across many separate requests at random.

Answer: B. Documents-first, query-last improves grounding and makes the stable prefix cacheable in one move. Moving the question into the system prompt (A) still precedes the evidence; interleaving (C) and random splitting (D) harm coherence and caching.

Q12 · A defensive parser catches invalid JSON and returns an empty result marked 'success' so the pipeline keeps running. Which anti-pattern is this and what should happen instead? (Select one)

A. It is fine; empty results are a safe default. B. Silently suppressing errors (anti-pattern #7); validate against the schema and either retry with the error fed back or fail loudly with diagnostic context. C. Self-report reliance; ask the model for its confidence. D. Over-engineering; remove the validation.

Answer: B. Returning empty-as-success hides failures downstream (anti-pattern #7). The correct behaviour is to surface the error — retry with the specific validation message or fail loudly. Empty defaults (A) mask the problem; confidence self-report (C) is a different anti-pattern (#4); removing validation (D) makes it worse.

Q13 · An extractor returns well-formatted JSON, but on scanned handwritten invoices the `total_cents` value is confidently wrong. Which set of changes BEST addresses this? (Select one)

A. Raise temperature so it explores more. B. Enable document citations, instruct ‘answer only from the provided document’, validate the value against the cited span, and add representative handwritten few-shot examples. C. Increase max_tokens. D. Switch to a strict: true schema and consider it solved.

Answer: B. A confident-but-wrong value is a grounding problem, addressed by citations, answer-only-from-context, value-vs-source validation and representative examples. Temperature (A) adds variability; max_tokens (C) is about truncation; a strict schema (D) fixes shape, not factual correctness.

Q14 · A developer wants to force JSON by prefilling the assistant turn with `{`, but the request also enables adaptive thinking, and it errors. Why, and what should they do? (Select one)

A. Prefill always works; the error is transient. B. You cannot prefill when thinking is enabled (the thinking block must come first); use output_config.format structured outputs instead, or run reasoning separately. C. Prefill requires tool_choice: 'any'. D. Lower max_tokens so the prefill fits.

Answer: B. With thinking on, the assistant turn must begin with the thinking block, so a prefill conflicts; structured outputs achieve constrained JSON without a prefill. The error is not transient (A); prefill does not require forced tools (C); max_tokens (D) is unrelated.

Q15 · Which statement best captures the relationship between schema-constrained output and correctness? (Select one)

A. A schema guarantees both the shape and the truth of the values. B. A schema guarantees the shape; values can still be wrong, so validate semantics and ground claims separately. C. Schemas are only cosmetic and do not constrain output. D. A schema removes the need for any validation.

Answer: B. output_config.format/strict constrain shape, not semantic correctness; grounding and value validation remain necessary. It does not guarantee truth (A); it does constrain output (C); and it does not remove semantic validation (D).

Q16 · A prompt asks the model to think step-by-step and, in the same block, emit only strict JSON. Output is inconsistent. What is the best fix? (Select one)

A. Increase max_tokens. B. Separate the reasoning from the structured answer — reason first (or use thinking), then produce the JSON via structured outputs — rather than mixing both in one block. C. Raise temperature for variety. D. Remove the schema entirely.

Answer: B. Mixing free-form reasoning and strict JSON in one block fights itself; separate the steps or use structured outputs for the final answer. max_tokens (A) and temperature (C) do not resolve the conflict; removing the schema (D) abandons the requirement.

Q17 · A validation-retry loop keeps failing because it re-sends the same prompt and never tells the model what was wrong. Which change fixes it? (Select one)

A. Bump the retry count to 15. B. Append the failing output plus the concrete error (‘currency must be one of USD/EUR/GBP’) as a new user turn so the next attempt corrects that exact defect. C. Switch to Opus 5 for retries. D. Lower max_tokens each retry.

Answer: B. Feeding the specific error back lets the model target the actual defect. Blindly repeating (A) reproduces the same mistake; a different model (C) still is not told what was wrong; lowering max_tokens (D) conveys nothing and risks truncation.

Q18 · A summariser handles common invoices well but fails on a rare handwritten format. Which TWO few-shot changes help most? (Select two)

A. Add representative examples that include the handwritten format. B. Increase example diversity to cover the range of formats. C. Add ten more common-format examples. D. Raise temperature. E. Remove all examples.

Answer: A and B. Edge-case failures need representative edge-case examples and greater diversity, not more of the common case (C). Temperature (D) adds variability; removing examples (E) discards the steering.

Key takeaways

  • Be clear and direct; use XML tags to separate instructions from data; add examples; use role, chain-of-thought and prefilling deliberately.
  • Prefer schema-constrained structured output (output_config.format, strict: true tools) over prompt-only JSON.
  • Parse defensively: validate against a schema and retry with the error; never swallow failures or trust self-reported confidence.
  • On Fable 5.1, use structured outputs rather than forcing a tool.
  • Prevent context bloat/drift with tool-output pruning, context editing, compaction, summarisation and subagent isolation.
  • Order long contexts documents-first, query-last – better grounding and cache-friendly layout.
  • Schema conformance (output_config.format / strict) guarantees shape, not values — ground claims with docs-first ordering, citations, RAG and value validation.
  • Validation-retry must feed the specific error back; re-sending the same prompt just reproduces the mistake.
  • You cannot prefill the assistant turn while thinking is enabled; use structured outputs, and never mix free-form reasoning with strict JSON in one block.
  • For edge-case failures, add representative edge-case few-shot examples and increase diversity rather than piling on more common-case examples.

Last updated Sep 18, 2026