# D4 · Prompt and Context Engineering

Prompt engineering principles, structured-output patterns, defensive parsing and validation-retry, and context engineering to prevent drift and bloat.

import { Accordions, AccordionItem, Tabs, TabItem } from '@prosefly/astro-components';

This domain is roughly **6 of 53 items**. It tests whether you can write prompts that reliably produce the shape and quality you need, parse output defensively (never trusting confident text), and engineer context so long-running work does not drift or bloat. The core lesson: **the prompt is the interface; the context is the working memory – engineer both.**

## Learning objectives

By the end of this page you should be able to:

1. Apply core **prompt-engineering principles**: clear/direct instructions, XML tags, examples/few-shot, role, chain-of-thought, prefilling, output format.
2. Produce **structured output** via `output_config.format` (JSON schema), tool-use-as-schema, and `strict: true`.
3. Parse output **defensively** with **validation-retry** and appropriate skepticism.
4. Prevent **context drift and bloat**: tool-output pruning, context editing, compaction, summarisation.
5. Use **context isolation** via subagents and correct **long-context ordering**.
6. Lay out prompts to be **caching-aware**.

---

## 4.1 Prompt-engineering principles

| Principle | What it does | Example lever |
| --- | --- | --- |
| **Clear and direct** | Removes ambiguity | State the task, constraints, and output shape explicitly |
| **XML tags** | Delimits sections; separates instructions from data | `<document>…</document>`, `<instructions>…</instructions>` |
| **Examples / few-shot** | Shows the desired pattern | 2–5 representative input→output pairs |
| **Role** | Sets perspective and tone | `system`: 'You are a senior tax analyst.' |
| **Chain-of-thought** | Improves reasoning on hard tasks | Ask for step-by-step reasoning, or use thinking |
| **Prefilling** | Constrains the start of the reply | Prefill the assistant turn with `{` to force JSON |
| **Output format** | Makes output parseable | Specify exact schema / structure |

```python
messages = [
    {"role": "user", "content": "Extract fields from <invoice>...</invoice> as JSON."},
    {"role": "assistant", "content": "{"},   # prefill forces JSON start
]
```

:::tip[Exam signal]
"Output is inconsistent / hard to parse" → specify the format, use XML tags, add few-shot examples, and consider prefilling. "Reasoning is shallow on a hard task" → chain-of-thought / thinking.
:::

---

## 4.2 Structured-output patterns

Three ways to get machine-readable output, strongest first:

1. **`output_config.format` with a JSON schema** – the model is constrained to emit JSON matching the schema.
2. **Tool-use-as-schema** – define a tool whose `input_schema` is your target shape; the `tool_use` block's input is your structured data. Add **`strict: true`** to enforce the schema.
3. **Prompt + prefill + validation** – ask for JSON, prefill `{`, then validate (weakest; needs defensive parsing).

```json
{
  "model": "claude-sonnet-5",
  "max_tokens": 512,
  "output_config": {
    "format": {
      "type": "json_schema",
      "schema": {
        "type": "object",
        "properties": {
          "total": {"type": "number"},
          "currency": {"type": "string"},
          "line_items": {"type": "array", "items": {"type": "string"}}
        },
        "required": ["total", "currency"]
      }
    }
  },
  "messages": [{"role": "user", "content": "Extract totals from the invoice."}]
}
```

:::caution[Fable 5.1]
Because Fable 5.1 rejects forced `tool_choice`, prefer **`output_config.format`** or `strict: true` schemas rather than forcing a tool to obtain structured output.
:::

---

## 4.3 Defensive parsing and validation-retry

Never trust that output is valid just because it looks confident. Validate against the schema and **retry with the error** on failure.

```python
import json
from jsonschema import validate, ValidationError

SCHEMA = {"type": "object", "required": ["total", "currency"],
          "properties": {"total": {"type": "number"}, "currency": {"type": "string"}}}

def extract(client, messages, attempts=3):
    for _ in range(attempts):
        resp = client.messages.create(model="claude-sonnet-5", max_tokens=512, messages=messages)
        text = "".join(b.text for b in resp.content if b.type == "text")
        try:
            data = json.loads(text)
            validate(data, SCHEMA)
            return data
        except (json.JSONDecodeError, ValidationError) as e:
            messages.append({"role": "assistant", "content": text})
            messages.append({"role": "user", "content": f"Invalid output: {e}. Return valid JSON only."})
    raise ValueError("could not obtain valid structured output")
```

:::danger[Skepticism toward confident output]
A fluent, confident answer is not a validated one. Parse defensively, validate against a schema, and surface failures – do not swallow them (anti-pattern #7). Do not rely on the model's self-reported confidence (anti-pattern #4).
:::

---

## 4.4 Context drift and bloat

As a conversation or agent runs, the context fills with tool outputs, intermediate reasoning and stale turns. This **bloats** cost and latency and causes **drift** (the model loses the thread or over-weights noise).

| Technique | What it does |
| --- | --- |
| **Tool-output pruning** | Drop or truncate large/old tool results no longer needed |
| **Context editing** | Programmatically clear stale content (e.g., old tool results) |
| **Compaction** | Server-side summarisation preserving the narrative |
| **Summarisation** | Replace long history with a running summary |
| **Context isolation** | Push subtasks to subagents with their own windows |

:::tip[Exam signal]
"Long-running agent slows down / loses focus / costs balloon" → context editing, compaction, summarisation, and subagent isolation – not simply a bigger `max_tokens`.
:::

---

## 4.5 Long-context ordering and caching-aware layout

For long inputs, **put stable, large content first and the query last**:

```text
[ system: role + rules ]        ← stable, cache here
[ tools ]                        ← stable, cache here
[ long documents ]               ← stable, cache here
[ conversation history ]         ← grows
[ the current question / task ]  ← variable, last
```

- Documents-first, query-last improves grounding and answer quality on long contexts.
- The same layout is **caching-friendly**: the stable prefix is cached (0.1× on hits), and only the variable tail is reprocessed.

---

## 4.6 Long-context ordering: a worked example

On long inputs, **place large stable content first and the query last**. Two reasons: grounding quality improves when the model reads the evidence before the question, and the stable prefix becomes cacheable. Below, the documents and instructions come first; the user's actual question is the final block.

```json
{
  "model": "claude-sonnet-5",
  "max_tokens": 1024,
  "system": [
    {"type": "text", "text": "You answer strictly from the provided documents. If the answer is not present, say so."}
  ],
  "messages": [
    {"role": "user", "content": [
      {"type": "text", "text": "<documents>\n<doc id='1'>...50 pages of policy text...</doc>\n<doc id='2'>...contract text...</doc>\n</documents>"},
      {"type": "text", "text": "<question>What is the termination notice period, and which clause states it?</question>"}
    ]}
  ]
}
```

The ordering is: **system rules → documents → question**. If you put the question first, the model reads it without evidence in view, weakening grounding, and the variable question would sit inside what you wanted to cache.

:::tip[Exam signal]
'Answers to long-document questions are poorly grounded' or 'the question is at the top of a huge prompt' → reorder to **documents first, query last**. This also sets up caching (next section).
:::

---

## 4.7 Caching-aware prompt layout

Prompt caching rewards a **stable prefix first**. Put the content that repeats byte-for-byte across requests (system prompt, tools, long reference documents) at the front and mark the last stable block with `cache_control`. Everything variable (the current question, fresh turns) goes after the cached breakpoint.

```json
{
  "model": "claude-sonnet-5",
  "system": [
    {"type": "text", "text": "Long stable system instructions and schema...",
     "cache_control": {"type": "ephemeral"}}
  ],
  "tools": [
    {"name": "search", "description": "...", "input_schema": {"type": "object"}}
  ],
  "messages": [
    {"role": "user", "content": [
      {"type": "text", "text": "...long reference document reused every request...",
       "cache_control": {"type": "ephemeral"}},
      {"type": "text", "text": "Current question: ..."}
    ]}
  ]
}
```

- Cache reads cost ≈ 0.1× base input; writes ≈ 1.25× (5-min TTL) or 2× (1-hour TTL).
- The minimum cacheable prefix is ~1024 tokens (**2048 on Haiku 4.5**); shorter prefixes will not cache.
- Order the layout: `system → tools → stable documents → [cache breakpoint] → variable question`.

```text
[ system + schema ]      ┐
[ tools ]                │  stable prefix → mark cache_control here (cached at 0.1×)
[ reference documents ]  ┘
──────────────────────────────  cache breakpoint
[ current question ]        variable tail → reprocessed each request
```

:::caution[A prefix that changes never caches]
If any byte of the prefix differs between requests (a timestamp, a per-user name spliced into the system prompt), the cache misses every time. Keep per-request variability out of the prefix and after the breakpoint.
:::

---

## 4.8 Context editing vs compaction

Long agents outgrow the window. Two server-side controls manage it, and the exam expects you to pick the right one.

| | Context editing | Compaction |
| --- | --- | --- |
| What it does | Programmatically **clears/trims** stale content (e.g., old tool results) | **Summarises** older turns server-side, **preserving the narrative** |
| Keeps narrative? | No — it removes content | Yes — it condenses while retaining the thread |
| Best when | Old tool outputs are large and no longer needed | The full history matters but is too long to keep verbatim |
| Risk | Removing something still needed | Summary loses a needed detail |

```json
{
  "model": "claude-sonnet-5",
  "context_management": {
    "edits": [
      {"type": "clear_tool_uses", "trigger": {"type": "input_tokens", "value": 100000}}
    ]
  }
}
```

```json
{
  "model": "claude-sonnet-5",
  "context_management": {
    "compaction": {"trigger": {"type": "input_tokens", "value": 150000}}
  }
}
```

:::tip[Exam signal]
'Old tool results are bloating the window but the conversation thread must stay intact' → **compaction** (summarise, preserve narrative). 'Large stale tool outputs are no longer needed' → **context editing** (clear them). Neither is 'just raise `max_tokens`', which only affects output length.
:::

---

## 4.9 Few-shot example selection

Few-shot examples steer output shape and quality — but only if they are **representative, diverse, and correct**. Poor example selection can hurt more than it helps.

| Principle | Why it matters |
| --- | --- |
| **Representative** | Examples should match the real input distribution, including the hard/edge cases |
| **Diverse** | Cover the range of categories/formats so the model does not overfit one pattern |
| **Correct and consistent** | A wrong or inconsistent example teaches the wrong pattern; keep formatting identical |
| **Ordered deliberately** | Group by type or escalate difficulty; keep them before the query (stable prefix, cacheable) |
| **Right number (2–5)** | Enough to show the pattern without bloating context; more is not always better |
| **Dynamic selection** | For varied inputs, retrieve the few examples most similar to the current input |

:::tip[Exam signal]
'The model handles common cases but fails edge cases' → add **representative edge-case examples**, not just more common ones. 'Output overfits one format' → increase **example diversity**. Place examples in the stable prefix so they stay cacheable.
:::

---

## 4.10 Prefilling, stop sequences and chain-of-thought placement

Three fine-grained levers shape *how* the model produces output, and the exam tests when each helps.

| Lever | What it does | When to use | Pitfall |
| --- | --- | --- | --- |
| **Prefill** | Seed the assistant turn's start (e.g. `{`) to constrain format | Force JSON start, force a specific opening | Cannot prefill when thinking is on (thinking must come first) |
| **`stop_sequences`** | Halt generation at a marker | Cut off after a delimiter; bound output | The marker text is not included in output |
| **Chain-of-thought / thinking** | Reason step-by-step before answering | Hard multi-step reasoning | Do not ask for reasoning *and* strict JSON in one block — separate them |

```python
# Prefill to force a JSON object start (thinking OFF)
messages = [
    {"role": "user", "content": "Return the invoice as JSON."},
    {"role": "assistant", "content": "{"},   # model continues valid JSON from here
]
```

:::caution[Prefill vs thinking]
You cannot prefill the assistant turn when extended/adaptive thinking is enabled — the `thinking` block must be first. If you need both reasoning and a constrained format, use **structured outputs** (`output_config.format`) rather than a prefill, or run reasoning in a separate step.
:::

---

## 4.11 Grounding and hallucination control

Format reliability (structured outputs) is separate from *factual* reliability. To keep answers grounded, engineer the context and the request, not just the schema.

| Technique | Effect |
| --- | --- |
| **Docs-first, query-last ordering** | Model reads evidence before the question → better grounding |
| **'Answer only from the provided documents; if absent, say so'** | Suppresses ungrounded guessing |
| **Document `citations`** | Forces the model to point at source spans it used |
| **Retrieval (RAG) of the right spans** | Puts the actual evidence in context |
| **Validation of claims against sources** | Catches fabricated values before they ship |

:::tip[Exam signal]
'Confident but wrong / fabricated facts' → grounding controls (docs-first, 'answer only from context', citations, RAG, claim validation), not `temperature` or a bigger `max_tokens`. Schema conformance does not prevent a wrong-but-well-formatted value.
:::

---

## 4.12 Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| Asking for JSON in the prompt is enough | Prompt-only JSON drifts; use `output_config.format` / `strict` + validation-retry | Structured-output questions |
| Confident output is validated output | Fluency ≠ validity; parse defensively | Anti-pattern #7 avoidance |
| Retrying with the same prompt fixes bad JSON | Feed the specific error back so the next attempt targets it | Validation-retry design |
| Schema conformance means the values are correct | Shape is guaranteed, not semantics; validate values and ground claims | Separates format from factual reliability |
| A bigger `max_tokens` fixes a bloated agent | Address context via editing/compaction/isolation | Context-management trap |
| Order of documents vs question does not matter | Docs-first, query-last improves grounding and caching | Long-context ordering |
| More few-shot examples always help | Representative, diverse, correct examples matter more than volume | Edge-case failure questions |
| You can prefill while thinking is enabled | You cannot; the thinking block must come first — use structured outputs instead | Prefill-vs-thinking trap |

---

## 4.13 Scenario walkthrough: a flaky invoice extractor

**Scenario.** An extraction endpoint on Sonnet 5 must return `{total_cents: int, currency: str}` for each invoice and cite the line it read the total from. In production it fails three ways: (1) about 3% of responses are almost-valid JSON with an extra trailing field; (2) a defensive `try/except` currently returns `{}` marked 'success' when parsing fails, so bad rows slip through silently; (3) on scanned handwritten invoices it confidently returns a plausible-but-wrong total. The team asks whether to raise `temperature`, switch to Opus 5, or increase `max_tokens`.

**Expert reasoning trace.**

1. **Enforce the shape, do not beg for it.** The extra-field drift is fixed by **`output_config.format` with the JSON schema** (or a `strict: true` tool), which constrains output to the schema. Prompt-only JSON is the weakest option.
2. **Stop the silent suppression.** Returning `{}`-as-success is **anti-pattern #7**. Replace it with a **validation-retry** loop that validates against the schema and, on failure, feeds the *specific* error back to the model; if it still fails, fail loudly with diagnostic context. Never report failure as empty success.
3. **Ground the factual reliability separately.** The confident-but-wrong total on handwritten invoices is a *grounding* problem, not a format problem. Enable **document citations**, instruct 'answer only from the provided document', and **validate the extracted total against the cited span**. Consider representative **few-shot examples that include handwritten cases** (diversity over volume).
4. **Reject the tempting levers.** 'Raise `temperature`' — increases variability, the opposite of what you want. 'Switch to Opus 5' — may help marginally but does not fix the missing schema constraint or the silent-suppression bug; it is an over-engineered non-fix. 'Increase `max_tokens`' — addresses truncation, which is not the reported failure.
5. **Order the prompt for grounding and caching.** Put the stable instructions/schema and the document first (cacheable prefix), the extraction request last.

**Correct decision.** Schema-constrained output plus a validation-retry loop that feeds errors back and fails loudly; document citations and 'answer only from context' plus value-vs-source validation for factual grounding; representative handwritten few-shot examples; docs-first, query-last layout. Temperature, model swap and `max_tokens` are rejected as wrong-lever fixes.

---

## Exam traps in this domain

| Trap | Why it is wrong |
| --- | --- |
| Trusting confident output without validation | Fluency ≠ validity; validate against a schema |
| Relying on self-reported confidence | Anti-pattern #4; use programmatic validation |
| Swallowing parse errors as empty success | Anti-pattern #7; surface and retry |
| Forcing a tool on Fable 5.1 for structured output | Returns 400; use `output_config.format` / `strict` |
| Only raising `max_tokens` to fix a bloated agent | Address context via editing/compaction/isolation |
| Putting the query before long documents | Hurts grounding and caching; docs first, query last |
| No examples/format spec when output must be parseable | Under-specified prompts drift in shape |
| Prompt-based enforcement of critical rules | Belongs in hooks/validation, not the prompt |
| Splicing per-request values (timestamp, user name) into the cached prefix | Any byte change misses the cache; keep variability after the breakpoint |
| Using compaction to drop large stale tool results | That is context editing's job; compaction summarises to preserve narrative |
| Adding more common-case few-shot examples to fix edge-case failures | Add representative edge-case examples; diversity beats volume |
| Placing few-shot examples after the query | Keep examples in the stable prefix so they stay cacheable and read before the task |
| Treating schema conformance as factual correctness | Shape is guaranteed, not values; ground claims and validate semantics |
| Raising `temperature` to fix confident-but-wrong output | Fabrication is a grounding problem; use docs-first, citations, RAG and claim validation |
| Trying to prefill the assistant turn while thinking is enabled | The thinking block must be first; use structured outputs instead |
| Retrying invalid JSON with the identical prompt | Feed the specific parse/schema error back so the next attempt targets it |
| Asking for step-by-step reasoning and strict JSON in one block | Separate reasoning from the structured answer, or use structured outputs |

---

## Practice questions

<Accordions>
  <AccordionItem title="Q1 · An extraction endpoint must return JSON matching a fixed schema every time. Which approach is most reliable? (Select one)">
    A. Ask nicely for JSON in the prompt and hope.
    B. Use `output_config.format` with a JSON schema (or a `strict: true` tool schema) plus validation-retry.
    C. Set `temperature: 0` only.
    D. Increase `max_tokens`.

    **Answer: B.** Schema-constrained structured output with validation-retry is the reliable pattern. Prompt-only (A) drifts; temperature (C) and `max_tokens` (D) do not enforce shape.
  </AccordionItem>

  <AccordionItem title="Q2 · A parser occasionally receives malformed JSON and currently returns an empty object as if successful. What is wrong and the fix? (Select one)">
    A. Nothing; empty is a safe default.
    B. It silently suppresses errors (anti-pattern #7); validate against the schema and retry with the error, or fail loudly.
    C. Lower temperature.
    D. Switch to Opus 5.

    **Answer: B.** Returning empty success hides failures. Defensive parsing validates and either retries with the error or surfaces it. Temperature (C) and model choice (D) do not fix silent suppression.
  </AccordionItem>

  <AccordionItem title="Q3 · A long-running agent gets slower and less focused as tool outputs accumulate. Which TWO techniques address this? (Select two)">
    A. Context editing to clear stale tool results.
    B. Increase `max_tokens`.
    C. Compaction / summarisation of history.
    D. Add more tools.
    E. Raise temperature.

    **Answer: A and C.** Context editing and compaction/summarisation reduce bloat and drift. `max_tokens` (B) caps output; more tools (D) and temperature (E) do not help.
  </AccordionItem>

  <AccordionItem title="Q4 · A prompt places the user's question before a 300-page document. Answers are poorly grounded. What ordering is better and why? (Select one)">
    A. Keep it; order does not matter.
    B. Put the document first and the question last, which improves grounding and enables caching of the stable prefix.
    C. Put both in the system prompt.
    D. Split the document across many messages randomly.

    **Answer: B.** Documents-first, query-last improves grounding on long context and makes the stable prefix cacheable. Order does matter (A); random splitting (D) hurts coherence.
  </AccordionItem>

  <AccordionItem title="Q5 · A team wants to force a specific extraction tool on Fable 5.1 to guarantee structured output. What should they do instead? (Select one)">
    A. Force the tool anyway; it works on all models.
    B. Use `output_config.format` with a JSON schema, since Fable 5.1 rejects forced `tool_choice`.
    C. Disable thinking.
    D. Lower `max_tokens`.

    **Answer: B.** Fable 5.1 returns 400 for forced tools; use `output_config.format` (or `strict: true` / `auto` + instruction). Forcing (A) fails; thinking (C) and `max_tokens` (D) are irrelevant.
  </AccordionItem>

  <AccordionItem title="Q6 · Output shape is inconsistent across runs, making it hard to parse. Which TWO prompt levers most directly improve consistency? (Select two)">
    A. Provide 2–5 few-shot input→output examples.
    B. Specify the exact output format (and consider prefilling the assistant turn).
    C. Increase temperature.
    D. Add a friendly tone instruction.
    E. Remove the system prompt.

    **Answer: A and B.** Examples and an explicit format spec (with prefill) constrain the shape. Higher temperature (C) increases variability; tone (D) and removing the system prompt (E) do not help shape.
  </AccordionItem>

  <AccordionItem title="Q7 · A validation-retry loop rejects the model's JSON but re-sends only the original prompt each time, so the same error recurs. What is the fix? (Select one)">
    A. Increase the retry count to 10.
    B. Feed the specific schema/parse error back to the model as a new user turn so it can correct that exact problem.
    C. Switch to Opus 5.
    D. Lower `max_tokens`.

    **Answer: B.** Effective validation-retry appends the failing output and the concrete error ('Invalid: total must be a number') so the next attempt targets the actual defect. Blindly retrying (A) repeats the same mistake; model choice (C) and `max_tokens` (D) do not convey what was wrong.
  </AccordionItem>

  <AccordionItem title="Q8 · An agent's window is dominated by large, old tool results that are no longer needed, but the conversation narrative must stay coherent for later reasoning. Which control fits, and which does not? (Select one)">
    A. Compaction, because it deletes the tool results outright.
    B. Context editing to clear the stale tool results; if the narrative itself were too long, compaction (summarise, preserve narrative) would be the choice.
    C. Just raise `max_tokens`.
    D. Restart the session and lose all state.

    **Answer: B.** Context editing clears stale tool outputs; compaction is for condensing a long narrative while preserving it. `max_tokens` (C) only caps output; restarting (D) discards needed state. Compaction does not simply delete (A).
  </AccordionItem>

  <AccordionItem title="Q9 · A prompt reused on every request splices the current timestamp into the system prompt, and cache hit rates are near zero. Why, and what is the fix? (Select one)">
    A. Caching is disabled on Sonnet 5.
    B. Any byte change in the prefix (the timestamp) invalidates the cache; move per-request values after the cache breakpoint and keep the prefix stable.
    C. The prefix is too short; pad it with whitespace.
    D. Raise the TTL to 1 hour.

    **Answer: B.** A caching prefix must be byte-identical to hit; a per-request timestamp changes it every call. Keep variable content in the tail after the breakpoint. Caching is not disabled (A); padding whitespace (C) does not fix a changing prefix; a longer TTL (D) still requires identical bytes.
  </AccordionItem>

  <AccordionItem title="Q10 · A summariser answers common documents well but repeatedly mishandles a rare edge-case format. Which TWO changes best improve it? (Select two)">
    A. Add representative few-shot examples that include the edge-case format.
    B. Add ten more examples of the common format.
    C. Increase example diversity to cover the range of formats.
    D. Raise temperature.
    E. Remove all examples.

    **Answer: A and C.** Edge-case failures call for representative edge-case examples and greater diversity, not more of the same common case (B). Temperature (D) adds variability; removing examples (E) discards the steering entirely.
  </AccordionItem>

  <AccordionItem title="Q11 · A long-document QA prompt places the user's question first, then 200 pages of source. Answers are poorly grounded and caching never helps. Which reordering fixes both? (Select one)">
    A. Keep the order; move the question into the system prompt.
    B. Put the documents (and stable instructions) first with a cache breakpoint, and the question last.
    C. Interleave the question between every page.
    D. Split the documents across many separate requests at random.

    **Answer: B.** Documents-first, query-last improves grounding and makes the stable prefix cacheable in one move. Moving the question into the system prompt (A) still precedes the evidence; interleaving (C) and random splitting (D) harm coherence and caching.
  </AccordionItem>

  <AccordionItem title="Q12 · A defensive parser catches invalid JSON and returns an empty result marked 'success' so the pipeline keeps running. Which anti-pattern is this and what should happen instead? (Select one)">
    A. It is fine; empty results are a safe default.
    B. Silently suppressing errors (anti-pattern #7); validate against the schema and either retry with the error fed back or fail loudly with diagnostic context.
    C. Self-report reliance; ask the model for its confidence.
    D. Over-engineering; remove the validation.

    **Answer: B.** Returning empty-as-success hides failures downstream (anti-pattern #7). The correct behaviour is to surface the error — retry with the specific validation message or fail loudly. Empty defaults (A) mask the problem; confidence self-report (C) is a different anti-pattern (#4); removing validation (D) makes it worse.
  </AccordionItem>

  <AccordionItem title="Q13 · An extractor returns well-formatted JSON, but on scanned handwritten invoices the `total_cents` value is confidently wrong. Which set of changes BEST addresses this? (Select one)">
    A. Raise `temperature` so it explores more.
    B. Enable document citations, instruct 'answer only from the provided document', validate the value against the cited span, and add representative handwritten few-shot examples.
    C. Increase `max_tokens`.
    D. Switch to a `strict: true` schema and consider it solved.

    **Answer: B.** A confident-but-wrong value is a grounding problem, addressed by citations, answer-only-from-context, value-vs-source validation and representative examples. Temperature (A) adds variability; `max_tokens` (C) is about truncation; a strict schema (D) fixes shape, not factual correctness.
  </AccordionItem>

  <AccordionItem title="Q14 · A developer wants to force JSON by prefilling the assistant turn with `{`, but the request also enables adaptive thinking, and it errors. Why, and what should they do? (Select one)">
    A. Prefill always works; the error is transient.
    B. You cannot prefill when thinking is enabled (the thinking block must come first); use `output_config.format` structured outputs instead, or run reasoning separately.
    C. Prefill requires `tool_choice: 'any'`.
    D. Lower `max_tokens` so the prefill fits.

    **Answer: B.** With thinking on, the assistant turn must begin with the thinking block, so a prefill conflicts; structured outputs achieve constrained JSON without a prefill. The error is not transient (A); prefill does not require forced tools (C); `max_tokens` (D) is unrelated.
  </AccordionItem>

  <AccordionItem title="Q15 · Which statement best captures the relationship between schema-constrained output and correctness? (Select one)">
    A. A schema guarantees both the shape and the truth of the values.
    B. A schema guarantees the shape; values can still be wrong, so validate semantics and ground claims separately.
    C. Schemas are only cosmetic and do not constrain output.
    D. A schema removes the need for any validation.

    **Answer: B.** `output_config.format`/`strict` constrain shape, not semantic correctness; grounding and value validation remain necessary. It does not guarantee truth (A); it does constrain output (C); and it does not remove semantic validation (D).
  </AccordionItem>

  <AccordionItem title="Q16 · A prompt asks the model to think step-by-step and, in the same block, emit only strict JSON. Output is inconsistent. What is the best fix? (Select one)">
    A. Increase `max_tokens`.
    B. Separate the reasoning from the structured answer — reason first (or use thinking), then produce the JSON via structured outputs — rather than mixing both in one block.
    C. Raise temperature for variety.
    D. Remove the schema entirely.

    **Answer: B.** Mixing free-form reasoning and strict JSON in one block fights itself; separate the steps or use structured outputs for the final answer. `max_tokens` (A) and temperature (C) do not resolve the conflict; removing the schema (D) abandons the requirement.
  </AccordionItem>

  <AccordionItem title="Q17 · A validation-retry loop keeps failing because it re-sends the same prompt and never tells the model what was wrong. Which change fixes it? (Select one)">
    A. Bump the retry count to 15.
    B. Append the failing output plus the concrete error ('currency must be one of USD/EUR/GBP') as a new user turn so the next attempt corrects that exact defect.
    C. Switch to Opus 5 for retries.
    D. Lower `max_tokens` each retry.

    **Answer: B.** Feeding the specific error back lets the model target the actual defect. Blindly repeating (A) reproduces the same mistake; a different model (C) still is not told what was wrong; lowering `max_tokens` (D) conveys nothing and risks truncation.
  </AccordionItem>

  <AccordionItem title="Q18 · A summariser handles common invoices well but fails on a rare handwritten format. Which TWO few-shot changes help most? (Select two)">
    A. Add representative examples that include the handwritten format.
    B. Increase example diversity to cover the range of formats.
    C. Add ten more common-format examples.
    D. Raise temperature.
    E. Remove all examples.

    **Answer: A and B.** Edge-case failures need representative edge-case examples and greater diversity, not more of the common case (C). Temperature (D) adds variability; removing examples (E) discards the steering.
  </AccordionItem>
</Accordions>

## Key takeaways

- Be clear and direct; use XML tags to separate instructions from data; add examples; use role, chain-of-thought and prefilling deliberately.
- Prefer schema-constrained structured output (`output_config.format`, `strict: true` tools) over prompt-only JSON.
- Parse defensively: validate against a schema and retry with the error; never swallow failures or trust self-reported confidence.
- On Fable 5.1, use structured outputs rather than forcing a tool.
- Prevent context bloat/drift with tool-output pruning, context editing, compaction, summarisation and subagent isolation.
- Order long contexts documents-first, query-last – better grounding and cache-friendly layout.
- Schema conformance (`output_config.format` / `strict`) guarantees shape, not values — ground claims with docs-first ordering, citations, RAG and value validation.
- Validation-retry must feed the *specific* error back; re-sending the same prompt just reproduces the mistake.
- You cannot prefill the assistant turn while thinking is enabled; use structured outputs, and never mix free-form reasoning with strict JSON in one block.
- For edge-case failures, add representative edge-case few-shot examples and increase diversity rather than piling on more common-case examples.
