Domains
D5 · Context Management and Reliability
Context window economics, context editing vs compaction, memory and external state, subagent isolation, prompt caching, Fable 5.1 append-only history, retries and backoff, graceful degradation and fallback models, provenance, escalation gates, monitoring, rate limits and timeouts.
This domain is worth roughly 9 of 60 items and underpins every scenario: a system that works in a demo but degrades as the context window fills, or falls over when the API returns 429, is not production-ready. It tests the economics of the context window, the two ways to reclaim it (context editing vs compaction), what must live in durable state, and the reliability patterns – caching, retries, fallback, monitoring – that keep an agentic system running.
Learning objectives
By the end of this page you should be able to:
- Reason about context window economics and what consumes context.
- Choose between context editing (clear tool results) and compaction (summarise preserving narrative).
- Use the memory tool and external state for anything that must survive compaction, and subagent isolation to protect the main context.
- Use prompt caching as a reliability and cost lever, and honour Fable 5.1’s append-only constraint.
- Design retries/backoff/idempotency, graceful degradation and fallback models – and know what breaks when falling back from Fable 5.1.
- Maintain provenance and citations, and design escalation / human-in-the-loop gates.
- Set up monitoring (latency p95, error rate, cache-hit rate, cost per task) and design for rate limits, timeouts and partial results.
5.1 Context window economics
| Model | Context window | Max output |
|---|---|---|
| Claude Fable 5.1 | 1M | 128k |
| Claude Opus 5 | 1M | 128k |
| Claude Sonnet 5 | 1M | 128k |
| Claude Haiku 4.5 | 200k | 64k |
The window is a budget, not free space. Everything shares it:
┌───────── context window budget ─────────┐│ system prompt │ stable│ tool definitions │ stable│ conversation history (turns) │ grows every turn│ tool results (can be huge) │ grows fast│ thinking tokens │ grows with reasoning└──────────────────────────────────────────┘As the window fills, latency and cost rise and quality can degrade (“context rot”). The architect’s job is to keep the working set small: cache the stable parts, evict what is no longer needed, and push durable facts to external state.
Exam signal
“Long-running agent whose quality degrades over time” or “context is filling up” points to context editing / compaction / subagent isolation / memory, not a bigger model. Haiku 4.5’s window is 200k, not 1M – a common distractor.
5.2 Context editing vs compaction
Two distinct mechanisms reclaim context. Choosing correctly is tested.
| Mechanism | What it does | Preserves | Use when |
|---|---|---|---|
| Context editing | Clears old tool results (and similar bulky blocks) from the window | The conversation narrative; drops stale tool output | Tool results are large and no longer needed, but the dialogue matters |
| Compaction | Summarises the conversation server-side, preserving the narrative in condensed form | The gist/narrative of the whole conversation | The conversation itself has grown too long |
Context editing: [sys][tools][turn1][BIG tool result ✗ cleared][turn2]… ← keeps dialogue, drops bulkCompaction: [sys][tools][==== summary of turns 1..40 ====][turn41]… ← condenses the narrativeExam signal
“Bulky tool outputs we no longer need” → context editing (clear tool results). “The whole conversation is too long but we must keep its thread” → compaction (summarise preserving narrative). Neither preserves everything – anything that must survive verbatim goes to external state or the memory tool.
5.3 Memory tool and external state
Compaction and context editing are lossy. Anything that must survive – a decision, a customer fact, a checkpoint, canonical business data – belongs in durable state:
- Memory tool – model-managed persistence across sessions.
- Files – artefacts and checkpoints you control.
- Database – canonical, transactional business state (orders, tickets).
The conversation references durable state; it does not own it. This is the same rule as Domain 1’s state-vs-conversation table.
5.4 Subagent isolation protects the main context
A subagent runs in its own context window. Delegating a bulky sub-task (reading 50 files, searching a large corpus) to a subagent means the raw material never enters the coordinator’s window – only the distilled result returns. This is one of the most effective context-protection patterns, and it is why explicit context passing (Domain 1) matters: the subagent starts clean.
5.5 Long-context ordering and prompt caching
- Ordering. Put stable, reusable content first (system, tools, long reference docs); put the variable task last. This both enables caching and keeps the model’s attention on the right material.
- Prompt caching as a reliability + cost lever. Cache reads ≈ 0.1× base input price; writes ≈ 1.25× (5-min TTL) or 2× (1-hour TTL). Mark the last stable block with
cache_control: {"type": "ephemeral"}. Minimum cacheable prefix ~1024 tokens (2048 on Haiku). A high cache-hit rate lowers cost and latency, improving reliability under load.
system = [ {"type": "text", "text": SYSTEM_RULES}, {"type": "text", "text": TOOLS_DOC}, {"type": "text", "text": REFERENCE, "cache_control": {"type": "ephemeral"}}, # cache boundary]5.6 Fable 5.1: append-only history and thinking-block binding
Fable 5.1 introduces breaking changes an architect must design around:
- Append-only history. Editing, reordering or removing earlier turns invalidates later thinking blocks. Harnesses must be append-only: freeze
systemandtools, put mid-session changes inrole: "system"messages, and trim server-side via context editing/compaction rather than rewriting the transcript. - Thinking-block binding. Thinking blocks are readable only by the producing model or a newer one. Falling back to an older model silently drops those thinking blocks.
- Operational constraints. Requires 30-day data retention; not in the Priority Tier; thinking always on.
Design implication
On Fable 5.1 you cannot “clean up” the transcript by editing old turns. Design the harness to only append, and reclaim space with context editing/compaction – which operate server-side without invalidating the append-only chain.
5.7 Retries, backoff, idempotency
Retry the transient errors with exponential backoff and jitter, honouring retry-after:
| Status | Meaning | Retry? |
|---|---|---|
| 400 invalid_request | Bad request | No — fix the request |
| 401 / 403 | Auth / permission | No |
| 413 request_too_large | Too big | No — reduce size |
| 429 rate_limit | Rate limited | Yes — backoff, respect retry-after |
| 500 api_error / 529 overloaded | Server-side | Yes — backoff + jitter |
Only retry idempotent operations safely; give writes an idempotency key (Domain 1).
def call_with_backoff(fn, max_attempts=5): for i in range(max_attempts): try: return fn() except (RateLimitError, APIError, OverloadedError) as e: if i == max_attempts - 1: raise delay = min(2 ** i, 30) + random.random() # backoff + jitter time.sleep(getattr(e, "retry_after", delay))5.8 Graceful degradation and fallback models
When the primary model is unavailable or overloaded, degrade gracefully: fall back to another model, serve a cached/partial result, or queue for later. Fallback has costs.
| Fallback | Consideration |
|---|---|
| Fable 5.1 → older model | Thinking blocks are silently dropped; behaviour changes; forced-tool-choice rules differ; validate the fallback path explicitly |
| Opus 5 → Sonnet 5 | Cheaper/faster, some capability loss; Sonnet 5 disallows mid-conversation system messages and task budgets |
| Any → Haiku 4.5 | 200k window (not 1M); still uses budget_tokens; big context inputs may not fit |
Falling back from Fable 5.1
Because thinking blocks are readable only by the producing or a newer model, a fallback to an older model drops them and can change behaviour subtly. Test the degraded path; do not assume the fallback is a drop-in replacement.
5.9 Provenance and citations
For any output that will be relied on, preserve where it came from: keep citations from web search / retrieval, attach source IDs to extracted facts, and surface provenance so a human can verify. Provenance is both a quality and a governance control (and pairs with Domain 3’s per-type metrics and validation).
5.10 Escalation and human-in-the-loop gates
Irreversible or high-stakes actions need a human gate – the same escalation logic as Domain 1 (explicit request → immediate; capability gap → after attempting; never sentiment/self-report), plus mandatory human approval before irreversible effects (payments, deletions, external sends). The gate is deterministic (a permission/hook), not a prompt request.
5.11 Monitoring
Instrument the system with the metrics that predict failure:
| Metric | Why it matters |
|---|---|
| Latency p50 / p95 | p95 catches the tail users actually feel |
| Error rate by category | Distinguishes rate-limit from validation from model failures |
| Cache-hit rate | Low hit rate silently inflates cost and latency |
| Cost per task | The unit economics that decide viability |
| Retry counts / timeouts | Rising retries signal upstream trouble |
Pair with per-agent traces and correlation IDs (Domain 1).
5.12 Rate limits, timeouts and partial results
- Rate limits are per-model-tier (RPM, ITPM, OTPM). Design concurrency and batching to stay under them; use the Batch API (50% discount, within 24h) for latency-tolerant work.
- Timeouts. Set them at every boundary (tool calls, subagents, the overall task); a hung tool must not hang the agent.
- Partial results. On timeout or partial failure, return what you have with provenance of the gap (Domain 1’s partial-failure handling), rather than nothing or a false-complete answer.
5.13 Context economics with token arithmetic
Treat the window as a budget you can compute against. A worked example for a long-running agent on Opus 5 (1M window, $5/M in, $25/M out):
Turn budget as the session grows (input tokens carried each turn): system + tools (stable) = 12,000 CLAUDE-style reference docs = 8,000 conversation history @ turn 30 = 60,000 accumulated tool results = 400,000 ← the runaway thinking (this turn) = 10,000 ---------------------------------------------- input this turn = 490,000 tokens cost of THIS input turn = 490,000 × $5 / 1e6 = $2.45 (and rising every turn)The accumulated tool results dominate. Three levers, with their arithmetic:
- Context editing clears the 400k of stale tool results → input drops to ~90k → this-turn input cost falls from $2.45 to ~$0.45.
- Prompt caching the stable 20k prefix (system + tools + docs) → those tokens read at ~0.1× (
20,000 × $0.5/1e6 = $0.01vs$0.10). - Subagent isolation: delegate the tool-heavy sub-task so the 400k never enters the main window at all — only the distilled result returns.
| Symptom | Mechanism | Rough effect |
|---|---|---|
| Bulky, stale tool results | Context editing | Removes the largest term |
| Whole dialogue too long | Compaction | Condenses history to a summary |
| Repeated stable prefix | Prompt caching | ~10× cheaper reads on the prefix |
| Sub-task needs huge raw input | Subagent isolation | Raw material never hits main window |
Exam signal
“Context is filling / cost rising every turn” is almost never solved by “a bigger model”. The math points to context editing (remove the biggest term), caching (cheapen the stable term), or subagent isolation (keep the big term out). Haiku 4.5’s window is 200k, so it is smaller, not a fix.
5.14 The Fable 5.1 harness: constraints and a compliant design
Fable 5.1 (claude-fable-5-1, 1M/128k, $10/$50) has the strictest harness rules on the exam. Design around all of them:
| Constraint | Consequence | Compliant design |
|---|---|---|
| Append-only history | Editing/reordering/removing earlier turns invalidates later thinking blocks | Never rewrite the transcript; only append |
| Thinking-block binding | Thinking readable only by the producing model or newer | A fallback to an older model silently drops thinking → test the degraded path |
No forced tool_choice | any / {type:"tool"} return 400 | Use auto + instruction, strict tools, or structured outputs |
| Thinking always on | Cannot disable thinking | Budget for thinking tokens; no budget_tokens (Haiku-only) |
| 30-day retention, no ZDR | Not zero-data-retention | Exclude for ZDR-required workloads |
| Not Priority Tier | No priority capacity guarantees | Plan capacity/fallback accordingly |
Compliant Fable 5.1 harness: ┌ freeze system + tools (never edited) ─────────────┐ │ append user/assistant turns only │ │ mid-session change? → append a role:"system" msg │ (NOT edit an old turn) │ window pressure? → server-side context editing │ │ / compaction (append-safe) │ │ need structured out?→ output_config.format / strict │ (never forced tool_choice) └────────────────────────────────────────────────────┘Do not “tidy” a Fable 5.1 transcript
Rewriting old turns to remove noise breaks thinking-block binding and produces inconsistent later responses. Reclaim space with server-side context editing/compaction, which do not invalidate the append-only chain.
5.15 Reliability budgets: timeouts, concurrency and rate-limit math
Reliability is quantitative. A quick capacity check for a real-time workload:
Tier limits (example): RPM = 4,000 ITPM = 400,000 input tokens/minPer request: ~10,000 input tokensToken-bound throughput: 400,000 / 10,000 = 40 requests/min ← ITPM binds firstRequest-bound: 4,000 RPM=> Effective ceiling = min(40, 4,000) = 40 rpm; the token limit dominates.If you need 10,000 latency-tolerant jobs, the real-time ceiling (40/min ≈ 4+ hours and rate-limit risk) argues for the Batch API (50% off, ≤24h) instead. Set timeouts at every boundary (tool, subagent, overall task) so one hung call cannot stall the system, and on timeout return partial results with provenance of the gap.
| Boundary | Timeout | On breach |
|---|---|---|
| Single tool call | seconds | Structured timeout error, retryable |
| Subagent | tens of seconds | Partial result + gap provenance |
| Overall task | task-appropriate | Escalate or serve partial with a note |
Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| “A bigger model fixes a filling context.” | The fix is context editing/compaction/subagents/caching; a bigger model just delays it. | The top D5 distractor. |
| “Haiku 4.5 has a 1M window.” | Haiku 4.5 is 200k; large inputs may not fit. | Fallback/window items. |
| “Compaction and context editing are the same.” | Compaction summarises the narrative; context editing clears tool results. | Mechanism-choice items. |
| “Must-survive facts are safe in the conversation.” | Compaction/editing are lossy; use the memory tool or a database. | State-vs-conversation items. |
| “Fallback from Fable 5.1 is a drop-in.” | Thinking blocks are dropped on older fallbacks; behaviour changes. | Graceful-degradation items. |
| “Retry any error with backoff.” | Only 429/5xx/529 are retryable; 400/401/403/413 must be fixed. | Retry-classification items. |
| “Average latency is enough to monitor.” | p95/p99 catch the tail users feel; averages hide it. | Monitoring items. |
| “Editing old turns tidies a Fable 5.1 session.” | It breaks thinking-block binding; the harness must be append-only. | Fable 5.1 harness items. |
Scenario walkthrough — a research agent that is fast in a demo and falls over in production
Situation. A multi-agent research agent works in demos but in production (1) degrades after ~30 turns as the window fills with large search results it no longer needs, while the dialogue thread must stay intact; (2) during an Anthropic capacity event it fails over from Fable 5.1 to an older model and starts behaving differently even with an identical prompt; (3) an SRE reports it is “fine on average” yet users occasionally wait 9 seconds and the monthly bill is a surprise; and (4) a nightly run of 10,000 extraction jobs keeps hitting rate limits.
Expert reasoning trace.
-
Problem 1 — filling window, keep dialogue. Bulky, stale tool results with the narrative preserved → context editing, not compaction (which would summarise the dialogue) and not a bigger model. Additionally, delegate the search-heavy work to subagents so raw results never enter the main window.
-
Problem 2 — fallback behaviour change. Fable 5.1 thinking blocks are readable only by that model or newer, so the older fallback silently drops them. The design must test the degraded path and not assume parity; window size and key expiry are red herrings.
-
Problem 3 — hidden tail and cost. Add p95/p99 latency (the 9-second tail averages hide) and cost per task (the surprise bill). Reject “total request count” and “model name” as non-diagnostic.
-
Problem 4 — rate limits on bulk. 10,000 latency-tolerant jobs belong on the Batch API (50% off, ≤24h), which eases rate-limit pressure. Reject real-time max concurrency (full price, limit risk) and one giant request (will not fit).
Exam-correct decision: context editing + subagent isolation for the window; a tested degraded path for the Fable 5.1 fallback; p95/p99 and cost-per-task monitoring; and the Batch API for bulk. Every rejected option is a named trap (bigger-model, parity-assumption, average-only monitoring, real-time bulk).
Exam traps in this domain
| Trap | Why it is wrong |
|---|---|
| “Use a bigger model” to fix a filling context | The fix is context editing/compaction/subagents/memory |
| Claim Haiku 4.5 has a 1M window | Haiku 4.5 is 200k |
| Keep must-survive facts only in conversation | Compaction/editing are lossy; use memory tool or DB |
| Use compaction to drop bulky tool results | That is context editing’s job; compaction summarises the narrative |
| Edit old turns to tidy a Fable 5.1 transcript | Invalidates later thinking blocks; harness must be append-only |
| Fall back from Fable 5.1 assuming parity | Thinking blocks are dropped; behaviour changes |
| Retry a 400/413 with backoff | Non-retryable; fix the request/size |
Retry without jitter or ignoring retry-after | Causes thundering-herd; respect the header |
| Monitor only average latency | p95 catches the tail; averages hide it |
| Return nothing on a partial failure | Return partial results with provenance of the gap |
| “Bigger model” for a cost-rising-every-turn session | Context editing removes the biggest term; caching cheapens the stable term |
| Ignore ITPM/RPM when sizing real-time throughput | The token limit often binds first; compute the effective ceiling |
| Use real-time high concurrency for 10k overnight jobs | The Batch API (50% off, ≤24h) fits and eases rate-limit pressure |
| Skip timeouts on tool/subagent boundaries | A hung call stalls the whole agent; set boundary timeouts |
| Use Fable 5.1 for a ZDR-required workload | Fable 5.1 has 30-day retention and no ZDR; exclude it |
Practice questions
Q1 · A long-running agent's answers degrade after many turns as the context fills with large tool outputs the agent no longer needs. What is the BEST fix? (Select one)
A. Switch to a model with a bigger context window. B. Use context editing to clear stale tool results while preserving the conversation narrative. C. Restart the conversation from scratch each time. D. Increase max_tokens.
Answer: B. Clearing bulky, no-longer-needed tool results is exactly context editing. A bigger window (A) delays the problem, restarting (C) loses state, and max_tokens (D) is unrelated.
Q2 · A conversation itself has grown very long but the thread must be preserved for the agent to stay coherent. Which mechanism fits? (Select one)
A. Context editing (clear tool results). B. Compaction: summarise the conversation server-side, preserving the narrative in condensed form. C. Delete the oldest turns and hope. D. Move to Haiku 4.5 for its larger window.
Answer: B. Compaction condenses the narrative when the conversation is the thing that is too long. Context editing (A) targets tool results, deleting turns (C) loses the thread (and breaks Fable 5.1 append-only), and Haiku 4.5 (D) has a smaller window.
Q3 · A customer preference must be available in a session next week. Where should it live? (Select one)
A. Conversation history. B. The memory tool or an external database — durable state that survives compaction and new sessions. C. A thinking block. D. The system prompt of the current session.
Answer: B. Durable, cross-session facts need the memory tool or external state. Conversation history (A) and thinking blocks (C) are lossy/ephemeral, and a single session’s system prompt (D) does not persist.
Q4 · On Claude Fable 5.1, a harness edits earlier turns to remove noise and later responses become inconsistent. What is the cause and fix? (Select one)
A. The model is faulty; open a ticket. B. Editing earlier turns invalidates later thinking blocks; make the harness append-only and reclaim space with context editing/compaction instead. C. Increase the context window. D. Turn thinking off.
Answer: B. Fable 5.1 is append-only; editing turns breaks thinking-block binding. The fix is an append-only harness with server-side trimming. It is not a bug (A), window size (C) is unrelated, and thinking cannot be turned off on Fable 5.1 (D).
Q5 · A system falls back from Fable 5.1 to an older model during an outage and behaviour changes unexpectedly. What is the MOST likely reason? (Select one)
A. The older model has a bigger window. B. Thinking blocks produced by Fable 5.1 are readable only by that model or newer, so the older fallback silently drops them, changing behaviour. C. The API key expired. D. The prompt cache was cold.
Answer: B. Thinking-block binding means older fallbacks drop the thinking, altering behaviour. Window size (A) does not cause this, key expiry (C) would error not change behaviour, and a cold cache (D) affects cost/latency not correctness.
Q6 · Which TWO errors should be retried with exponential backoff and jitter? (Select two)
A. 429 rate_limit. B. 400 invalid_request. C. 529 overloaded. D. 401 authentication. E. 413 request_too_large.
Answer: A and C. 429 and 529 are transient and retryable with backoff (respect retry-after). 400 (B), 401 (D) and 413 (E) are client-side and must be fixed, not retried.
Q7 · A stakeholder says 'just use Haiku 4.5 as the fallback; it has the same 1M window'. What is the correct correction? (Select one)
A. They are right. B. Haiku 4.5’s context window is 200k, not 1M, so large inputs may not fit; the fallback path must be validated. C. Haiku 4.5 has a 2M window. D. Haiku 4.5 cannot be used as a fallback at all.
Answer: B. Haiku 4.5 is 200k. Large-context requests that fit Fable/Opus/Sonnet’s 1M may overflow Haiku. It is not 1M (A), not 2M (C), and can be a fallback if inputs fit (D is too absolute).
Q8 · To improve reliability and cost under load, an architect wants a high cache-hit rate. Which arrangement achieves this? (Select one)
A. Put the variable user input first and the stable system prompt last.
B. Put the stable system prompt, tools and reference docs first with cache_control on the last stable block; keep the variable task after.
C. Disable caching to avoid stale reads.
D. Randomise the prompt order each call.
Answer: B. Caching needs stable content first with a cache boundary; the variable task follows. Reversed order (A) prevents hits, disabling caching (C) raises cost, and randomising (D) destroys the cacheable prefix.
Q9 · A latency-tolerant batch of 10,000 extraction jobs must be processed cheaply. What is the BEST choice? (Select one)
A. Real-time Messages API calls with high concurrency. B. The Message Batches API (50% discount, results within 24h) for latency-tolerant workloads. C. One giant single request. D. Haiku 4.5 in a tight synchronous loop.
Answer: B. The Batch API is designed for latency-tolerant bulk work at 50% off. Real-time high concurrency (A) risks rate limits and costs more, one request (C) will not fit, and a synchronous loop (D) is slow and rate-limited.
Q10 · A subagent must read 50 large files to answer one sub-question. How does delegating this protect reliability? (Select one)
A. It does not; put all 50 files in the coordinator’s context. B. The subagent’s isolated context holds the raw files; only the distilled result returns to the coordinator, keeping the main window small. C. It doubles the cost with no benefit. D. It removes the need for error handling.
Answer: B. Subagent isolation keeps bulky raw material out of the main window and returns only the result. Loading all files into the coordinator (A) bloats context, and B is a real benefit not pure cost (C); error handling still matters (D).
Q11 · An SRE monitors only average latency and misses that some users see 8-second responses. What metric should be added? (Select one)
A. Total request count. B. p95 (and p99) latency, which captures the tail that averages hide. C. The number of tools. D. Model name.
Answer: B. p95/p99 expose the tail experience averages mask. Request count (A), tool count (C) and model name (D) do not reveal tail latency.
Q13 · At turn 30 an agent carries 20k stable tokens, 60k history, 400k accumulated tool results and 10k thinking, and per-turn input cost is rising. Which change reduces cost MOST directly while keeping the dialogue? (Select one)
A. Switch to Haiku 4.5 for a bigger window. B. Use context editing to clear the 400k of stale tool results (the dominant term), dropping this-turn input from ~490k to ~90k, while preserving the narrative. C. Increase max_tokens. D. Disable prompt caching.
Answer: B. The tool-result term dominates; clearing it via context editing removes most of the cost while keeping the dialogue. Haiku 4.5 (A) has a smaller 200k window, max_tokens (C) is output not input, and disabling caching (D) raises cost.
Q14 · A ZDR (zero-data-retention) requirement applies to a workload, and an architect proposes Fable 5.1 for its reasoning. What is the correct critique? (Select one)
A. Fable 5.1 is fine; all models support ZDR.
B. Fable 5.1 requires 30-day retention and is not ZDR, so it must be excluded for a ZDR-required workload; choose a compliant model/configuration.
C. Enable budget_tokens to turn on ZDR.
D. ZDR is only about the context window size.
Answer: B. Fable 5.1 has 30-day retention and no ZDR, disqualifying it where ZDR is required. All models do not support ZDR (A), budget_tokens is unrelated (C), and ZDR is a data-retention property, not window size (D).
Q15 · A tier allows 4,000 RPM and 400,000 input tokens/min; each request uses ~10,000 input tokens. What is the effective real-time throughput ceiling and the implication for 10,000 latency-tolerant jobs? (Select one)
A. 4,000 rpm; run them all in real time. B. ~40 rpm (the token limit binds first: 400,000 / 10,000), so real-time would be slow and limit-prone; use the Batch API (50% off, ≤24h) for the bulk run. C. Unlimited, because tokens do not count. D. 400,000 rpm.
Answer: B. ITPM binds first at ~40 rpm, making real-time bulk slow and risky; the Batch API fits. RPM alone (A) ignores the token limit, tokens always count (C), and D confuses tokens with requests.
Q16 · On Fable 5.1, an architect wants to reclaim window space by rewriting several early turns into a shorter summary in place. Why is this wrong and what is correct? (Select one)
A. It is correct and the cheapest option. B. Rewriting turns breaks Fable 5.1’s append-only history and invalidates later thinking blocks; instead reclaim space with server-side context editing/compaction and keep the harness append-only. C. It is fine if you also lower max_tokens. D. Switch off thinking to allow the rewrite.
Answer: B. Fable 5.1 is append-only; editing turns invalidates thinking blocks. Server-side editing/compaction reclaim space without breaking the chain. It is not correct (A, C), and thinking cannot be disabled on Fable 5.1 (D).
Q17 · A subagent gathering data times out midway. What should the system return to the coordinator? (Select one)
A. Nothing, to be safe. B. A structured partial result with provenance of what is missing (and a retryable timeout error), so the coordinator can retry, proceed with a quorum noting the gap, or escalate. C. A false ‘complete’ answer using only what arrived. D. A bare ‘error’ string.
Answer: B. Partial results plus gap provenance enable an explicit decision. Returning nothing (A) wastes work, a silent false-complete (C) is #7, and a bare error string (D) is #6.
Q18 · An SRE dashboard shows only average latency and total request count, and the team is surprised by both slow tail responses and cost. Which TWO metrics should be added FIRST? (Select two)
A. p95/p99 latency to capture the tail averages hide. B. Cost per task, the unit economics that decide viability. C. The number of tools per agent. D. The model’s name. E. Total prompt character count.
Answer: A and B. p95/p99 exposes the tail and cost-per-task surfaces the economics the team is missing. Tool count (C), model name (D) and character totals (E) do not reveal tail latency or cost.
Key takeaways
- The context window is a shared budget consumed by system, tools, history, tool results and thinking; keep the working set small.
- Context editing clears stale tool results; compaction summarises the narrative — neither preserves everything, so durable facts go to the memory tool or a database.
- Subagent isolation keeps bulky raw material out of the main window; put stable content first and cache it for ~10× cheaper reads and lower latency.
- Do the arithmetic: the accumulated-tool-results term usually dominates a filling window, so context editing (or subagent isolation) — not a bigger model — is the cost-effective fix.
- Fable 5.1 is append-only (never edit old turns), thinking is always on, forced
tool_choiceis a 400, it has 30-day retention/no ZDR, and it is not Priority Tier — design around all of these. - Retry only transient errors (429/5xx/529) with backoff, jitter and
retry-after; only retry idempotent writes. - Fallback is not free — from Fable 5.1 thinking blocks are dropped; Haiku 4.5 is 200k, not 1M; validate the degraded path.
- Size real-time throughput against ITPM/RPM (the token limit often binds first); use the Batch API for latency-tolerant bulk; set boundary timeouts and return partial results with provenance.
- Monitor p95/p99 latency, error rate by category, cache-hit rate and cost per task; gate irreversible actions on humans.
Last updated Sep 18, 2026