Domains
D3 · Integration (incl. RAG)
RAG pipeline design end to end, chunking and embeddings, retrieval and reranking, grounding and citations, retrieval evaluation and debugging, tool least-privilege and authz, MCP vs API selection, observability, and enterprise integration patterns.
This is the largest domain on the exam – 19%, roughly 12 of 63 items. It is dominated by RAG architecture but also covers tool/agent capability governance, identity and authorization, integration-mechanism selection, observability at scale, and enterprise integration patterns. The recurring judgement being tested: when an answer is confidently wrong, suspect retrieval and indexing before the model; and when an agent has many tools, apply least privilege by removing capabilities, not by logging or confirming them.
Learning objectives
By the end of this page you should be able to:
- Design a RAG pipeline stage by stage: ingestion → chunking → embedding → indexing → retrieval → rerank → grounding.
- Match a chunking strategy to the data’s shape.
- Choose dense vs sparse vs hybrid retrieval and an index/vector store with metadata filters.
- Improve retrieval with reranking, query rewriting, HyDE, multi-query, MMR.
- Evaluate retrieval (recall@k, MRR, faithfulness, answer relevance) and debug confident-but-wrong answers.
- Decide between RAG, long context (1M) and fine-tuning; design agentic RAG.
- Run capability-bloat / least-privilege analysis and close authn/authz gaps.
- Choose integration mechanisms (MCP vs API/CLI vs agent-to-agent), design observability at scale, and apply enterprise integration patterns (queues, webhooks, idempotency, batch, freshness).
3.1 The RAG pipeline, stage by stage
Retrieval-Augmented Generation grounds the model in your data by retrieving relevant passages and placing them in context at query time. Learn the pipeline as a fixed skeleton; every design decision slots into one stage.
INGESTION INDEX-BUILD (offline) QUERY-TIME (online) ┌──────────┐ ┌───────────────────────────────┐ ┌───────────────────────────────────┐ │ sources │ │ clean → chunk → embed → │ │ query → (rewrite/HyDE/multi-query) │ │ (docs, │──▶│ write vectors + metadata to │ │ → retrieve top-k (dense+sparse) │ │ DB, web)│ │ vector store / search index │ │ → rerank → assemble context │ └──────────┘ └───────────────────────────────┘ │ → Claude generates grounded answer │ │ freshness / refresh pipeline ▲ │ → cite sources → validate │ └──────────────────────────────┘ └───────────────────────────────────┘| Stage | Responsibility | Primary failure mode |
|---|---|---|
| Ingestion | Pull, clean, normalise, de-duplicate source data | Stale/duplicate content; lost structure |
| Chunking | Split into retrievable units | Chunks too large (dilute) or too small (fragment) |
| Embedding | Turn chunks into vectors | Wrong/mismatched embedding model |
| Indexing | Store vectors + metadata for search | Not re-indexed after refresh → stale hits |
| Retrieval | Fetch candidate chunks for a query | Low recall; wrong ranking |
| Rerank | Reorder candidates by true relevance | Skipped → best passage buried below top-k |
| Grounding | Constrain the answer to retrieved text + cite | Model answers from parametric memory, not context |
3.2 Chunking strategies matched to data shape
Chunking is the highest-leverage RAG decision. The right strategy depends on the structure of the source.
| Strategy | How it splits | Best for | Weakness |
|---|---|---|---|
| Fixed-size | N tokens with overlap | Uniform prose, quick baseline | Cuts mid-sentence/mid-idea |
| Recursive | Split on paragraph → sentence → token boundaries | General text with some structure | Still boundary-blind to meaning |
| Semantic | Split where embedding similarity drops | Topically dense text where ideas shift | More compute to build |
| Structural / document-aware | Split on headings, sections, table rows, code blocks | Manuals, contracts, Markdown, HTML, code | Requires a parser per format |
| Parent–child | Retrieve small child chunks, return larger parent for context | Precise matching + rich context to the model | More storage/bookkeeping |
| Late chunking | Embed the whole doc, then pool per-chunk from token embeddings | Long docs where cross-chunk context matters | Needs long-context embedding support |
Document shape? ├─ Structured (headings/tables/code) → structural / document-aware chunking ├─ Long, cross-referential prose → parent–child or late chunking ├─ Topically shifting prose → semantic chunking └─ Uniform prose, need a baseline → recursive (fall back to fixed)Exam signal
“Answers cite the wrong section / lose the surrounding context” → the chunks are too small or boundary-blind; move to parent–child or structural chunking. “Contracts / manuals / code” in the stem → document-aware chunking, not fixed-size.
3.3 Embeddings and indexing
Dense vs sparse vs hybrid
| Retrieval type | Signal | Strong at | Weak at |
|---|---|---|---|
| Dense (embeddings) | semantic similarity | paraphrase, synonyms, concepts | exact IDs, rare tokens, codes |
| Sparse (BM25 / keyword) | lexical overlap | exact terms, part numbers, names | synonyms, paraphrase |
| Hybrid (dense + sparse, fused) | both, score-fused | most enterprise corpora | slightly more infra |
Most production RAG uses hybrid retrieval (e.g. reciprocal-rank fusion of BM25 and dense) because real queries mix concepts and exact identifiers.
Vector store and metadata filters
| Choice driver | Guidance |
|---|---|
| Scale (billions of vectors) | Purpose-built vector DB or managed service |
| Already on a cloud | Use its managed vector/search offering (IAM, residency inherited) |
| Need lexical + vector in one | A search engine with hybrid support |
| Filtering | Store metadata (tenant, doc type, ACL, date) and filter before/along vector search |
Metadata filters are also a security control: filtering by tenant_id and per-user ACL at retrieval time is how you prevent cross-tenant/permission leakage (see 3.9).
3.4 Retrieval and re-ranking techniques
| Technique | What it does | When to add it |
|---|---|---|
| top-k | fetch the k nearest candidates | always; tune k for recall vs noise |
| MMR (max marginal relevance) | diversify results, reduce redundancy | near-duplicate chunks crowd top-k |
| Reranking | a cross-encoder re-scores candidates for true relevance | precision matters; retrieve wide, rerank to few |
| Query rewriting | rephrase/expand the query | conversational or underspecified queries |
| HyDE | generate a hypothetical answer, embed that to retrieve | sparse queries where the answer’s vocabulary differs from the question’s |
| Multi-query | issue several query variants, union results | recall-critical retrieval |
A common high-precision recipe: retrieve wide (k=50, hybrid) → rerank → keep top 5–8 → ground. Reranking is often the single biggest precision win.
query ─▶ [rewrite / multi-query] ─▶ hybrid retrieve (k=50) │ ▼ reranker ─▶ top 6 ─▶ context ─▶ Claude3.5 Grounding and citations
Grounding means the answer is constrained to the retrieved passages, and every claim is traceable.
- Instruct the model to answer only from the provided context and to say when the context is insufficient (never invent).
- Use the Citations feature / structured references so each claim links to its source chunk and location.
- Return
insufficient contextrather than a parametric-memory guess when retrieval fails — a grounded system prefers “I don’t have that” to a confident fabrication.
{ "system": "Answer ONLY from <context>. Cite the source id for each claim. If the context does not contain the answer, say so.", "messages": [ {"role": "user", "content": "<context>{retrieved_chunks_with_ids}</context>\n\nQuestion: What is the termination notice period?"} ]}3.6 Evaluating retrieval
You cannot fix what you do not measure. Retrieval and generation are evaluated separately so you know which stage failed.
| Metric | Measures | Stage |
|---|---|---|
| Recall@k | Did the relevant chunk make it into the top-k? | Retrieval |
| MRR | How high did the first relevant chunk rank? | Retrieval / rerank |
| Precision@k | How much of top-k is actually relevant? | Retrieval / rerank |
| Faithfulness / groundedness | Is the answer supported by retrieved context (no fabrication)? | Generation |
| Answer relevance | Does the answer address the question? | Generation |
Exam signal
If recall@k is high but answers are wrong, the failure is in generation/grounding (or reranking), not retrieval. If recall@k is low, fix chunking/embeddings/retrieval first. Measuring only end-to-end accuracy hides which stage broke — an aggregate-metric trap.
3.7 Debugging confident-but-wrong answers after a document refresh
This is a signature exam scenario. A document is updated; the assistant keeps giving the old answer, confidently. The instinct to blame the prompt or the model is wrong.
-
Suspect the index first. Was the refreshed document re-chunked, re-embedded and re-indexed? A refresh that updates the source store but not the vector index leaves stale vectors that retrieval faithfully returns.
-
Inspect what was retrieved. Log the retrieved chunk IDs and text for the failing query. If they are the old content, it’s an indexing/freshness bug, not a model bug.
-
Check chunk/version metadata. Stale
updated_ator a missing re-index job confirms it. -
Only then examine grounding/prompt. If retrieval returned the new content but the answer used the old, the grounding instruction or reranking is at fault.
Confident-but-wrong after a refresh? 1. What did retrieval return? ──stale── ▶ re-chunk/re-embed/re-index (fix freshness pipeline) 2. Returned fresh content? ──yes──── ▶ check grounding instruction / reranking 3. Neither? ────────── ▶ check embedding-model mismatch / query rewriteThe wrong instinct
Distractors will suggest “rewrite the system prompt”, “raise effort”, or “switch to a bigger model”. None fix stale retrieval. The correct first move is to inspect and repair the retrieval/indexing layer.
3.8 RAG vs long context vs fine-tuning
| Approach | Best when | Cost/ops profile | Fails when |
|---|---|---|---|
| RAG | Corpus is large, changes often, needs citations/freshness | Retrieval infra + refresh pipeline | Retrieval quality is poor |
| Long context (1M) | Small/bounded corpus fits the window; simplicity valued | High per-call token cost; no infra | Corpus too big/costly; needle degradation |
| Fine-tuning | Stable domain style/format/behaviour to bake in | Training + retraining cadence | Facts change often (retrain infeasible) |
Does the knowledge change frequently or need citations? ── yes ─▶ RAGIs the corpus small, stable, and fits comfortably in context? ── yes ─▶ long contextIs it about behaviour/format/style rather than changing facts? ── yes ─▶ fine-tuningThese are not mutually exclusive: fine-tune for style, RAG for facts, long context for a bounded session.
3.9 Agentic RAG, tool capability-bloat and least privilege
Agentic RAG lets the model decide when and what to retrieve, issue follow-up queries, and combine sources — more powerful than one-shot retrieval, but with more cost and failure surface. Use it when queries are multi-hop or exploratory; use plain RAG when a single retrieval suffices.
Capability-bloat and least privilege
An agent with too many tools (18 vs a recommended 4–5) is slower, more error-prone, and more dangerous. The correct remediation for an unneeded destructive tool is to remove it, not to log its use or add a confirmation prompt.
| Symptom | Wrong fix | Right fix |
|---|---|---|
Agent has refund, delete_account, issue_credit it never legitimately needs | Log the calls; add “are you sure?” | Remove the tools from the agent’s allowlist (least privilege) |
| 18 tools, model picks wrong ones | Longer prompt describing each | Cut to 4–5; use tool search + defer_loading for large catalogues |
| Occasional dangerous action | Prompt “never do X” | Programmatic permission hook denies X |
Least privilege is removal, not observation
Logging a dangerous capability or confirming it still leaves the capability present and reachable (excessive agency). The exam’s correct answer removes unneeded tools so the agent cannot invoke them at all.
Authn / authz gap analysis
Tools act on real systems, so identity and permissions must propagate to the tool call — the agent must act as the user, with the user’s permissions, not as an omnipotent service account.
| Gap | Risk | Control |
|---|---|---|
| Agent uses one service account for all users | One user reaches another’s data | Propagate end-user identity; per-user scoping |
| Tool has no per-user permission check | Privilege escalation | Enforce ACLs inside the tool, not just in the prompt |
| Remote MCP server unauthenticated | Anyone can call powerful tools | OAuth 2.1 on the remote MCP server |
| Retrieval ignores ACLs | Cross-tenant leakage | Metadata/ACL filter at retrieval time (3.3) |
3.10 Choosing the integration mechanism
| Mechanism | Use when | Notes |
|---|---|---|
| MCP | Reusable tools/resources shared across many agents/clients; standard protocol | JSON-RPC 2.0; stdio local / Streamable HTTP remote; OAuth 2.1 for remote; MCP connector lets Messages API call remote servers |
| Direct API / CLI tool | A one-off or app-specific capability; tightest control/latency | Define as a tool in the request; no protocol overhead |
| Agent-to-agent | Decompose across specialised agents with isolated context | Coordinator/subagent; higher cost/latency |
Will many clients/agents reuse this capability? ── yes ─▶ MCP server (OAuth if remote)One app, tight control/latency? ── yes ─▶ direct API/CLI toolNeed a specialised agent with its own context? ── yes ─▶ agent-to-agent (subagent)Progressive discovery vs monolithic context
Do not load every tool and document into context up front. Use progressive discovery: tool search + defer_loading: true for large tool catalogues, and Skills loaded on demand. A monolithic context is expensive, cache-hostile, and degrades tool-selection accuracy.
3.11 Observability at scale
At production scale you cannot debug what you cannot trace. Instrument every request end to end.
| Signal | What to capture |
|---|---|
| Traces | Full span tree: retrieval, rerank, model call, tool calls, validation |
| Correlation IDs | One ID threaded through app → model → tools → downstream systems |
| Token/cost telemetry | Input/output/thinking tokens and cost per request, per segment, per model |
| Retrieval telemetry | Query, retrieved chunk IDs, scores, rerank order |
| Errors & stop reasons | stop_reason, error codes, retries, fallbacks, degraded-mode flags |
[req id: 9f3a] ── app ──▶ retrieve(k=50) ──▶ rerank(6) ──▶ opus-5 ──▶ tool:lookup ──▶ validate ──▶ resp correlation id 9f3a threads through every span; token+cost tagged per spanPer-segment cost/latency telemetry is what lets you route, cache and optimise (D4) with evidence rather than guesses.
3.12 Enterprise integration patterns and data freshness
| Pattern | Purpose | Claude-specific note |
|---|---|---|
| Queue | Decouple bursty producers from rate-limited model calls | Smooths load against RPM/ITPM tiers |
| Webhook | React to external events (ticket created, doc updated) | Trigger ingestion/refresh and agent runs |
| Idempotency | Safe retries without duplicate side effects | Idempotency keys on tool actions; critical with backoff retries |
| Batch | Latency-tolerant bulk work | Message Batches API: 50% discount, results within 24 h |
| Refresh / freshness | Keep the index current | Event- or schedule-driven re-chunk/re-embed/re-index (ties to 3.7) |
Exam signal
“Retried request charged/acted twice” → idempotency keys. “Thousands of documents to classify overnight, cost matters” → Batch API. “Document changed but the answer didn’t” → freshness/refresh pipeline and re-indexing.
3.13 A concrete RAG pipeline configuration
Item writers reward candidates who can read a config and predict its behaviour. A defensible enterprise baseline:
ingestion: sources: [confluence, s3_pdfs, postgres_kb] dedup: content_hash pii_scrub: truechunking: strategy: structural # split on headings/sections fallback: recursive target_tokens: 400 overlap_tokens: 60 parent_child: true # match child (~400), return parent (~1500)embedding: model: text-embedding-3-large # keep query + index model identical dimensions: 1024index: store: managed_vector_db metadata: [tenant_id, doc_type, acl, updated_at, source_id] filters_before_search: [tenant_id, acl] # security + precisionretrieval: mode: hybrid # BM25 + dense, RRF fusion k: 50 rerank: model: cross_encoder_reranker keep_top: 6 query_transform: [rewrite, multi_query] # for conversational/sparse queriesgeneration: model: claude-sonnet-5 grounding: answer_only_from_context citations: true on_insufficient_context: "say so; do not use parametric memory"refresh: trigger: [webhook_on_update, nightly_schedule] action: re_chunk_re_embed_re_index| Knob | Effect if too low | Effect if too high |
|---|---|---|
target_tokens | fragments ideas; loses context | dilutes relevance; buries the answer |
overlap_tokens | boundary facts lost | storage/cost bloat, duplicate hits |
k (pre-rerank) | low recall | noise; slower rerank |
keep_top | best passage may be dropped | context bloat, higher cost, needle dilution |
Exam signal
“Retrieve wide (k≈50 hybrid) → rerank → keep 6 → ground with citations” is the canonical high-precision recipe. Distractors that skip reranking, use dense-only, or raise keep_top to 50 are the wrong answers.
3.14 Retrieval evaluation with worked numbers
You must be able to compute the metrics, not just name them.
Setup. 5 queries; for each we know the single relevant chunk and where it ranked in the retrieved list:
| Query | Rank of relevant chunk | In top-3? | Reciprocal rank |
|---|---|---|---|
| Q1 | 1 | yes | 1/1 = 1.00 |
| Q2 | 4 | no | 1/4 = 0.25 |
| Q3 | 2 | yes | 1/2 = 0.50 |
| Q4 | (not retrieved) | no | 0 |
| Q5 | 1 | yes | 1/1 = 1.00 |
recall@3 = (#queries whose relevant chunk is in top-3) / total = 3/5 = 0.60MRR = mean reciprocal rank = (1.00 + 0.25 + 0.50 + 0 + 1.00) / 5 = 0.55Now suppose adding a reranker moves Q2’s relevant chunk to rank 2 and retrieves Q4’s at rank 3:
recall@3 → 5/5 = 1.00 (both now in top-3)MRR → (1.00 + 0.50 + 0.50 + 0.333 + 1.00)/5 = 0.667The reranker lifted recall@3 from 0.60 → 1.00 and MRR from 0.55 → 0.67 — the single biggest precision/ranking win, without touching chunking or embeddings.
| Metric | Formula | Reads as |
|---|---|---|
| recall@k | relevant-in-top-k ÷ total queries | “did we retrieve the answer at all?” |
| MRR | mean(1 ÷ rank of first relevant) | “how high did it rank?” |
| precision@k | relevant-in-top-k ÷ k | “how clean is the top-k?” |
| faithfulness | supported claims ÷ total claims | “did the answer stay grounded?” |
Interpreting the numbers
High recall but low faithfulness ⇒ retrieval is fine, generation/grounding is broken. Low recall ⇒ fix chunking/embeddings/retrieval first. Reporting only end-to-end accuracy hides which of these is true — the aggregate-metric trap.
3.15 Multi-tenant isolation and ACL-aware retrieval
In a shared index, retrieval is a security boundary. A query must never surface another tenant’s or an unauthorised user’s chunk.
user (tenant_42, role: support) asks a question │ attach identity: {tenant_id: 42, acl_groups: [support]} ▼filter BEFORE/ALONGSIDE vector search: WHERE tenant_id = 42 AND acl IN ('public','support') ▼hybrid retrieve → rerank → ground (only authorised chunks ever enter context)| Gap | Leak | Control |
|---|---|---|
No tenant_id filter | Cross-tenant data exposure | Mandatory pre-filter on tenant_id |
| ACL applied only in the prompt | Model can be talked past it | Enforce ACL at the retrieval layer, not in prose |
| Shared service account for retrieval | User sees data they can’t access | Propagate end-user identity into the query filter |
| Rerank/cache ignores tenant | Cached cross-tenant hit | Key caches by tenant + ACL |
Retrieval-time filtering is the control
Filtering after generation (“the model won’t mention it”) is not a control — the unauthorised chunk already entered context and can leak. The exam’s correct answer filters by tenant_id/ACL before the vector search.
3.16 Scenario walkthrough: a leaking, stale, over-privileged support RAG agent
Scenario. A B2B SaaS support agent serves 300 tenants from one shared vector index. Three complaints arrive: (1) a tenant occasionally sees another tenant’s runbook in an answer; (2) after customers update their docs, the agent still quotes yesterday’s version, confidently; (3) the agent has 22 tools including delete_ticket and refund it should never call. Retrieval is dense-only, k=8, no reranker.
Expert reasoning trace.
-
Isolation first (highest severity). Cross-tenant exposure is a security incident. Add a mandatory
tenant_id(and ACL) pre-filter on every query, key caches by tenant, and propagate end-user identity. Do not rely on prompt wording to keep tenants apart. -
Freshness second. Confident-wrong-after-update is stale retrieval: the refresh updated the source store but not the vector index. Inspect retrieved chunk IDs (they’ll be old), then fix the re-chunk/re-embed/re-index pipeline with webhook + nightly triggers.
-
Least privilege third. Remove
delete_ticket,refund, and other unneeded tools from the allowlist — do not merely log or add “are you sure?”. Cut to ~4–5 tools; use tool search +defer_loadingif the legitimate catalogue is large. -
Then quality. Dense-only + k=8 + no rerank underperforms on exact identifiers and precision. Move to hybrid retrieval, k=50, rerank, keep top 6, and measure recall@k and faithfulness per tenant segment.
-
Close the loop. Instrument correlation IDs and per-segment retrieval telemetry so the next incident is reconstructable.
Why the tempting alternatives are wrong: “add a system-prompt rule not to reveal other tenants” is prompt-as-enforcement and leaves the leak; “switch to a bigger model” fixes neither isolation nor freshness; “log the dangerous tools” leaves excessive agency; “raise k to 500” adds noise instead of adding a reranker.
3.17 Common misconceptions
| Misconception | Reality | Why it matters on the exam |
|---|---|---|
| “Confident-but-wrong means the model is bad.” | After a data change it usually means stale retrieval/indexing. | Inspect retrieval first; prompt/model fixes are distractors. |
| “Dense embeddings retrieve everything.” | Dense is weak on exact IDs/codes; hybrid adds sparse. | Part-number/SKU stems require hybrid. |
| “A bigger context window replaces RAG.” | It costs more per call, can’t cite, and degrades on the needle. | Long-context-stuffing is the wrong answer for large/changing corpora. |
| “Fine-tuning is how you add knowledge.” | Fine-tuning bakes behaviour/style; changing facts need RAG. | Weekly-changing-facts stems reject fine-tuning. |
| “Logging a dangerous tool makes it safe.” | The capability is still reachable — excessive agency. | Least privilege means removal, not observation. |
| “Prompt rules keep tenants isolated.” | Isolation must be enforced at the retrieval filter. | Retrieval-time ACL filtering is the correct control. |
| “MCP servers are authenticated by default.” | Remote MCP needs OAuth 2.1 + per-user checks. | Unauthenticated remote MCP is a security trap. |
| “Retrying is always safe.” | Non-idempotent actions double up; add idempotency keys. | Double-charge stems test idempotency. |
Exam traps in this domain
| Trap | Why it is wrong |
|---|---|
| Blaming the prompt/model for confident-wrong answers after a refresh | The cause is usually stale retrieval/indexing; inspect what was retrieved first |
| Fixing an unneeded dangerous tool by logging or confirming it | Leaves excessive agency; least privilege means removing the tool |
| Giving an agent 18 tools “for flexibility” | Slower, error-prone; cut to 4–5, use tool search + defer_loading |
| Using dense-only retrieval for part numbers / exact IDs | Dense misses exact tokens; use sparse/hybrid |
| Measuring only end-to-end accuracy | Hides whether retrieval or generation failed; measure recall@k and faithfulness separately |
| Long-context stuffing a large, changing corpus | Costly, no citations, needle degradation; use RAG |
| Fine-tuning to inject frequently-changing facts | Retraining cadence infeasible; use RAG |
| Skipping reranking | Best passage buried below top-k; precision suffers |
| One service account for all users’ tool calls | Cross-user data exposure; propagate identity + per-user ACLs |
| Unauthenticated remote MCP server | Anyone can invoke powerful tools; require OAuth 2.1 |
| Retrying non-idempotent tool actions without keys | Duplicate side effects (double refund/charge) |
| Loading all tools/docs into context up front | Expensive, cache-hostile, worse tool selection; use progressive discovery |
| Enforcing multi-tenant isolation with a system-prompt rule | Isolation must be a retrieval-time tenant_id/ACL filter, not prose |
| Filtering unauthorised chunks after generation | The chunk already entered context; filter before the vector search |
| Raising k to hundreds instead of adding a reranker | Adds noise; reranking is the precision/ranking win |
| Reporting only end-to-end accuracy for a RAG system | Hides whether retrieval or grounding failed; measure recall@k and faithfulness |
| Using the same embedding model for query and index inconsistently | Query/index model mismatch wrecks similarity; keep them identical |
| Caching retrieval results without keying by tenant/ACL | Risks serving a cross-tenant cached hit |
Practice questions
Q1 · A policy assistant kept answering with the OLD figure after a policy document was updated last night, and it sounds completely confident. What should the architect investigate FIRST? (Select one)
A. Rewrite the system prompt to be more accurate. B. Inspect what retrieval actually returned; the updated document was likely not re-chunked/re-embedded/re-indexed, so retrieval is serving stale vectors. C. Switch to Opus 5 with xhigh effort. D. Add more few-shot examples.
Answer: B. Confident-but-wrong immediately after a refresh points to stale retrieval/indexing, not the model. The first move is to log the retrieved chunks and confirm whether the refresh updated the vector index. Prompt rewrites (A, D) and a bigger model (C) cannot fix stale retrieval.
Q2 · A support agent has 18 tools, including `delete_account` and `issue_refund`, which its role should never use. What is the correct remediation? (Select one)
A. Keep the tools but log every call for audit. B. Remove the unneeded tools from the agent’s allowlist (least privilege) so it cannot invoke them at all. C. Add a confirmation prompt before those tools run. D. Add a system-prompt sentence forbidding their use.
Answer: B. Least privilege means the capability should not be present. Removing the tools eliminates the excessive agency. Logging (A) and confirmation (C) leave the capability reachable; a prompt rule (D) is prompt-as-enforcement and can be bypassed.
Q3 · Users search a parts catalogue by exact part numbers AND by descriptions. Dense-only retrieval misses many exact-number queries. What is the BEST fix? (Select one)
A. Increase k to 500. B. Use hybrid retrieval (BM25 + dense with rank fusion) so exact identifiers and semantic matches both rank well. C. Switch to a bigger generation model. D. Remove metadata filters.
Answer: B. Dense embeddings are weak on exact tokens/codes; sparse (BM25) handles them, and hybrid fusion covers both query types. A huge k (A) adds noise without fixing lexical matching; a bigger model (C) doesn’t change what is retrieved; removing filters (D) hurts precision and security.
Q4 · Retrieval eval shows recall@10 = 0.95 but faithfulness is low and answers include facts not in the retrieved chunks. Where is the failure and the fix? (Select one)
A. Retrieval; lower k. B. Generation/grounding; strengthen the instruction to answer only from context, add citations, and consider reranking so the best passage is on top. C. Embeddings; change the model. D. Indexing; re-index everything.
Answer: B. High recall means the right chunks are retrieved, so the fault is in generation/grounding — the model is answering from parametric memory. Grounding instructions, citations and reranking address it. The retrieval-side fixes (A, C, D) target a stage that is already performing well.
Q5 · A 2M-document knowledge base changes weekly and answers must cite the exact source clause. Which approach is BEST? (Select one)
A. Fine-tune a model on the corpus each week. B. RAG with hybrid retrieval, reranking and citations, plus a scheduled refresh pipeline. C. Stuff the whole corpus into the 1M context per query. D. Long context plus fine-tuning combined.
Answer: B. Large, frequently-changing, citation-requiring corpora are the canonical RAG case. Weekly fine-tuning (A) is an infeasible retraining cadence for facts. The corpus exceeds the window and would cost too much and lose citations (C). (D) inherits both problems.
Q6 · A contract-QA system chunks contracts every 500 tokens with fixed size. Answers cite the wrong sub-clause and lose surrounding context. Which TWO changes help most? (Select two)
A. Use structural/document-aware chunking that splits on clauses/sections. B. Use parent–child chunking: match on small child chunks but return the larger parent clause for context. C. Switch to dense-only retrieval. D. Increase temperature. E. Remove citations.
Answer: A and B. Contracts are structured, so clause/section-aware chunking preserves boundaries, and parent–child gives precise matching with enough surrounding context. Dense-only (C), temperature (D) and removing citations (E) do not address the chunking problem — the last two make it worse.
Q7 · A remote MCP server exposes powerful tools over Streamable HTTP with no auth. What must the architect add? (Select one)
A. Nothing; MCP is safe by default. B. OAuth 2.1 authentication on the remote MCP server, plus per-user permission checks inside the tools. C. A longer system prompt. D. A higher rate-limit tier.
Answer: B. Remote MCP servers require OAuth 2.1, and tools must enforce per-user permissions so the agent acts with the caller’s authority. MCP is not authenticated by default (A); prompts (C) and rate limits (D) do not address authorization.
Q8 · A nightly job must classify 200,000 documents; latency is not important but cost is. Which mechanism is BEST? (Select one)
A. Real-time synchronous calls in a tight loop. B. The Message Batches API for a 50% discount with results within 24 hours. C. A multi-agent system. D. Fine-tuning.
Answer: B. Latency-tolerant bulk work is exactly the Batch API’s use case (50% discount, results within 24 h). Synchronous loops (A) hit rate limits and cost more. Multi-agent (C) adds cost/complexity; fine-tuning (D) is unrelated to a classification batch.
Q9 · After enabling backoff retries, some refunds are issued twice. What is the correct fix? (Select one)
A. Disable retries entirely. B. Add idempotency keys to the refund tool so retried calls are de-duplicated and produce no duplicate side effect. C. Lower the model temperature. D. Log the duplicates and reconcile later.
Answer: B. Retries on non-idempotent actions cause duplicate side effects; idempotency keys make retries safe. Disabling retries (A) harms resilience. Temperature (C) is irrelevant. Reconciling after the fact (D) still charged customers twice.
Q10 · A single retrieval sometimes isn't enough — some questions need follow-up lookups combining multiple sources. Which design fits, and what is the trade-off? (Select one)
A. Agentic RAG, where the model decides when/what to retrieve and can issue follow-up queries, at the cost of more latency and tokens. B. Stuff everything into context to avoid retrieval. C. Fine-tune on the multi-hop questions. D. Remove reranking to speed things up.
Answer: A. Multi-hop, exploratory queries justify agentic RAG’s model-driven retrieval loop; the trade-off is higher cost and latency versus one-shot RAG. Context stuffing (B) doesn’t scale; fine-tuning (C) can’t hold changing facts; removing reranking (D) hurts precision.
Q11 · A capability will be reused by many agents and clients across the company and must follow a standard protocol. Which integration mechanism is BEST? (Select one)
A. Hard-code it as a per-app CLI tool in each service. B. Build it as an MCP server (with OAuth 2.1 if remote) so many clients reuse it via a standard protocol. C. Implement it as an agent-to-agent handoff only. D. Paste its logic into every system prompt.
Answer: B. Reuse across many clients under a standard protocol is MCP’s purpose. Per-app CLI tools (A) fragment the implementation; agent-to-agent (C) is for specialised context isolation, not shared capability; prompt-pasting (D) is unmaintainable.
Q12 · An agent loads all 40 tools and 30 reference documents into context on every request; cost is high, cache hit rate is low, and tool selection is error-prone. What is the BEST remedy? (Select two)
A. Use tool search with defer_loading: true so only relevant tools are loaded.
B. Package reference material as Skills loaded progressively on demand.
C. Increase the context window to 1M and keep loading everything.
D. Escalate every request to Opus 5.
E. Disable prompt caching.
Answer: A and B. Progressive discovery — tool search with deferred loading and on-demand Skills — keeps context lean, restores a stable cache prefix, and improves tool selection. A bigger window (C) still pays for the bloat; escalating models (D) raises cost; disabling caching (E) is the opposite of the fix.
Q13 · A B2B assistant serves 300 tenants from one shared vector index; a tenant occasionally sees another tenant's document in an answer. What is the correct control? (Select one)
A. Add a system-prompt rule telling the model not to reveal other tenants’ data.
B. Apply a mandatory tenant_id (and ACL) filter at retrieval time, before/alongside the vector search, so only authorised chunks ever enter context.
C. Filter the answer after generation to remove other tenants’ data.
D. Give each tenant a bigger model.
Answer: B. Isolation is a retrieval-time security boundary: filter by tenant_id/ACL before the vector search. A prompt rule (A) is bypassable prompt-as-enforcement; post-generation filtering (C) is too late — the chunk already entered context; a bigger model (D) doesn’t isolate data.
Q14 · Across 5 queries the relevant chunk ranked 1, 4, 2, not-retrieved, 1. What are recall@3 and MRR? (Select one)
A. recall@3 = 1.00; MRR = 1.00. B. recall@3 = 0.60; MRR = 0.55. C. recall@3 = 0.55; MRR = 0.60. D. recall@3 = 0.80; MRR = 0.70.
Answer: B. Three of five relevant chunks are in the top 3 (ranks 1, 2, 1) → recall@3 = 3/5 = 0.60. Reciprocal ranks are 1, 0.25, 0.5, 0, 1 → MRR = 2.75/5 = 0.55. Option A ignores the misses; C swaps the two values; D is arithmetically wrong.
Q15 · recall@10 = 0.62 (low) and answers are frequently missing the needed fact. Where should the architect work FIRST? (Select one)
A. Grounding; tighten the answer-only-from-context instruction. B. Retrieval: fix chunking/embeddings/hybrid and add reranking, because low recall means the relevant chunk often isn’t retrieved at all. C. Add more citations. D. Switch the generation model to Opus 5.
Answer: B. Low recall means the right chunk isn’t reaching the top-k, so the failure is upstream in retrieval. Grounding/citations (A, C) and a bigger generation model (D) can’t help if the answer was never retrieved.
Q16 · A RAG system indexes with `text-embedding-3-large` but a new service queries with a different embedding model. Similarity scores look random. What is the cause and fix? (Select one)
A. The vector store is corrupt; rebuild hardware. B. Query/index embedding-model mismatch; use the identical embedding model for both indexing and querying. C. k is too low; raise it to 1000. D. The generation model is too small.
Answer: B. Embeddings from different models live in different vector spaces, so cross-model similarity is meaningless; query and index must use the same embedding model. It isn’t hardware (A); raising k (C) can’t fix incompatible vectors; the generation model (D) is unrelated to retrieval similarity.
Q17 · An agent retrieves k=8 with dense-only and no reranker; precision is poor and exact part numbers are missed. Which TWO changes give the biggest quality lift? (Select two)
A. Switch to hybrid retrieval (BM25 + dense) so exact identifiers rank. B. Retrieve wide (k≈50) and add a reranker, keeping the top 6. C. Increase temperature. D. Remove citations to speed responses. E. Move to a 1M-token context and stuff everything.
Answer: A and B. Hybrid retrieval fixes exact-identifier misses and retrieve-wide-then-rerank fixes precision — the two canonical levers. Temperature (C) is irrelevant to retrieval; removing citations (D) harms grounding traceability; stuffing context (E) inflates cost without improving ranking.
Q18 · Which TWO signals let you reconstruct a failed multi-step RAG request end to end? (Select two)
A. A correlation ID threaded through app → retrieval → model → tools → downstream.
B. Per-span traces capturing retrieved chunk IDs, rerank order, token/cost, and stop_reason.
C. Only the final HTTP status code.
D. The model’s own assessment that it did fine.
E. A daily aggregate request count.
Answer: A and B. A correlation ID plus per-span traces (with retrieval detail and cost) make an incident reconstructable. A status code (C) and a daily count (E) are too coarse; self-assessment (D) is unreliable (self-report anti-pattern).
Key takeaways
- RAG is a fixed pipeline: ingest → chunk → embed → index → retrieve → rerank → ground → cite; every decision slots into a stage.
- Match chunking to data shape: document-aware for structured text, parent–child/late for cross-referential prose, semantic for shifting topics.
- Use hybrid retrieval (dense + sparse) for real corpora; rerank for precision; add rewriting/HyDE/multi-query for recall.
- Ground answers to retrieved context, cite sources, and prefer “insufficient context” over fabrication.
- Evaluate retrieval and generation separately (recall@k, MRR vs faithfulness); confident-wrong-after-refresh means inspect retrieval/indexing first.
- Choose RAG for changing/cited facts, long context for small stable corpora, fine-tuning for behaviour/style.
- Apply least privilege by removing unneeded (especially destructive) tools; keep agents to ~4–5 tools with tool search for larger catalogues.
- Propagate user identity to tools, enforce per-user ACLs, and require OAuth 2.1 on remote MCP servers.
- Instrument traces, correlation IDs and token/cost telemetry; use queues, idempotency keys, webhooks, Batch API and a freshness pipeline for enterprise integration.
- Compute retrieval metrics: recall@k = relevant-in-top-k ÷ queries; MRR = mean(1÷rank); a reranker typically lifts both the most.
- In a shared index, retrieval is a security boundary: filter by
tenant_id/ACL before the vector search, propagate end-user identity, and key caches by tenant. - Keep the query and index embedding models identical; the canonical recipe is retrieve wide (k≈50 hybrid) → rerank → keep ~6 → ground with citations.
Last updated Sep 18, 2026