AI Cert Prep
Type to search documentation.

Appendix · Claude

RAG Cookbook

Retrieval-augmented generation recipes — chunking parameters, hybrid retrieval config, reranking, query rewriting, eval metrics with worked numbers, and a debugging playbook.

RAG grounds answers in retrieved chunks at query time. Reach for it when the corpus is large or changing and you need citations. Prefer long context for a small stable corpus that fits the window, and fine-tuning for style/format, not facts.

Exam signal

“Confident but wrong after a document refresh” → suspect retrieval/indexing first (stale index, chunk drift), not the prompt or model. Changing the prompt is the distractor.

When RAG vs alternatives

SituationChooseWhy
Thousands of docs, updated often, must citeRAGFresh, grounded, auditable
50-page handbook, rarely changes, fits 1M windowLong contextNo retrieval infra; simplest
“Match our tone / always this format”Fine-tuning or few-shotBehaviour, not facts
Model must decide when/what to retrieveAgentic RAG (retrieval as a tool)Iterative reformulation

Chunking recipes

StrategyParameters (start here)Best forTrade-off
Fixed-size512 tokens, 50-token overlapUniform proseSplits mid-idea
Recursive (by separator)400–800 tokens, split on \n\n → \n → sentenceMixed proseUneven sizes
SemanticBreak at embedding-similarity dipsTopic-shifting docsCompute cost
Structural / document-awareOne chunk per section/headingManuals, contracts, wikisNeeds good structure
Parent-child (small-to-big)Child 150–300 tok for match; return parent 800–1200 tokPrecise hit + broad contextTwo-tier store
Late chunkingEmbed full doc, pool per chunkChunks needing doc-level contextRequires long-context embedder

Overlap rule of thumb: 10–20% of chunk size. Too little loses boundary context; too much inflates the index and duplicates hits.

Chunk drift

When documents are re-ingested with a different chunker or size, retrieval quality shifts silently. Version your chunking config and re-run retrieval evals after any change.

Hybrid retrieval config

Dense (embeddings) catches paraphrase; sparse (BM25) catches exact IDs, codes and rare terms. Fuse them, then rerank.

yaml
retrieval:
dense:
model: text-embedding-large
top_k: 40
sparse:
algorithm: bm25
top_k: 40
fusion:
method: reciprocal_rank_fusion # RRF
k: 60 # RRF constant
rerank:
model: cross-encoder-reranker
top_n: 8 # feed the LLM the top 8
diversity:
method: mmr # optional, reduce redundancy
lambda: 0.5

Reciprocal Rank Fusion: score of a doc = Σ over lists of 1 / (k + rank). With k=60, a doc ranked #1 in both lists scores 1/61 + 1/61 ≈ 0.0328; a doc ranked #1 dense but absent in sparse scores 1/61 ≈ 0.0164. RRF needs no score calibration between systems — it ranks by position.

Reranking

A cross-encoder reranker reads the query and each candidate together, giving far better precision than the bi-encoder used for first-stage retrieval — at higher cost. Retrieve broadly (top_k 40–100), rerank to a small top_n (5–10) for the LLM.

Bi-encoder (retrieval)Cross-encoder (rerank)
EncodesQuery and docs separately (precomputed)Query+doc jointly, per pair
SpeedFast, index-timeSlow, query-time
PrecisionGood recallBest precision
RoleFirst-stage, top-kSecond-stage, top-n

Query rewriting

TechniqueWhat it doesUse when
Rewriting / normalisationClean, expand abbreviations, resolve pronounsChat follow-ups (“and the second one?”)
Multi-queryGenerate several paraphrases, retrieve for each, unionVague or broad questions
HyDEGenerate a hypothetical answer, embed it, retrieve by similaritySparse queries where the answer’s vocabulary differs
DecompositionSplit a multi-part question into sub-queriesCompound questions

Eval metrics with worked numbers

Separate retrieval quality from generation quality — a good answer can hide bad retrieval and vice-versa.

MetricMeasuresFormula / example
Recall@kDid we retrieve the relevant chunk in top-k?8 of 10 queries had their gold chunk in top-5 → Recall@5 = 0.80
MRRRank of first relevant hitRanks 1, 3, 2 → (1 + 1/3 + 1/2)/3 = 0.611
Precision@kFraction of top-k that are relevant2 relevant of top-5 → 0.40
nDCG@kRank-weighted relevanceRewards relevant hits near the top
FaithfulnessIs the answer supported by retrieved context?47 of 50 claims grounded → 0.94
Answer relevanceDoes the answer address the question?Human/LLM-judge rubric
Context precisionAre retrieved chunks actually used?Low → retrieving noise

Worked comparison — swapping in a reranker on a 200-query eval:

ConfigRecall@5MRRFaithfulnessAnswer relevance
Dense only0.710.520.860.83
Hybrid (RRF)0.830.610.900.86
Hybrid + reranker0.830.740.940.90

Reranking barely moves Recall@5 (same candidates) but lifts MRR and faithfulness sharply by ordering the right chunk first — the LLM sees it earlier.

Debugging playbook

  1. Reproduce with the trace. Log the query, rewritten query, retrieved chunk IDs+scores, reranked order, and the final prompt. Most “wrong answer” bugs are visible here.
  2. Is the gold chunk retrieved at all? If not in top-k → retrieval problem: check chunking (drift), embeddings (wrong model/dim), or query rewriting.
  3. Retrieved but ranked low? Add/adjust the reranker; tune RRF k; verify hybrid weighting.
  4. Retrieved and top-ranked but answer wrong? Now it is generation: check the prompt (are chunks actually passed?), context ordering (docs first), and faithfulness (is the model ignoring context?).
  5. Confident-but-wrong after a refresh? Suspect the index: re-embedded with a different model, changed chunker, or stale vectors. Re-run retrieval evals.
  6. Per-segment check. Break metrics down by document type/source; an aggregate 90% can hide one source at 40%.

Common misconceptions

MisconceptionRealityWhy it matters on the exam
“More context always helps”Noise lowers faithfulness; rerank to a small top-nOver-retrieval distractor
“Dense retrieval covers everything”It misses exact IDs/codes; add BM25Hybrid-search signal
“Wrong answer means fix the prompt”Check retrieval first (drift, index)Root-cause distractor
“Aggregate faithfulness of 0.9 is fine”One source may be failing; report per-segmentAggregate-metric anti-pattern
“Fine-tune to add new facts”Fine-tuning is for style/format; RAG for factsRAG-vs-fine-tuning distractor
“Reranking improves recall”It improves ordering/precision, not recallMetric-confusion distractor
“Bigger chunks = better context”They dilute matches; use parent-childChunking distractor

Scenario walkthrough

A support assistant over 40,000 KB articles (updated weekly) started giving confident, wrong answers this week. Users ask both natural-language questions and exact error codes. What is the FIRST step and the right architecture?

  1. FIRST step — pull traces and check whether the gold chunk is retrieved. It surfaced that this week’s re-ingest changed the chunker (drift) — a retrieval problem, not a prompt problem.
  2. Chunking — structural/parent-child for articles: child chunks for precise matches, parent sections for context. Version the config.
  3. Retrieval — hybrid (dense + BM25): error codes need exact-term matching that dense misses.
  4. Reranking — cross-encoder to top-8; the gold article was retrieved but ranked #14 pre-rerank.
  5. Eval — recall@k, MRR, faithfulness, per source; gate re-ingests on these.
  6. Citations — attach spans so answers are auditable (provenance test).

Rejected alternatives: rewriting the system prompt (the retrieval was broken), fine-tuning on the KB (facts change weekly — RAG’s job), dense-only retrieval (misses error codes), and trusting the aggregate score (masked the failing source).

Key takeaways

  • RAG for large/changing corpora + citations; long context for small stable; fine-tuning for style, not facts.
  • Chunk to match data shape and query pattern; parent-child balances precise hits with context; version the config.
  • Hybrid (dense + BM25) + cross-encoder reranking beats either alone; RRF fuses without score calibration.
  • Separate retrieval metrics (recall@k, MRR) from generation metrics (faithfulness, answer relevance), and report per-segment.
  • Debug from the trace: gold chunk retrieved? ranked? passed? used? Suspect the index first after a refresh.

Last updated Sep 18, 2026