# RAG Cookbook

Retrieval-augmented generation recipes — chunking parameters, hybrid retrieval config, reranking, query rewriting, eval metrics with worked numbers, and a debugging playbook.

import { Steps } from '@prosefly/astro-components';

RAG grounds answers in retrieved chunks at query time. Reach for it when the corpus is **large or changing** and you need **citations**. Prefer long context for a small stable corpus that fits the window, and fine-tuning for **style/format**, not facts.

:::tip[Exam signal]
"Confident but wrong after a document refresh" → suspect **retrieval/indexing first** (stale index, chunk drift), not the prompt or model. Changing the prompt is the distractor.
:::

## When RAG vs alternatives

| Situation | Choose | Why |
| --- | --- | --- |
| Thousands of docs, updated often, must cite | RAG | Fresh, grounded, auditable |
| 50-page handbook, rarely changes, fits 1M window | Long context | No retrieval infra; simplest |
| "Match our tone / always this format" | Fine-tuning or few-shot | Behaviour, not facts |
| Model must decide when/what to retrieve | Agentic RAG (retrieval as a tool) | Iterative reformulation |

## Chunking recipes

| Strategy | Parameters (start here) | Best for | Trade-off |
| --- | --- | --- | --- |
| Fixed-size | 512 tokens, 50-token overlap | Uniform prose | Splits mid-idea |
| Recursive (by separator) | 400–800 tokens, split on `\n\n` → `\n` → sentence | Mixed prose | Uneven sizes |
| Semantic | Break at embedding-similarity dips | Topic-shifting docs | Compute cost |
| Structural / document-aware | One chunk per section/heading | Manuals, contracts, wikis | Needs good structure |
| Parent-child (small-to-big) | Child 150–300 tok for match; return parent 800–1200 tok | Precise hit + broad context | Two-tier store |
| Late chunking | Embed full doc, pool per chunk | Chunks needing doc-level context | Requires long-context embedder |

**Overlap** rule of thumb: 10–20% of chunk size. Too little loses boundary context; too much inflates the index and duplicates hits.

:::caution[Chunk drift]
When documents are re-ingested with a different chunker or size, retrieval quality shifts silently. Version your chunking config and re-run retrieval evals after any change.
:::

## Hybrid retrieval config

Dense (embeddings) catches paraphrase; sparse (BM25) catches exact IDs, codes and rare terms. Fuse them, then rerank.

```yaml
retrieval:
  dense:
    model: text-embedding-large
    top_k: 40
  sparse:
    algorithm: bm25
    top_k: 40
  fusion:
    method: reciprocal_rank_fusion   # RRF
    k: 60                            # RRF constant
  rerank:
    model: cross-encoder-reranker
    top_n: 8                         # feed the LLM the top 8
  diversity:
    method: mmr                      # optional, reduce redundancy
    lambda: 0.5
```

**Reciprocal Rank Fusion**: score of a doc = Σ over lists of `1 / (k + rank)`. With `k=60`, a doc ranked #1 in both lists scores `1/61 + 1/61 ≈ 0.0328`; a doc ranked #1 dense but absent in sparse scores `1/61 ≈ 0.0164`. RRF needs no score calibration between systems — it ranks by position.

## Reranking

A **cross-encoder** reranker reads the query and each candidate *together*, giving far better precision than the bi-encoder used for first-stage retrieval — at higher cost. Retrieve broadly (`top_k` 40–100), rerank to a small `top_n` (5–10) for the LLM.

| | Bi-encoder (retrieval) | Cross-encoder (rerank) |
| --- | --- | --- |
| Encodes | Query and docs separately (precomputed) | Query+doc jointly, per pair |
| Speed | Fast, index-time | Slow, query-time |
| Precision | Good recall | Best precision |
| Role | First-stage, top-k | Second-stage, top-n |

## Query rewriting

| Technique | What it does | Use when |
| --- | --- | --- |
| Rewriting / normalisation | Clean, expand abbreviations, resolve pronouns | Chat follow-ups ("and the second one?") |
| Multi-query | Generate several paraphrases, retrieve for each, union | Vague or broad questions |
| HyDE | Generate a hypothetical answer, embed it, retrieve by similarity | Sparse queries where the answer's vocabulary differs |
| Decomposition | Split a multi-part question into sub-queries | Compound questions |

## Eval metrics with worked numbers

Separate **retrieval** quality from **generation** quality — a good answer can hide bad retrieval and vice-versa.

| Metric | Measures | Formula / example |
| --- | --- | --- |
| Recall@k | Did we retrieve the relevant chunk in top-k? | 8 of 10 queries had their gold chunk in top-5 → Recall@5 = 0.80 |
| MRR | Rank of first relevant hit | Ranks 1, 3, 2 → (1 + 1/3 + 1/2)/3 = 0.611 |
| Precision@k | Fraction of top-k that are relevant | 2 relevant of top-5 → 0.40 |
| nDCG@k | Rank-weighted relevance | Rewards relevant hits near the top |
| Faithfulness | Is the answer supported by retrieved context? | 47 of 50 claims grounded → 0.94 |
| Answer relevance | Does the answer address the question? | Human/LLM-judge rubric |
| Context precision | Are retrieved chunks actually used? | Low → retrieving noise |

Worked comparison — swapping in a reranker on a 200-query eval:

| Config | Recall@5 | MRR | Faithfulness | Answer relevance |
| --- | --- | --- | --- | --- |
| Dense only | 0.71 | 0.52 | 0.86 | 0.83 |
| Hybrid (RRF) | 0.83 | 0.61 | 0.90 | 0.86 |
| Hybrid + reranker | 0.83 | 0.74 | 0.94 | 0.90 |

Reranking barely moves Recall@5 (same candidates) but lifts MRR and faithfulness sharply by ordering the *right* chunk first — the LLM sees it earlier.

## Debugging playbook

<Steps>
1. **Reproduce with the trace.** Log the query, rewritten query, retrieved chunk IDs+scores, reranked order, and the final prompt. Most "wrong answer" bugs are visible here.
2. **Is the gold chunk retrieved at all?** If not in top-k → retrieval problem: check chunking (drift), embeddings (wrong model/dim), or query rewriting.
3. **Retrieved but ranked low?** Add/adjust the reranker; tune RRF `k`; verify hybrid weighting.
4. **Retrieved and top-ranked but answer wrong?** Now it is generation: check the prompt (are chunks actually passed?), context ordering (docs first), and faithfulness (is the model ignoring context?).
5. **Confident-but-wrong after a refresh?** Suspect the index: re-embedded with a different model, changed chunker, or stale vectors. Re-run retrieval evals.
6. **Per-segment check.** Break metrics down by document type/source; an aggregate 90% can hide one source at 40%.
</Steps>

## Common misconceptions

| Misconception | Reality | Why it matters on the exam |
| --- | --- | --- |
| "More context always helps" | Noise lowers faithfulness; rerank to a small top-n | Over-retrieval distractor |
| "Dense retrieval covers everything" | It misses exact IDs/codes; add BM25 | Hybrid-search signal |
| "Wrong answer means fix the prompt" | Check retrieval first (drift, index) | Root-cause distractor |
| "Aggregate faithfulness of 0.9 is fine" | One source may be failing; report per-segment | Aggregate-metric anti-pattern |
| "Fine-tune to add new facts" | Fine-tuning is for style/format; RAG for facts | RAG-vs-fine-tuning distractor |
| "Reranking improves recall" | It improves ordering/precision, not recall | Metric-confusion distractor |
| "Bigger chunks = better context" | They dilute matches; use parent-child | Chunking distractor |

## Scenario walkthrough

A support assistant over 40,000 KB articles (updated weekly) started giving confident, wrong answers this week. Users ask both natural-language questions and exact error codes. What is the FIRST step and the right architecture?

<Steps>
1. **FIRST step** — pull traces and check whether the gold chunk is retrieved. It surfaced that this week's re-ingest changed the chunker (drift) — a retrieval problem, not a prompt problem.
2. **Chunking** — structural/parent-child for articles: child chunks for precise matches, parent sections for context. Version the config.
3. **Retrieval** — hybrid (dense + BM25): error codes need exact-term matching that dense misses.
4. **Reranking** — cross-encoder to top-8; the gold article was retrieved but ranked #14 pre-rerank.
5. **Eval** — recall@k, MRR, faithfulness, per source; gate re-ingests on these.
6. **Citations** — attach spans so answers are auditable (provenance test).
</Steps>

Rejected alternatives: rewriting the system prompt (the retrieval was broken), fine-tuning on the KB (facts change weekly — RAG's job), dense-only retrieval (misses error codes), and trusting the aggregate score (masked the failing source).

## Key takeaways

- RAG for large/changing corpora + citations; long context for small stable; fine-tuning for style, not facts.
- Chunk to match data shape and query pattern; parent-child balances precise hits with context; version the config.
- Hybrid (dense + BM25) + cross-encoder reranking beats either alone; RRF fuses without score calibration.
- Separate retrieval metrics (recall@k, MRR) from generation metrics (faithfulness, answer relevance), and report per-segment.
- Debug from the trace: gold chunk retrieved? ranked? passed? used? Suspect the index first after a refresh.
