# RAG Cookbook (OpenAI)

Retrieval-augmented generation recipes for the OpenAI stack — pipeline anatomy, chunking trade-offs, embeddings and vector stores, the file search tool versus a self-managed store, hybrid retrieval and reranking, grounding, permissions, freshness, evaluating retrieval separately from generation, a failure taxonomy, and a cost and latency budget.

import { Steps } from '@prosefly/astro-components';

RAG grounds answers in retrieved passages at query time. Reach for it when the corpus is **large or changing** and you need **citations**. Prefer long context for a small stable corpus that fits comfortably in the window; prefer fine-tuning for **style and format**, not for facts. This cookbook mirrors the objectives of the Academy *Build with Retrieval-Augmented Generation* course; it is independent preparation built from published learning objectives.

:::tip[Assessment signal]
"Confident but wrong after a document refresh" points at **retrieval or indexing first** — a stale index, a changed chunker, re-embedded with a different model — not the prompt or the model. Rewriting the prompt is the distractor.
:::

## Pipeline anatomy

```text
 ingest ──► chunk ──► embed ──► store        (build time, offline)
                                   │
 query ──► rewrite ──► retrieve ──► rerank ──► assemble ──► generate ──► cite
                                                              │
                                                        (query time, online)
```

Two halves fail for different reasons. The **build-time** half (chunk, embed, store) fails silently: a bad chunker or a changed embedding model degrades every future query. The **query-time** half (retrieve, rerank, generate) fails visibly, per request. Instrument both, and evaluate them separately.

| Stage | Owns | Common failure |
| --- | --- | --- |
| Chunk | Splitting docs into retrievable units | Splits mid-idea; drift after re-ingest |
| Embed | Turning text into vectors | Wrong or mismatched model between index and query |
| Store | Indexing and nearest-neighbour search | Stale vectors; missing permission metadata |
| Retrieve | Fetching candidates | Misses exact IDs (dense only); wrong top-k |
| Rerank | Ordering candidates by relevance | Absent, so the gold passage sits too low |
| Generate | Writing the grounded answer | Ignores context; no citation |

## When RAG vs alternatives

| Situation | Choose | Why |
| --- | --- | --- |
| Thousands of docs, updated often, must cite | RAG | Fresh, grounded, auditable |
| A 50-page handbook that rarely changes and fits the window | Long context | No retrieval infrastructure; simplest |
| "Always answer in our house style / format" | Fine-tuning or few-shot | Behaviour, not facts |
| Model must decide when and what to retrieve | Agentic RAG (retrieval as a tool) | Iterative reformulation |

The GPT-5.6 and GPT-6 models carry a large context window, which tempts teams to "just paste everything". That works for a small stable corpus, but for a large or changing one it is more expensive per query, harder to keep fresh, and gives you no citations to audit.

## Chunking strategies with trade-offs

| Strategy | Start here | Best for | Trade-off |
| --- | --- | --- | --- |
| Fixed-size | 512 tokens, 50-token overlap | Uniform prose | Splits mid-idea |
| Recursive (by separator) | 400–800 tokens, split on `\n\n` then `\n` then sentence | Mixed prose | Uneven sizes |
| Semantic | Break where embedding similarity dips | Topic-shifting docs | Compute cost at ingest |
| Structural / document-aware | One chunk per section or heading | Manuals, contracts, wikis | Needs clean structure |
| Parent-child (small-to-big) | Child 150–300 tok to match, return parent 800–1200 tok | Precise hit plus broad context | Two-tier store |

**Overlap** rule of thumb: 10–20% of the chunk size. Too little loses boundary context; too much inflates the index and produces duplicate hits.

:::caution[Chunk drift]
When documents are re-ingested with a different chunker or size, retrieval quality shifts silently and answers go confidently wrong. Version your chunking configuration and re-run retrieval evals after any change to it.
:::

## Embeddings and vector stores

Embeddings turn text into vectors so that semantically similar passages sit near each other. Two rules dominate:

<Steps>
1. **The index and the query must use the same embedding model and dimension.** Mixing models is the classic silent corruption: cosine similarity between vectors from two different models is meaningless.
2. **Re-embed the whole corpus when you change the model.** A half-migrated index returns garbage for the un-migrated half. Treat an embedding-model change like a schema migration: all or nothing, gated on an eval.
</Steps>

The store holds vectors plus metadata (source, section, ACL, timestamp) and does approximate nearest-neighbour search. On the OpenAI stack you can either let the platform manage this for you or run your own.

## The file search tool vs a self-managed store

| | File search tool (managed) | Self-managed vector store |
| --- | --- | --- |
| Who chunks and embeds | The platform | You |
| Who indexes and retrieves | The platform, called as a tool | Your database and code |
| Control over chunking / reranking | Limited to the tool's options | Full |
| Best when | You want retrieval fast, with less infrastructure | You need custom chunking, hybrid search, your own reranker, or your own store of record |
| Trade-off | Less tuning surface | You own freshness, permissions and evals end to end |

The decision rule: reach for the **file search tool** when you want grounded answers over your files without standing up a retrieval stack, and reach for a **self-managed store** when you need control the tool does not expose — a specific chunking scheme, hybrid retrieval, a cross-encoder reranker, or permission-aware retrieval integrated with your identity system.

## Hybrid retrieval and reranking

Dense (embedding) retrieval catches paraphrase; sparse (BM25) retrieval catches exact IDs, error codes and rare terms. Combine them, then rerank.

```yaml
retrieval:
  dense:   { top_k: 40 }
  sparse:  { algorithm: bm25, top_k: 40 }
  fusion:  { method: reciprocal_rank_fusion, k: 60 }
  rerank:  { model: cross-encoder, top_n: 8 }
```

**Reciprocal Rank Fusion**: a document's score is the sum over lists of `1 / (k + rank)`. With `k=60`, a document ranked #1 in both lists scores about `1/61 + 1/61 = 0.033`; ranked #1 in one list and absent in the other scores about `1/61 = 0.016`. RRF needs no score calibration between the two systems — it fuses by rank position alone.

A **cross-encoder reranker** reads the query and each candidate together, giving far better precision than the first-stage retriever, at higher per-query cost. Retrieve broadly (top-k 40–100), then rerank down to a small top-n (5–10) for the model.

| | First-stage retrieval | Cross-encoder rerank |
| --- | --- | --- |
| Encodes | Query and docs separately, precomputed | Query and doc jointly, per pair |
| Speed | Fast, index-time | Slow, query-time |
| Strength | Recall | Precision / ordering |
| Role | Fetch top-k | Order to top-n |

## Grounding and citation formatting

Grounding is only real if the answer can be traced to a passage. Two disciplines:

- **Constrain the answer to the retrieved context**, and provide a sentinel for "not found": `Answer only from the passages below. If they do not contain the answer, reply NOT COVERED.` The sentinel lets code detect a non-answer instead of shipping a guess.
- **Return citations the reader can verify**: a source identifier and a section or span for each claim. When the model must synthesise across passages, ask it to cite each supporting passage rather than one blanket citation at the end.

## Permission-aware retrieval

The retriever must never surface a passage the asking user is not allowed to see. This is an ingest-and-retrieve concern, not a prompt concern.

<Steps>
1. **Attach access metadata at ingest** — store each chunk's ACL (owner, group, sensitivity) alongside its vector.
2. **Filter at query time by the asking user's identity** — apply the permission filter as part of retrieval, before the model ever sees a candidate.
3. **Never rely on a prompt instruction** like "do not reveal restricted content" — an unfiltered restricted passage in context is already a leak, whatever the prompt says.
4. **Re-check on permission changes** — when a user loses access, their future queries must stop returning those chunks; permission state lives with the data, not the session.
</Steps>

:::caution[Permissions are retrieval, not prompting]
The trap is a stem where sensitive data leaks and the tempting fix is a stronger system-prompt rule. If a restricted passage was retrieved into context, the prompt is irrelevant — the fix is a permission filter at retrieval time.
:::

## Freshness

A RAG system is only as current as its index. Decide, per corpus, how fresh it must be and build for that:

| Freshness need | Approach |
| --- | --- |
| Minutes (prices, tickets) | Retrieve from the source of record at query time, or stream updates into the index |
| Hours to a day | Scheduled re-ingest; version the index and swap atomically |
| Rarely changes | Periodic full rebuild; gate the swap on a retrieval eval |

Whatever the cadence, **gate re-ingests on an eval** so a chunker or model change cannot silently degrade retrieval, and store a build timestamp so "confident but wrong after a refresh" is diagnosable.

## Evaluating retrieval separately from generation

A good answer can hide bad retrieval and vice versa, so measure them apart.

| Metric | Measures | Example |
| --- | --- | --- |
| Recall@k | Was the gold passage in the top-k? | 8 of 10 queries had it in top-5 → 0.80 |
| MRR | Rank of the first relevant hit | Ranks 1, 3, 2 → (1 + 1/3 + 1/2)/3 = 0.61 |
| Precision@k | Fraction of top-k that are relevant | 2 of top-5 relevant → 0.40 |
| Faithfulness | Is the answer supported by the retrieved context? | 47 of 50 claims grounded → 0.94 |
| Answer relevance | Does the answer address the question? | Rubric or model-graded |

Worked comparison on a 200-query set:

| Config | Recall@5 | MRR | Faithfulness | Answer relevance |
| --- | --- | --- | --- | --- |
| Dense only | 0.71 | 0.52 | 0.86 | 0.83 |
| Hybrid (RRF) | 0.83 | 0.61 | 0.90 | 0.86 |
| Hybrid + reranker | 0.83 | 0.74 | 0.94 | 0.90 |

Reranking barely moves Recall@5 (same candidates) but lifts MRR and faithfulness sharply by putting the right passage first, where the model reads it earliest. Report every metric **per segment** (document type, source, language) — an aggregate 0.90 can hide one source at 0.40.

## Failure taxonomy with fixes

| Symptom | Likely cause | Fix |
| --- | --- | --- |
| Answer wrong, gold passage not retrieved | Chunk drift, wrong embedding model, poor query | Version chunking; align embed models; add query rewriting |
| Gold passage retrieved but ranked low | No reranker; bad fusion weights | Add a cross-encoder reranker; tune RRF `k` |
| Exact codes / IDs missed | Dense-only retrieval | Add BM25 (hybrid) |
| Retrieved and top-ranked but answer wrong | Model ignoring context; context not passed | Check the prompt actually includes chunks; constrain to context |
| Confident but wrong after a refresh | Stale or re-chunked index | Re-run retrieval evals; check build timestamp |
| One source consistently bad | Aggregate hid it | Report per-segment; fix that source's ingest |
| Restricted content surfaced | No permission filter at retrieval | Filter by identity at query time |

## Cost and latency budget worked example

A support assistant over 40,000 articles, target end-to-end latency under 2 seconds at p95, `gpt-5.6-terra` for generation.

```text
Per query budget (p95 target: 1900 ms)
  query rewrite (gpt-5.6-luna, low effort)   ~120 ms
  hybrid retrieve (dense + BM25, top_k 40)   ~140 ms
  cross-encoder rerank (40 -> 8)             ~180 ms
  generate grounded answer (terra, 8 chunks) ~1200 ms
  citation assembly (code)                    ~20 ms
                                             --------
  total (p95)                                ~1660 ms  ✓ under 1900
```

Cost levers, cheapest first: cut the number of chunks fed to the model (8 → 5 with a stronger reranker) before touching the model tier; use `gpt-5.6-luna` for the rewrite step; cache the stable system prefix on the API; and reserve a larger model for only the queries a router flags as hard. The wrong instinct — "use a bigger model to fix wrong answers" — spends money on a generation problem that is usually a retrieval problem.

## Common misconceptions

| Misconception | Reality | Why it matters on the assessment |
| --- | --- | --- |
| "More context always helps" | Noise lowers faithfulness; rerank to a small top-n | Over-retrieval distractor |
| "Dense retrieval covers everything" | It misses exact IDs and codes; add BM25 | Hybrid-search signal |
| "Wrong answer means fix the prompt" | Check retrieval first after a refresh | Root-cause distractor |
| "A prompt rule keeps restricted data out" | Permissions are enforced at retrieval, not by prompt | Permission trap |
| "Fine-tune to add new facts" | Fine-tuning is for style; RAG for facts | RAG-vs-fine-tuning distractor |
| "Reranking improves recall" | It improves ordering and precision, not recall | Metric-confusion distractor |

## Scenario walkthrough

A team runs a KB assistant over 40,000 articles refreshed weekly. This week it began giving confident wrong answers. Users ask both natural-language questions and exact error codes. What is the FIRST step and the right architecture?

<Steps>
1. **FIRST step** — pull traces and check whether the gold passage is even retrieved. It surfaced that this week's re-ingest changed the chunker (drift). This is a retrieval problem, so rewriting the prompt would waste the cycle.
2. **Chunking** — structural or parent-child for articles: child chunks for precise matches, parent sections for context; version the config and gate re-ingests on an eval.
3. **Retrieval** — hybrid dense + BM25, because error codes need the exact-term matching that dense retrieval misses.
4. **Reranking** — cross-encoder to top-8; the gold article was retrieved but sat at rank 14 before reranking.
5. **Permissions** — filter by the asking user's identity at query time, since some articles are internal-only.
6. **Eval** — recall@k, MRR and faithfulness, reported per source, so a single failing source cannot hide behind the aggregate.
</Steps>

Rejected alternatives: rewriting the system prompt (retrieval was broken), fine-tuning on the KB (facts change weekly — that is RAG's job), dense-only retrieval (misses error codes), and trusting the aggregate score (it masked the failing source).

## Key takeaways

- RAG for large or changing corpora with citations; long context for small stable corpora; fine-tuning for style, not facts.
- Chunk to match the data shape; parent-child balances precise hits with context; version the config and gate re-ingests.
- Keep the index and query on the same embedding model; re-embed the whole corpus on any model change.
- Choose the file search tool for speed with less infrastructure, a self-managed store when you need chunking, hybrid retrieval, a reranker or permission integration you control.
- Hybrid (dense + BM25) plus a cross-encoder reranker beats either alone; RRF fuses without score calibration.
- Enforce permissions and freshness at retrieval time, and evaluate retrieval (recall@k, MRR) separately from generation (faithfulness), per segment.
