Back to System Design Index

Free Course15 lessonsLive diagramsKabir

RAG in Production: 15 Lessons From the Retrieval Bench

Your demo answers from three perfect documents. Production answers from three million messy ones, half stale, one poisoned. This course builds the pipeline between those worlds. Chunk it, ground it, price it. Every lesson ships an artifact you can run.

TL;DR: Fifteen workshop lessons for shipping retrieval pipelines, from chunking to eval gates plus tenant isolation. Built from production postmortems with live diagrams plus calculators.

By Kabir · AI engineering course · Updated September 19, 2026

What you will be able to do

How interviews test this course: draw first, price second, confess one retrieval failure third. Lesson 14 rehearses the bench.

Lesson 01 · Foundations · The retrieval pipeline

Retrieval Is the Product

Friday night. Your demo answers brilliantly from three hand-picked documents. Monday morning it meets three million real ones. An answer engine once paid for every question twice, retrieval plus inference, because the pipeline is the product now. The model is the last step, never the first.

query retrieval $0.02 inference $0.31 cache: skip cited answer
  • Draw the pipeline before the demo. Query plus retrieval plus rerank plus generation plus citations, with the skip path priced first.
  • Price one query end to end with the token bill calculator. Retrieval you skip is margin you keep.
  • Deep dive: Perplexity Answer Engine, paying twice per question with receipts.

Your artifact this lesson: one pipeline diagram with a price tag on every arrow. No price, no ship. The demo answered. Production is the exam.

Interview room

Q1. Sketch a RAG pipeline for refund policies with citations. Seen at: AI product plus platform loops.

Q2. Retrieval returns nothing useful. What does the user see? Seen at: applied AI loops.

Pipeline drawn. Now the documents themselves need cutting. Lesson 02: Chunking →

Lesson 02 · Chunking · Cut for questions

Chunk Like the Question Asks

Tuesday. Your retriever returns half a procedure plus half a warning from the next section. The answer stitches both into confident nonsense. Chunks cut for storage serve the index. Chunks cut for questions serve the user. Only one of them answers correctly.

long doc frag A frag B whole answer overlap 15 pct metadata
  • Split by meaning, not by length. Headings plus procedures plus tables stay whole or answers fracture.
  • Overlap generously at boundaries. The sentence that matters always sits on the cut line.
  • Tag every chunk with source plus date. Stale chunks answer with fresh confidence otherwise.

Your artifact: a chunking spec, target size plus overlap plus splitter plus metadata fields. Run ten real questions against two chunk shapes. The shape with whole answers wins, whatever the theory says.

Interview room

Q1. Chunk a 200-page policy manual for refund questions. Seen at: AI engineering loops.

Q2. Answers cite step 4 without steps 1 to 3. What broke? Seen at: applied AI loops.

Chunks cut. Now the vectors themselves start to rot. Lesson 03: Drift →

Lesson 03 · Embeddings · Version plus reindex

Embeddings Drift

Wednesday. The vendor ships a better embedding model overnight. Your index still speaks the old dialect. Queries in the new tongue match nothing, because vectors from two models share a space the way two cities share a street name. Same words, different map.

model v1 space reindex job model v2 space never mix dialects in one index
  • Pin the embedding model version like a dependency. Upgrades reindex everything or serve nothing.
  • Budget the reindex before adopting the model. Millions of chunks re-embed on your invoice, not theirs.
  • Deep dive: Cursor Token Exhaustion, what version moves cost in practice.

Your artifact: an embedding manifest, model name plus version plus dimension plus reindex date. Watch the packet cross the bridge above, then schedule your own crossing. Drift you plan is migration. Drift you discover is an outage.

Interview room

Q1. The embedding vendor deprecates your model. Walk through the migration. Seen at: AI platform loops.

Q2. Half the index is v1 plus half is v2. What do users notice? Seen at: backend plus AI loops.

Vectors versioned. Now stop trusting vectors alone. Lesson 04: Hybrid search →

Lesson 04 · Retrieval · Keywords meet vectors

Vector Search Is Not Search

Thursday. A user asks about form 1099-B. Vector search returns a soulful essay about tax anxiety. Keyword search would have found the exact form in milliseconds. Meaning matches intent. Tokens match names, codes, plus error strings. Production questions carry both.

query vector: meaning keyword: tokens fusion rank top 20
  • Run both lanes on every query that carries names, codes, or error strings. Vectors paraphrase. Keywords pinpoint.
  • Fuse with weights you measured, not defaults you inherited. The mix is a dial, not a doctrine.
  • Deep dive: Notion Warehouse, where vector tiering meets keyword truth.

Your artifact: a hybrid test with twenty queries, ten semantic plus ten exact-token. Score each lane separately, then the fusion. Ship the mix that wins both halves, not the lane with the louder vendor.

Interview room

Q1. Users search by ticket ID plus by symptom description. Design retrieval. Seen at: AI engineering loops.

Q2. When does keyword search beat vectors outright? Seen at: search plus platform loops.

Two lanes merged. Now spend wisely on the final ordering. Lesson 05: Rerank budget →

Lesson 05 · Reranking · Precision with a meter

Rerank With a Budget

Friday. Fusion returns twenty candidates. The cross-encoder reranker reads all twenty slowly plus bills every one. Top-3 accuracy jumps. P99 doubles. Reranking is the most honest trade in retrieval: named precision for metered milliseconds. Spend it where answers matter.

top 20 rerank: slow top 5 generate rerank depth is a latency dial, not a default
  • Rerank depth trades milliseconds for placement. Measure the curve, then pick the knee, not the max.
  • Skip reranking for cached plus navigational queries. Precision you already own needs no second opinion.
  • Deep dive: Perplexity Answer Engine, where every stage carries a price.

Your artifact: a rerank curve, depth versus accuracy versus p99, drawn from your own golden set. Follow the packet above: twenty in, five out, the meter running between. Chosen depth is architecture. Default depth is a donation.

Interview room

Q1. Reranking adds 300ms per query. Where do you win it back? Seen at: latency-aware AI loops.

Q2. When would you drop the reranker entirely? Seen at: AI engineering loops.

Ordering bought. Now count what the shelves cost. Lesson 06: Index bill →

Lesson 06 · Economics · Storage plus serving

The Index Has a Bill

Monday. Three million chunks times fifteen hundred dimensions times four bytes lands near eighteen gigabytes before replicas plus overhead. Nobody priced the serving fleet either. The index is a database with a GPU habit. Model it like one.

3M chunks 18 GB vec replicas x3 QPS fleet bytes times replicas times queries, every month
  • Size vectors from chunk math, not vibes. Chunks times dimensions times bytes times replicas is the floor.
  • Quantize deliberately. Smaller vectors cost less plus recall slightly less. Measure the recall you sell.
  • Scale the serving fleet with the RPS envelope calculator before launch, not after the invoice.

Your artifact: an index price card, storage plus replicas plus serving fleet, then the monthly total beside it. Follow the waterfall above to the red box. That box is your quarter.

Interview room

Q1. Size a vector index for ten million chunks with replicas. Seen at: AI infra plus startup loops.

Q2. Cut index cost 40 percent without touching recall. Seen at: FinOps-flavored AI loops.

Shelves priced. Now stop paying for repeats. Lesson 07: Cache repeat →

Lesson 07 · Caching · Same question twice

Cache the Repeat

Tuesday. A voice service bills every hello at 75 milliseconds of GPU while the cached hello answers in 5. Retrieval repeats the same way: the same ten questions arrive all day, each one re-searched plus re-ranked plus re-billed. Same question twice should cost once.

query repeat: 5ms novel: full pipe fill on miss TTL guard
  • Cache at three levels: exact query, near-duplicate embedding, plus hot passages. Each level catches a different repeat.
  • Expire by document change, not by clock alone. Stale citations spend like real answers.
  • Deep dives: The Voice That Bills Per Hello plus Perplexity Answer Engine.

Your artifact: a cache dashboard, hit rate per level plus dollars skipped per day. Watch repeats take the green path above, then measure your own skip rate. Uncached loops answer fast. Cached loops answer free.

Interview room

Q1. Design a semantic cache for support questions with invalidation. Seen at: AI engineering loops.

Q2. Hit rate falls from 60 to 10 percent overnight. Debug it live. Seen at: SRE plus AI loops.

Repeats skipped. Now make the answers trustworthy. Lesson 08: Grounding →

Lesson 08 · Grounding · Citations or fiction

Ground or Hallucinate

Wednesday. Your pipeline answers a refund question beautifully, citing a policy that never existed. Generation without retrieval is confident fiction. The fix is mechanical, not moral: every claim points at a passage, every passage points at a source, no pointer means no sentence.

draft claims cite gate cited: ship bare: block answer
  • Require passage pointers on factual claims. The gate is a verifier, not a vibe check.
  • Show citations to users. Visible sources earn corrections before errors earn churn.
  • Deep dive: Perplexity Answer Engine, grounding with receipts attached.

Your artifact: a citation gate with a measured grounded rate on your golden set. Watch bare claims stop at the red box above, then make your gate earn the same halt. Politeness is not accuracy.

Interview room

Q1. Your RAG cites sources that contradict the answer. Walk through the fix. Seen at: AI engineering loops.

Q2. When should the pipeline refuse to answer at all? Seen at: AI product loops.

Answers grounded. Now assume the documents attack. Lesson 09: Poisoned docs →

Lesson 09 · Adversarial · Retrieval as attack path

Poisoned Retrieval

Thursday. A candidate uploads a CV. The summary comes back with your API keys inside a rendered image. Retrieved documents are untrusted input wearing a trusted uniform. Your pipeline reads the room. The room reads your secrets.

poisoned doc hidden orders scan gate model safe
  • Treat every retrieved passage as untrusted input. Documents, pages, plus uploads can all carry instructions.
  • Scope secrets so theft buys little. Short TTLs plus least privilege turn breaches into footnotes.
  • Deep dive: AgentFlayer Zero-Click, the CV that downloaded keys.

Your artifact: three poisoned-document cases, one injection plus one exfil plus one instruction override, run monthly. Watch the packet stop at the scanner above, then make your scanner earn the same blink. Attackers read your docs too.

Interview room

Q1. A retrieved page tells the model to ignore policy. Walk through the defense. Seen at: AI safety loops.

Q2. Who can write to your index, plus what stops them? Seen at: security-minded AI loops.

Attackers welcomed. Now wall off the tenants. Lesson 10: Tenant walls →

Lesson 10 · Isolation · Shared index, separate truths

Tenants Share Nothing

Friday. Your retriever serves two companies from one index. A query from tenant A returns a passage stamped tenant B, quoted word for word. Memory never says it is unsure. Indexes never say whose. Ground every recall plus isolate every namespace or bill the apology.

tenant A tenant B filter gate part A part B none: refuse
  • Filter by tenant before ranking, never after. Post-filtering leaks through scores plus snippets.
  • Test isolation with adversarial queries monthly. One tenant recall must never answer another tenant question.
  • Deep dive: Cross-Account Channel, the isolation failure with a receipt.

Your artifact: an isolation test suite, ten cross-tenant probes that must all refuse. Follow the packets above: each tenant reaches only its partition, the rest meet refusal. Sharing infrastructure is fine. Sharing answers is a breach.

Interview room

Q1. Design multi-tenant RAG with provable isolation. Seen at: security-minded AI loops.

Q2. A shared passage is useful to both tenants. Who owns it? Seen at: AI platform loops.

Tenants walled. Now keep the index fresh. Lesson 11: Freshness →

Lesson 11 · Freshness · Streams beat rebuilds

Freshness Versus Rebuild

Monday. Policies change daily. Your weekly full rebuild serves last Tuesday all week. A workspace platform once streamed every block change through a pipeline instead of re-dumping the database nightly. Freshness is a stream, not a schedule.

doc edits stream: minutes rebuild: weekly fresh index stale: week
  • Stream document changes into the index in minutes. Rebuilds are backstops, not freshness plans.
  • Delete deliberately. Removed documents must vanish from retrieval, not linger as ghosts with citations.
  • Deep dive: Notion Warehouse, streaming change capture at block scale.

Your artifact: a freshness SLO, edit-to-searchable latency with a dashboard plus an alert. Follow the green path above, then measure your own lag. Fresh indexes answer today. Rebuilt ones answer last week.

Interview room

Q1. A policy changes at noon. When do answers change? Prove it. Seen at: AI engineering loops.

Q2. Deleted documents still surface in answers. Diagnose live. Seen at: backend plus AI loops.

Index fresh. Now prove it retrieves. Lesson 12: Eval retriever →

Lesson 12 · Evaluation · Golden sets for retrieval

Eval the Retriever

Tuesday. A prompt change ships at noon. By two the refund pipeline cites with new confidence about policies that do not exist. Nobody changed the tests because there were no tests. Retrieval changes are code changes now, so gate them like code or debug them like folklore.

index edit recall gate fail: block pass: ship prod calm
  • Gate every chunking plus ranking change on recall at fixed depths. Green suite ships, red suite waits, no exceptions.
  • Judge answers separately from retrieval. A good retriever feeding a loose generator still hallucinates.
  • Bridge course: AI Evals That Hold, where gates become a full program.

Your artifact: one recall gate in CI with twenty golden queries that must pass before merge. Watch the packet take the green path above, then make your gate earn the same calm. Ungated retrieval is folklore with latency.

Interview room

Q1. A chunking change drops answer quality silently. How does CI catch it? Seen at: AI engineering loops.

Q2. Recall is high plus answers are wrong. Where is the fault? Seen at: evals-minded AI loops.

Retriever gated. Now count what passages cost the window. Lesson 13: Context budget →

Lesson 13 · Context · Passages crowd the window

Retrieval Eats Context

Wednesday. Your pipeline stuffs twelve passages into the window, then the model forgets the question. A coding agent once hit a paywall at 11:47 AM because compaction plus re-reading billed the same tokens twice. Retrieved tokens crowd out reasoning the same way. Budget the window like money.

question 5 passages reasoning overflow five passages in, reasoning squeezed out

Your artifact: a per-query token budget, passages plus reasoning plus headroom, written down before tuning. Watch reasoning squeeze at the red box above, then set your cap where answers stay sharp. Chosen caps are architecture. Surprise caps are incidents.

Interview room

Q1. More passages help recall but hurt answers. Find the sweet spot live. Seen at: applied AI loops.

Q2. When does a bigger window beat fewer passages? Seen at: LLM platform loops.

Window budgeted. Now rehearse the interview. Lesson 14: Interview bench →

Lesson 14 · Interviews · Draw plus price plus confess

The Interview Bench

Thursday. The panel asks for a support copilot over two million tickets. The candidate who draws the pipeline, prices one query out loud, plus confesses a poisoning incident passes. The candidate who names three vector databases does not. Builders draw. Tourists describe.

draw pipe price query failure told offer
  • Draw first, talk second. The pipeline with a price tag opens every strong answer.
  • Price one million queries out loud with the token bill calculator. Panels score math over adjectives.
  • Close with your poisoning story. Attacks you survived beat features you shipped.

Rehearse three answers this week: a chunking design, a hybrid retrieval plan with rerank math, plus a tenant-isolation proof with adversarial tests. Bring diagrams. Builders draw. Tourists describe.

Interview room

Q1. Design RAG over two million tickets end to end in 45 minutes. Seen at: AI product plus startup loops.

Q2. Tell me about a retrieval failure you caused plus fixed. Seen at: behavioral plus AI loops.

Answers rehearsed. Now assemble everything into one ship. Lesson 15: Ship pipeline →

Lesson 15 · Capstone · Chunk plus ground plus price

Ship One Pipeline

Sunday. Pick one corpus from your own work: support tickets, runbooks, or policy docs. Chunk it, index it, fuse the lanes, hang the rerank budget, cache the repeats, gate the citations, scan the poisons, wall the tenants, stream the freshness, eval the recall, budget the window. Then ship it to five users, not five thousand.

corpus cut gates hung cache warm 5 users, priced five users first, five thousand later
  • Ship the artifacts from lessons 1 through 13 in one folder. Gaps are visible at a glance.
  • Pilot with five users plus a kill switch. Small blast radius, full telemetry, honest notes.
  • Close with the invoice. Queries per month times modeled unit price, labeled as estimates.

Course complete. Price the meter next: LLM Inference Economics, then browse all courses plus all case studies. New corpora ship weekly. So will new questions.

The demo answered. Production is the exam. You just built the study guide. Earn the grade.

Sources plus method

This course teaches from published Buildopsy postmortems linked under each lesson: Perplexity inference bills, Cursor token economics, Notion warehouse plus vector tiering, AgentFlayer exfiltration, cross-account isolation, plus the agent eval program. Incident details follow those accounts. Cost framing is modeled from public list prices.