RAG in Production: 15 Lessons From the Retrieval Bench
Your demo answers from three perfect documents. Production answers from three million messy ones, half stale, one poisoned. This course builds the pipeline between those worlds. Chunk it, ground it, price it. Every lesson ships an artifact you can run.
TL;DR: Fifteen workshop lessons for shipping retrieval pipelines, from chunking to eval gates plus tenant isolation. Built from production postmortems with live diagrams plus calculators.
By Kabir · AI engineering course · Updated September 19, 2026
What you will be able to do
- Chunk plus index a messy corpus so real questions find whole answers, not confident fragments.
- Combine vector search with keyword search plus reranking inside a latency budget you measured.
- Ground every answer in cited passages while treating retrieved text as untrusted input.
- Isolate tenants in shared indexes, stream freshness without full rebuilds, plus price every query.
- Answer RAG interviews with a drawn pipeline, priced math, a poisoning story, plus one shipped pilot.
How interviews test this course: draw first, price second, confess one retrieval failure third. Lesson 14 rehearses the bench.
Retrieval Is the Product
Friday night. Your demo answers brilliantly from three hand-picked documents. Monday morning it meets three million real ones. An answer engine once paid for every question twice, retrieval plus inference, because the pipeline is the product now. The model is the last step, never the first.
- Draw the pipeline before the demo. Query plus retrieval plus rerank plus generation plus citations, with the skip path priced first.
- Price one query end to end with the token bill calculator. Retrieval you skip is margin you keep.
- Deep dive: Perplexity Answer Engine, paying twice per question with receipts.
Your artifact this lesson: one pipeline diagram with a price tag on every arrow. No price, no ship. The demo answered. Production is the exam.
Q1. Sketch a RAG pipeline for refund policies with citations. Seen at: AI product plus platform loops.
Q2. Retrieval returns nothing useful. What does the user see? Seen at: applied AI loops.
Pipeline drawn. Now the documents themselves need cutting. Lesson 02: Chunking →
Chunk Like the Question Asks
Tuesday. Your retriever returns half a procedure plus half a warning from the next section. The answer stitches both into confident nonsense. Chunks cut for storage serve the index. Chunks cut for questions serve the user. Only one of them answers correctly.
- Split by meaning, not by length. Headings plus procedures plus tables stay whole or answers fracture.
- Overlap generously at boundaries. The sentence that matters always sits on the cut line.
- Tag every chunk with source plus date. Stale chunks answer with fresh confidence otherwise.
Your artifact: a chunking spec, target size plus overlap plus splitter plus metadata fields. Run ten real questions against two chunk shapes. The shape with whole answers wins, whatever the theory says.
Q1. Chunk a 200-page policy manual for refund questions. Seen at: AI engineering loops.
Q2. Answers cite step 4 without steps 1 to 3. What broke? Seen at: applied AI loops.
Chunks cut. Now the vectors themselves start to rot. Lesson 03: Drift →
Embeddings Drift
Wednesday. The vendor ships a better embedding model overnight. Your index still speaks the old dialect. Queries in the new tongue match nothing, because vectors from two models share a space the way two cities share a street name. Same words, different map.
- Pin the embedding model version like a dependency. Upgrades reindex everything or serve nothing.
- Budget the reindex before adopting the model. Millions of chunks re-embed on your invoice, not theirs.
- Deep dive: Cursor Token Exhaustion, what version moves cost in practice.
Your artifact: an embedding manifest, model name plus version plus dimension plus reindex date. Watch the packet cross the bridge above, then schedule your own crossing. Drift you plan is migration. Drift you discover is an outage.
Q1. The embedding vendor deprecates your model. Walk through the migration. Seen at: AI platform loops.
Q2. Half the index is v1 plus half is v2. What do users notice? Seen at: backend plus AI loops.
Vectors versioned. Now stop trusting vectors alone. Lesson 04: Hybrid search →
Vector Search Is Not Search
Thursday. A user asks about form 1099-B. Vector search returns a soulful essay about tax anxiety. Keyword search would have found the exact form in milliseconds. Meaning matches intent. Tokens match names, codes, plus error strings. Production questions carry both.
- Run both lanes on every query that carries names, codes, or error strings. Vectors paraphrase. Keywords pinpoint.
- Fuse with weights you measured, not defaults you inherited. The mix is a dial, not a doctrine.
- Deep dive: Notion Warehouse, where vector tiering meets keyword truth.
Your artifact: a hybrid test with twenty queries, ten semantic plus ten exact-token. Score each lane separately, then the fusion. Ship the mix that wins both halves, not the lane with the louder vendor.
Q1. Users search by ticket ID plus by symptom description. Design retrieval. Seen at: AI engineering loops.
Q2. When does keyword search beat vectors outright? Seen at: search plus platform loops.
Two lanes merged. Now spend wisely on the final ordering. Lesson 05: Rerank budget →
Rerank With a Budget
Friday. Fusion returns twenty candidates. The cross-encoder reranker reads all twenty slowly plus bills every one. Top-3 accuracy jumps. P99 doubles. Reranking is the most honest trade in retrieval: named precision for metered milliseconds. Spend it where answers matter.
- Rerank depth trades milliseconds for placement. Measure the curve, then pick the knee, not the max.
- Skip reranking for cached plus navigational queries. Precision you already own needs no second opinion.
- Deep dive: Perplexity Answer Engine, where every stage carries a price.
Your artifact: a rerank curve, depth versus accuracy versus p99, drawn from your own golden set. Follow the packet above: twenty in, five out, the meter running between. Chosen depth is architecture. Default depth is a donation.
Q1. Reranking adds 300ms per query. Where do you win it back? Seen at: latency-aware AI loops.
Q2. When would you drop the reranker entirely? Seen at: AI engineering loops.
Ordering bought. Now count what the shelves cost. Lesson 06: Index bill →
The Index Has a Bill
Monday. Three million chunks times fifteen hundred dimensions times four bytes lands near eighteen gigabytes before replicas plus overhead. Nobody priced the serving fleet either. The index is a database with a GPU habit. Model it like one.
- Size vectors from chunk math, not vibes. Chunks times dimensions times bytes times replicas is the floor.
- Quantize deliberately. Smaller vectors cost less plus recall slightly less. Measure the recall you sell.
- Scale the serving fleet with the RPS envelope calculator before launch, not after the invoice.
Your artifact: an index price card, storage plus replicas plus serving fleet, then the monthly total beside it. Follow the waterfall above to the red box. That box is your quarter.
Q1. Size a vector index for ten million chunks with replicas. Seen at: AI infra plus startup loops.
Q2. Cut index cost 40 percent without touching recall. Seen at: FinOps-flavored AI loops.
Shelves priced. Now stop paying for repeats. Lesson 07: Cache repeat →
Cache the Repeat
Tuesday. A voice service bills every hello at 75 milliseconds of GPU while the cached hello answers in 5. Retrieval repeats the same way: the same ten questions arrive all day, each one re-searched plus re-ranked plus re-billed. Same question twice should cost once.
- Cache at three levels: exact query, near-duplicate embedding, plus hot passages. Each level catches a different repeat.
- Expire by document change, not by clock alone. Stale citations spend like real answers.
- Deep dives: The Voice That Bills Per Hello plus Perplexity Answer Engine.
Your artifact: a cache dashboard, hit rate per level plus dollars skipped per day. Watch repeats take the green path above, then measure your own skip rate. Uncached loops answer fast. Cached loops answer free.
Q1. Design a semantic cache for support questions with invalidation. Seen at: AI engineering loops.
Q2. Hit rate falls from 60 to 10 percent overnight. Debug it live. Seen at: SRE plus AI loops.
Repeats skipped. Now make the answers trustworthy. Lesson 08: Grounding →
Ground or Hallucinate
Wednesday. Your pipeline answers a refund question beautifully, citing a policy that never existed. Generation without retrieval is confident fiction. The fix is mechanical, not moral: every claim points at a passage, every passage points at a source, no pointer means no sentence.
- Require passage pointers on factual claims. The gate is a verifier, not a vibe check.
- Show citations to users. Visible sources earn corrections before errors earn churn.
- Deep dive: Perplexity Answer Engine, grounding with receipts attached.
Your artifact: a citation gate with a measured grounded rate on your golden set. Watch bare claims stop at the red box above, then make your gate earn the same halt. Politeness is not accuracy.
Q1. Your RAG cites sources that contradict the answer. Walk through the fix. Seen at: AI engineering loops.
Q2. When should the pipeline refuse to answer at all? Seen at: AI product loops.
Answers grounded. Now assume the documents attack. Lesson 09: Poisoned docs →
Poisoned Retrieval
Thursday. A candidate uploads a CV. The summary comes back with your API keys inside a rendered image. Retrieved documents are untrusted input wearing a trusted uniform. Your pipeline reads the room. The room reads your secrets.
- Treat every retrieved passage as untrusted input. Documents, pages, plus uploads can all carry instructions.
- Scope secrets so theft buys little. Short TTLs plus least privilege turn breaches into footnotes.
- Deep dive: AgentFlayer Zero-Click, the CV that downloaded keys.
Your artifact: three poisoned-document cases, one injection plus one exfil plus one instruction override, run monthly. Watch the packet stop at the scanner above, then make your scanner earn the same blink. Attackers read your docs too.
Q1. A retrieved page tells the model to ignore policy. Walk through the defense. Seen at: AI safety loops.
Q2. Who can write to your index, plus what stops them? Seen at: security-minded AI loops.
Attackers welcomed. Now wall off the tenants. Lesson 10: Tenant walls →
Tenants Share Nothing
Friday. Your retriever serves two companies from one index. A query from tenant A returns a passage stamped tenant B, quoted word for word. Memory never says it is unsure. Indexes never say whose. Ground every recall plus isolate every namespace or bill the apology.
- Filter by tenant before ranking, never after. Post-filtering leaks through scores plus snippets.
- Test isolation with adversarial queries monthly. One tenant recall must never answer another tenant question.
- Deep dive: Cross-Account Channel, the isolation failure with a receipt.
Your artifact: an isolation test suite, ten cross-tenant probes that must all refuse. Follow the packets above: each tenant reaches only its partition, the rest meet refusal. Sharing infrastructure is fine. Sharing answers is a breach.
Q1. Design multi-tenant RAG with provable isolation. Seen at: security-minded AI loops.
Q2. A shared passage is useful to both tenants. Who owns it? Seen at: AI platform loops.
Tenants walled. Now keep the index fresh. Lesson 11: Freshness →
Freshness Versus Rebuild
Monday. Policies change daily. Your weekly full rebuild serves last Tuesday all week. A workspace platform once streamed every block change through a pipeline instead of re-dumping the database nightly. Freshness is a stream, not a schedule.
- Stream document changes into the index in minutes. Rebuilds are backstops, not freshness plans.
- Delete deliberately. Removed documents must vanish from retrieval, not linger as ghosts with citations.
- Deep dive: Notion Warehouse, streaming change capture at block scale.
Your artifact: a freshness SLO, edit-to-searchable latency with a dashboard plus an alert. Follow the green path above, then measure your own lag. Fresh indexes answer today. Rebuilt ones answer last week.
Q1. A policy changes at noon. When do answers change? Prove it. Seen at: AI engineering loops.
Q2. Deleted documents still surface in answers. Diagnose live. Seen at: backend plus AI loops.
Index fresh. Now prove it retrieves. Lesson 12: Eval retriever →
Eval the Retriever
Tuesday. A prompt change ships at noon. By two the refund pipeline cites with new confidence about policies that do not exist. Nobody changed the tests because there were no tests. Retrieval changes are code changes now, so gate them like code or debug them like folklore.
- Gate every chunking plus ranking change on recall at fixed depths. Green suite ships, red suite waits, no exceptions.
- Judge answers separately from retrieval. A good retriever feeding a loose generator still hallucinates.
- Bridge course: AI Evals That Hold, where gates become a full program.
Your artifact: one recall gate in CI with twenty golden queries that must pass before merge. Watch the packet take the green path above, then make your gate earn the same calm. Ungated retrieval is folklore with latency.
Q1. A chunking change drops answer quality silently. How does CI catch it? Seen at: AI engineering loops.
Q2. Recall is high plus answers are wrong. Where is the fault? Seen at: evals-minded AI loops.
Retriever gated. Now count what passages cost the window. Lesson 13: Context budget →
Retrieval Eats Context
Wednesday. Your pipeline stuffs twelve passages into the window, then the model forgets the question. A coding agent once hit a paywall at 11:47 AM because compaction plus re-reading billed the same tokens twice. Retrieved tokens crowd out reasoning the same way. Budget the window like money.
- Cap passages per query plus compress the rest. Five cited passages beat twelve skimmed ones.
- Price the window with the compaction calculator plus the token bill calculator. Long prompts bill twice.
- Deep dive: Cursor Token Exhaustion, the paywall with a timestamp.
Your artifact: a per-query token budget, passages plus reasoning plus headroom, written down before tuning. Watch reasoning squeeze at the red box above, then set your cap where answers stay sharp. Chosen caps are architecture. Surprise caps are incidents.
Q1. More passages help recall but hurt answers. Find the sweet spot live. Seen at: applied AI loops.
Q2. When does a bigger window beat fewer passages? Seen at: LLM platform loops.
Window budgeted. Now rehearse the interview. Lesson 14: Interview bench →
The Interview Bench
Thursday. The panel asks for a support copilot over two million tickets. The candidate who draws the pipeline, prices one query out loud, plus confesses a poisoning incident passes. The candidate who names three vector databases does not. Builders draw. Tourists describe.
- Draw first, talk second. The pipeline with a price tag opens every strong answer.
- Price one million queries out loud with the token bill calculator. Panels score math over adjectives.
- Close with your poisoning story. Attacks you survived beat features you shipped.
Rehearse three answers this week: a chunking design, a hybrid retrieval plan with rerank math, plus a tenant-isolation proof with adversarial tests. Bring diagrams. Builders draw. Tourists describe.
Q1. Design RAG over two million tickets end to end in 45 minutes. Seen at: AI product plus startup loops.
Q2. Tell me about a retrieval failure you caused plus fixed. Seen at: behavioral plus AI loops.
Answers rehearsed. Now assemble everything into one ship. Lesson 15: Ship pipeline →
Ship One Pipeline
Sunday. Pick one corpus from your own work: support tickets, runbooks, or policy docs. Chunk it, index it, fuse the lanes, hang the rerank budget, cache the repeats, gate the citations, scan the poisons, wall the tenants, stream the freshness, eval the recall, budget the window. Then ship it to five users, not five thousand.
- Ship the artifacts from lessons 1 through 13 in one folder. Gaps are visible at a glance.
- Pilot with five users plus a kill switch. Small blast radius, full telemetry, honest notes.
- Close with the invoice. Queries per month times modeled unit price, labeled as estimates.
Course complete. Price the meter next: LLM Inference Economics, then browse all courses plus all case studies. New corpora ship weekly. So will new questions.
The demo answered. Production is the exam. You just built the study guide. Earn the grade.
Sources plus method
This course teaches from published Buildopsy postmortems linked under each lesson: Perplexity inference bills, Cursor token economics, Notion warehouse plus vector tiering, AgentFlayer exfiltration, cross-account isolation, plus the agent eval program. Incident details follow those accounts. Cost framing is modeled from public list prices.
