LLM Inference Economics: 15 Lessons From the Token Meter
Your demo answers beautifully at thirty cents a turn. One million turns later the invoice arrives with its own weather system. This course teaches the meter behind the magic. Meter it, batch it, cache it. Every lesson ships an artifact you can run.
TL;DR: Fifteen workshop lessons for pricing plus taming inference spend, from KV cache to vendor pacts plus the RPS envelope. Built from production postmortems with live diagrams plus calculators.
By Kabir · AI engineering course · Updated September 19, 2026
What you will be able to do
- Price one turn end to end, then scale the math to one million turns without flinching.
- Win back margin with KV cache reuse, batching, request caching, plus right-sized model routing.
- Survive compaction cliffs, vendor price moves, retry multipliers, plus idle billing with hedges you chose.
- Meter every turn in production with per-endpoint dashboards plus alerts that fire before the invoice.
- Answer inference interviews with a drawn meter, priced math, a bill shock story, plus one shipped endpoint.
How interviews test this course: draw first, price second, confess one bill shock third. Lesson 14 rehearses the bench.
The Turn Has a Price
Friday night. Your agent answers beautifully at thirty cents a turn. One million turns later the invoice arrives with its own weather system. An answer engine once paid for every question twice, retrieval plus inference, until caching made repeats nearly free. Price one turn before scaling to millions.
- Price one turn end to end with the token bill calculator. Cache hits turn per-turn math into margin.
- Scale to one million turns on paper before launch. Envelopes surprise only the unpriced.
- Deep dive: Perplexity Answer Engine, paying twice per question with receipts.
Your artifact this lesson: a one-turn price card, input plus cache rate plus output plus total, then the million-turn envelope beside it. No price, no ship. The demo charmed. Production invoices.
Q1. Price one million support turns monthly, line by line, out loud. Seen at: startup plus CTO-round loops.
Q2. What breaks first when turns triple overnight? Seen at: applied AI loops.
Turn priced. Now split the meter in two. Lesson 02: Prefill decode →
Prefill Is Not Decode
Monday. Your long prompt bills heavily before generating a single word. That is prefill: the model reads everything at once, compute-bound plus hungry. Decode then drips tokens one by one, memory-bound plus slow. One request, two physics, two prices. Tune them separately.
- Shorten prompts to cut prefill. Every repeated instruction block is a tax you chose.
- Cap output length deliberately. Rambling decodes bill per token plus bore per reader.
- Deep dive: Perplexity Answer Engine, where prefill flaws taxed every query.
Your artifact: a phase split for your heaviest endpoint, prefill share versus decode share, measured not guessed. Watch the packet slow at each stage above, then attack the bigger half first. Averages lie. Phases confess.
Q1. Long prompts make your endpoint slow. Is it prefill or decode? Prove it. Seen at: AI infra loops.
Q2. When does a shorter prompt beat a bigger GPU? Seen at: performance-aware AI loops.
Phases split. Now reuse what prefill computed. Lesson 03: KV margin →
KV Cache Is Margin
Tuesday. Ten thousand users share one system prompt plus the fleet recomputes it ten thousand times. The key-value cache stores that shared prefix once, then reuses it for every turn. Same answers. A fraction of the compute. Cache hits turn per-turn math into margin.
- Front-load shared instructions so prefixes match. Cache hits need identical beginnings, not similar intentions.
- Measure hit rate per endpoint. Unmeasured caches are rumors with dashboards.
- Deep dive: Perplexity Answer Engine, KV reuse done right plus prefill done wrong.
Your artifact: a prefix audit, shared tokens per endpoint plus measured hit rate plus dollars skipped. Follow the green path above, then standardize your own prefixes. Matching beginnings print money. Creative ones burn it.
Q1. Design prompts for maximum KV cache reuse across users. Seen at: AI infra loops.
Q2. Cache hit rate is 5 percent. What do you change first? Seen at: applied AI loops.
Prefix cached. Now pack the GPU tighter. Lesson 04: Batching →
Batch or Bleed
Wednesday. Your GPUs serve one request at a time while thirty wait in line. Utilization naps at 20 percent. Continuous batching packs waiting requests into every forward pass, so the same silicon serves multiples. Idle GPUs bill like busy ones. Pack them or pay for standing still.
- Batch continuously, not statically. Requests join plus leave mid-generation instead of waiting for stragglers.
- Guard tail latency while packing. Throughput that tortures p99 is a transfer, not a win.
- Deep dive: The Voice That Bills Per Hello, throughput math at a billion-scale bill.
Your artifact: a batching test, throughput versus p99 at four batch sizes, drawn as one curve. Follow the green box above, then pick the size at the knee. Packed GPUs print turns. Lonely ones print invoices.
Q1. GPU utilization sits at 20 percent with a queue forming. Fix it live. Seen at: AI infra loops.
Q2. Batching helps throughput but hurts latency. Where is the line? Seen at: performance-aware AI loops.
GPU packed. Now kill the most repeated work. Lesson 05: Hello tax →
Hello Costs 75ms
Thursday. Voice AI bills every hello at 75 milliseconds of GPU. The cached hello answers in 5. Same word. Fifteen times cheaper. Somewhere a quarter of all greetings repeat, which means a quarter of the invoice is optional. Your pipeline has hellos too. Find them.
- Log plus rank repeated inputs weekly. The top ten repeats fund the whole cache layer.
- Cache audio, embeddings, plus full answers separately. Each repeat shape needs its own shelf.
- Deep dive: The Voice That Bills Per Hello, the hello with a receipt.
Your artifact: a repeat report, top repeated inputs plus their billed cost plus cache design. Follow the green path above, then kill your own hellos. Repeated work is a subscription you never signed.
Q1. A quarter of traffic repeats daily greetings. Design the cache. Seen at: AI product loops.
Q2. Cached answers go stale on voice. How do you version them? Seen at: applied AI loops.
Hellos cached. Now survive the cliff at 11:47 AM. Lesson 06: Compaction →
Compaction Bills Twice
Friday. Your coding agent flies all morning, then hits a paywall at 11:47 AM with the sprint half done. Compaction crushed hundreds of thousands of tokens into thousands, then billed dozens of re-reading steps to rebuild what was lost. The meter runs twice. The auditor naps through both.
- Compact on purpose with named facts, not on accident at the cliff edge. Summaries you design survive.
- Price the rebuild with the compaction calculator plus the token bill calculator. Two meters, one incident.
- Deep dive: Cursor Token Exhaustion, the paywall with a timestamp.
Your artifact: a context budget per session with a compaction trigger you chose. Watch the packet cross the cliff above, then decide where your cliff sits. Chosen cliffs are architecture. Surprise cliffs are incidents.
Q1. Your agent degrades after 50 turns plus bills double. Diagnose live. Seen at: applied AI loops.
Q2. When does a bigger window beat a better summary? Seen at: LLM platform loops.
Cliff priced. Now watch vendors move the prices. Lesson 07: Price pact →
The Price Pact Broke
Monday. Three labs agree to slow down plus publish safety pacts. By Friday one ships a cheaper flagship that undercuts the pact by half. Price lists move faster than promises. Your margin plan must survive the vendor you chose plus the three you did not.
- Abstract the model behind your router, never behind hope. Swaps should take a config change, not a rewrite.
- Track price per million monthly like a dependency version. Moves you watch cost little. Moves you miss cost quarters.
- Deep dive: Muse Price Pact Fracture, the week pacts met price lists.
Your artifact: a vendor sheet, price per million per model plus switch cost plus trigger thresholds. Follow the packet down to the green tier above, then set your own shift rules. Loyalty is a strategy. Lock-in is a bill.
Q1. Your vendor doubles prices with 30 days notice. What moves first? Seen at: startup plus CTO-round loops.
Q2. When is single-vendor depth worth the lock-in risk? Seen at: platform plus AI loops.
Pact read. Now route every query to its cheapest sufficient mind. Lesson 08: Routing →
Route Small, Spend Small
Tuesday. Eighty percent of your queries need a sledgehammer, says nobody who measured. Most turns classify, summarize, or extract, work a small model does at a tenth of the price. Route by difficulty. Pay flagship rates only for flagship problems.
- Classify difficulty before generating. A cheap judge routing to a cheaper worker beats hope at scale.
- Escalate with evidence. Misrouted hard queries cost more in retries than the flagship would have.
- Deep dive: Perplexity Answer Engine, tiered work with tiered bills.
Your artifact: a routing table, query shape plus model plus observed quality plus unit price. Watch easy traffic take the green path above, then route your own easy majority down. Flagship brains for flagship problems. Mini brains for everything else.
Q1. Cut inference spend 50 percent without touching quality. Show the routing. Seen at: FinOps-flavored AI loops.
Q2. The classifier misroutes hard queries down. How do you catch it? Seen at: evals-minded AI loops.
Traffic routed. Now scale the math to the month. Lesson 09: RPS envelope →
The RPS Envelope
Wednesday. Thirty cents a turn feels trivial until one million turns a month turn it into a headcount. The envelope math is brutal plus simple: requests per second times seconds per month times unit price. Every pricing conversation that skips this step ends in a surprise.
- Run your shape through the RPS envelope calculator plus the token bill calculator. Two calculators, one honest number.
- Present envelopes as ranges with stated assumptions. Precision without assumptions is theater.
- Deep dive: The Voice That Bills Per Hello, envelope math at voice scale.
Your artifact: a monthly envelope for your endpoint, base plus growth plus peak, all labeled as estimates. Follow the waterfall above to the green total. That total is your budget conversation.
Q1. Price 100 RPS of agent traffic for a year, out loud. Seen at: startup plus CTO-round loops.
Q2. Growth triples RPS but the budget doubles. What gives? Seen at: platform plus AI loops.
Envelope sealed. Now add the quality line items. Lesson 10: Guardrail cost →
Guardrails Cost Too
Thursday. Evals, judges, plus red-team runs all consume the product they protect. A golden set of five hundred cases re-run on every merge is a second inference bill wearing a quality costume. Budget it openly or quality becomes the first cut when money tightens.
- Budget eval spend per merge plus per release. Quality with a meter survives budget season.
- Verify mechanically where possible. Proofs plus schemas cost less than judges per check.
- Bridge course: AI Evals That Hold, where the meter becomes a program.
Your artifact: an eval budget, cases times runs times unit price per month. Watch spend accumulate at the red box above, then defend it as insurance with a number. Unpriced quality gets cut first.
Q1. The CFO asks to cut eval spend in half. What do you protect? Seen at: startup plus leadership loops.
Q2. When is a code check enough plus when do you need a judge? Seen at: evals-minded loops.
Quality budgeted. Now stop retries from re-billing. Lesson 11: Retry meter →
Retries Multiply the Meter
Friday. Your agents retried failed deploys until retries cost more than features. Every retry is a second billed turn arriving exactly when the system is weakest. At two extra retries per stalled call with no backoff, recovery load lands near six times normal. Failed turns bill like successful ones.
- Budget retries like capacity. Run the modeled shape through the retry storm calculator before the incident.
- Cap attempts plus open breakers while errors rage. Unbounded retries are a self-inflicted invoice.
- Deep dive: Replit Agent Runtime, where retries cost more than features.
Your artifact: a retry policy with numbers, max attempts plus base delay plus jitter plus breaker threshold, each with a billed cost. Follow the red boxes above to the green one. That green box is your quarter.
Q1. Prove your agent retries cannot triple your inference bill. Seen at: Amazon-style plus AI infra loops.
Q2. Design a breaker for billed tool calls with states included. Seen at: backend plus agent loops.
Retries capped. Now catch the billing for standing still. Lesson 12: Idle billing →
Idle Still Bills
Monday. Serverless bills the wait between your function and someone else's API. One provider normalized a million requests per second into trillions of billed events per month, most of them waiting. Dependencies are slow. Your invoice pays for their pace. Keep them fast or pay for standing still.
- Time every dependency per turn. The slowest one sets the billed duration, not your code.
- Cap timeouts aggressively plus cache slow answers. Waiting is the most boring line item to defend.
- Deep dive: Vercel Fluid Compute, the bill that waited for a response.
Your artifact: a wait audit, billed wait seconds per turn per dependency, with caps set. Follow the red path above, then starve it. Fast dependencies are margin. Slow ones are rent.
Q1. A slow vendor API triples your serverless bill. What changes? Seen at: backend plus AI loops.
Q2. When does provisioned capacity beat serverless for agents? Seen at: platform plus FinOps loops.
Waiting capped. Now watch the meter live. Lesson 13: Metering →
Meter Every Turn
Tuesday. The invoice arrives. Nobody can say which endpoint, which team, or which experiment spent it. Unmetered turns are anonymous spending with full privileges. Tag every turn with endpoint plus model plus cache status, then alert on burn rate before finance does.
- Emit cost per turn with endpoint plus model plus cache tags. Attribution is the whole game.
- Alert on burn rate, not on invoices. Invoices are autopsies. Burn alerts are checkups.
- Deep dive: Cursor Token Exhaustion, the meter nobody watched.
Your artifact: a burn dashboard, spend per endpoint per day with an alert threshold plus an owner. Watch the hot endpoint trip the green alert above, then set your own tripwire. Measured spend is a budget. Unmeasured spend is weather.
Q1. Spend doubles overnight with no deploy. How do you find the cause? Seen at: SRE plus AI loops.
Q2. Design cost attribution for ten teams sharing one gateway. Seen at: platform loops.
Meter live. Now rehearse the interview. Lesson 14: Interview bench →
The Interview Bench
Wednesday. The panel asks for a voice feature at a million requests per second. The candidate who draws the meter, prices the month out loud, plus confesses a bill shock passes. The candidate who names three model versions does not. Builders draw. Tourists describe.
- Draw first, talk second. The meter with a price card opens every strong answer.
- Price the envelope out loud with the RPS envelope calculator. Panels score math over adjectives.
- Close with your bill shock. Invoices you survived beat features you shipped.
Rehearse three answers this week: a KV cache design, a routing plan with switch math, plus a burn dashboard with a tripwire story. Bring diagrams. Builders draw. Tourists describe.
Q1. Price a voice feature at a million requests per second, out loud. Seen at: startup plus CTO-round loops.
Q2. Tell me about an inference bill you caused plus fixed. Seen at: behavioral plus AI loops.
Answers rehearsed. Now assemble everything into one ship. Lesson 15: Ship endpoint →
Ship One Endpoint
Sunday. Pick one endpoint from your own work: support answers, voice greetings, or agentic drafts. Price the turn, split the phases, warm the cache, pack the batch, kill the hellos, dodge the cliff, hedge the vendor, route the easy, envelope the month, budget the evals, cap the retries, starve the wait, meter everything. Then ship it to five users, not five thousand.
- Ship the artifacts from lessons 1 through 13 in one folder. Gaps are visible at a glance.
- Pilot with five users plus a burn alert. Small blast radius, full telemetry, honest notes.
- Close with the invoice. Turns per month times modeled unit price, labeled as estimates.
Course complete. Ground the answers next: RAG in Production, then browse all courses plus all case studies. New models ship weekly. So will new price lists.
The demo charmed. Production invoices. You just built the study guide. Earn the grade.
Sources plus method
This course teaches from published Buildopsy postmortems linked under each lesson: Perplexity inference bills, Cursor token economics, ElevenLabs voice bills, Muse vendor pricing, Replit agent retries, Vercel wait billing, plus the agent eval program. Incident details follow those accounts. Cost framing is modeled from public list prices.
