Back to System Design Index

Free Course15 lessonsLive diagramsKabir

LLM Inference Economics: 15 Lessons From the Token Meter

Your demo answers beautifully at thirty cents a turn. One million turns later the invoice arrives with its own weather system. This course teaches the meter behind the magic. Meter it, batch it, cache it. Every lesson ships an artifact you can run.

TL;DR: Fifteen workshop lessons for pricing plus taming inference spend, from KV cache to vendor pacts plus the RPS envelope. Built from production postmortems with live diagrams plus calculators.

By Kabir · AI engineering course · Updated September 19, 2026

What you will be able to do

How interviews test this course: draw first, price second, confess one bill shock third. Lesson 14 rehearses the bench.

Lesson 01 · Foundations · Price one turn

The Turn Has a Price

Friday night. Your agent answers beautifully at thirty cents a turn. One million turns later the invoice arrives with its own weather system. An answer engine once paid for every question twice, retrieval plus inference, until caching made repeats nearly free. Price one turn before scaling to millions.

request $0 retrieval $0.02 inference $0.31 cache: -$0.28 find the red box first, optimize second
  • Price one turn end to end with the token bill calculator. Cache hits turn per-turn math into margin.
  • Scale to one million turns on paper before launch. Envelopes surprise only the unpriced.
  • Deep dive: Perplexity Answer Engine, paying twice per question with receipts.

Your artifact this lesson: a one-turn price card, input plus cache rate plus output plus total, then the million-turn envelope beside it. No price, no ship. The demo charmed. Production invoices.

Interview room

Q1. Price one million support turns monthly, line by line, out loud. Seen at: startup plus CTO-round loops.

Q2. What breaks first when turns triple overnight? Seen at: applied AI loops.

Turn priced. Now split the meter in two. Lesson 02: Prefill decode →

Lesson 02 · Mechanics · Two phases, two bills

Prefill Is Not Decode

Monday. Your long prompt bills heavily before generating a single word. That is prefill: the model reads everything at once, compute-bound plus hungry. Decode then drips tokens one by one, memory-bound plus slow. One request, two physics, two prices. Tune them separately.

long prompt prefill: hungry decode: drips tokens one request, two physics, two prices
  • Shorten prompts to cut prefill. Every repeated instruction block is a tax you chose.
  • Cap output length deliberately. Rambling decodes bill per token plus bore per reader.
  • Deep dive: Perplexity Answer Engine, where prefill flaws taxed every query.

Your artifact: a phase split for your heaviest endpoint, prefill share versus decode share, measured not guessed. Watch the packet slow at each stage above, then attack the bigger half first. Averages lie. Phases confess.

Interview room

Q1. Long prompts make your endpoint slow. Is it prefill or decode? Prove it. Seen at: AI infra loops.

Q2. When does a shorter prompt beat a bigger GPU? Seen at: performance-aware AI loops.

Phases split. Now reuse what prefill computed. Lesson 03: KV margin →

Lesson 03 · Caching · Keys plus values kept

KV Cache Is Margin

Tuesday. Ten thousand users share one system prompt plus the fleet recomputes it ten thousand times. The key-value cache stores that shared prefix once, then reuses it for every turn. Same answers. A fraction of the compute. Cache hits turn per-turn math into margin.

shared prefix KV cache turns: reuse suffix: fresh margin
  • Front-load shared instructions so prefixes match. Cache hits need identical beginnings, not similar intentions.
  • Measure hit rate per endpoint. Unmeasured caches are rumors with dashboards.
  • Deep dive: Perplexity Answer Engine, KV reuse done right plus prefill done wrong.

Your artifact: a prefix audit, shared tokens per endpoint plus measured hit rate plus dollars skipped. Follow the green path above, then standardize your own prefixes. Matching beginnings print money. Creative ones burn it.

Interview room

Q1. Design prompts for maximum KV cache reuse across users. Seen at: AI infra loops.

Q2. Cache hit rate is 5 percent. What do you change first? Seen at: applied AI loops.

Prefix cached. Now pack the GPU tighter. Lesson 04: Batching →

Lesson 04 · Throughput · Pack the batch

Batch or Bleed

Wednesday. Your GPUs serve one request at a time while thirty wait in line. Utilization naps at 20 percent. Continuous batching packs waiting requests into every forward pass, so the same silicon serves multiples. Idle GPUs bill like busy ones. Pack them or pay for standing still.

alone: 20 pct waiting: 30 batch: packed same GPU, 4x turns latency: guarded throughput is a packing problem first
  • Batch continuously, not statically. Requests join plus leave mid-generation instead of waiting for stragglers.
  • Guard tail latency while packing. Throughput that tortures p99 is a transfer, not a win.
  • Deep dive: The Voice That Bills Per Hello, throughput math at a billion-scale bill.

Your artifact: a batching test, throughput versus p99 at four batch sizes, drawn as one curve. Follow the green box above, then pick the size at the knee. Packed GPUs print turns. Lonely ones print invoices.

Interview room

Q1. GPU utilization sits at 20 percent with a queue forming. Fix it live. Seen at: AI infra loops.

Q2. Batching helps throughput but hurts latency. Where is the line? Seen at: performance-aware AI loops.

GPU packed. Now kill the most repeated work. Lesson 05: Hello tax →

Lesson 05 · Repetition · Cache the hello

Hello Costs 75ms

Thursday. Voice AI bills every hello at 75 milliseconds of GPU. The cached hello answers in 5. Same word. Fifteen times cheaper. Somewhere a quarter of all greetings repeat, which means a quarter of the invoice is optional. Your pipeline has hellos too. Find them.

hello? cached: 5ms fresh: 75ms repeat rate saved 25pct
  • Log plus rank repeated inputs weekly. The top ten repeats fund the whole cache layer.
  • Cache audio, embeddings, plus full answers separately. Each repeat shape needs its own shelf.
  • Deep dive: The Voice That Bills Per Hello, the hello with a receipt.

Your artifact: a repeat report, top repeated inputs plus their billed cost plus cache design. Follow the green path above, then kill your own hellos. Repeated work is a subscription you never signed.

Interview room

Q1. A quarter of traffic repeats daily greetings. Design the cache. Seen at: AI product loops.

Q2. Cached answers go stale on voice. How do you version them? Seen at: applied AI loops.

Hellos cached. Now survive the cliff at 11:47 AM. Lesson 06: Compaction →

Lesson 06 · Context · The double bill

Compaction Bills Twice

Friday. Your coding agent flies all morning, then hits a paywall at 11:47 AM with the sprint half done. Compaction crushed hundreds of thousands of tokens into thousands, then billed dozens of re-reading steps to rebuild what was lost. The meter runs twice. The auditor naps through both.

context: full cliff: crush thin recap rebuild

Your artifact: a context budget per session with a compaction trigger you chose. Watch the packet cross the cliff above, then decide where your cliff sits. Chosen cliffs are architecture. Surprise cliffs are incidents.

Interview room

Q1. Your agent degrades after 50 turns plus bills double. Diagnose live. Seen at: applied AI loops.

Q2. When does a bigger window beat a better summary? Seen at: LLM platform loops.

Cliff priced. Now watch vendors move the prices. Lesson 07: Price pact →

Lesson 07 · Vendors · Pacts break, lists ship

The Price Pact Broke

Monday. Three labs agree to slow down plus publish safety pacts. By Friday one ships a cheaper flagship that undercuts the pact by half. Price lists move faster than promises. Your margin plan must survive the vendor you chose plus the three you did not.

flagship: $15 mid: $3 mini: $1.25 your router traffic shifts
  • Abstract the model behind your router, never behind hope. Swaps should take a config change, not a rewrite.
  • Track price per million monthly like a dependency version. Moves you watch cost little. Moves you miss cost quarters.
  • Deep dive: Muse Price Pact Fracture, the week pacts met price lists.

Your artifact: a vendor sheet, price per million per model plus switch cost plus trigger thresholds. Follow the packet down to the green tier above, then set your own shift rules. Loyalty is a strategy. Lock-in is a bill.

Interview room

Q1. Your vendor doubles prices with 30 days notice. What moves first? Seen at: startup plus CTO-round loops.

Q2. When is single-vendor depth worth the lock-in risk? Seen at: platform plus AI loops.

Pact read. Now route every query to its cheapest sufficient mind. Lesson 08: Routing →

Lesson 08 · Routing · Small minds, big savings

Route Small, Spend Small

Tuesday. Eighty percent of your queries need a sledgehammer, says nobody who measured. Most turns classify, summarize, or extract, work a small model does at a tenth of the price. Route by difficulty. Pay flagship rates only for flagship problems.

query classifier easy: mini hard: flagship save 70pct
  • Classify difficulty before generating. A cheap judge routing to a cheaper worker beats hope at scale.
  • Escalate with evidence. Misrouted hard queries cost more in retries than the flagship would have.
  • Deep dive: Perplexity Answer Engine, tiered work with tiered bills.

Your artifact: a routing table, query shape plus model plus observed quality plus unit price. Watch easy traffic take the green path above, then route your own easy majority down. Flagship brains for flagship problems. Mini brains for everything else.

Interview room

Q1. Cut inference spend 50 percent without touching quality. Show the routing. Seen at: FinOps-flavored AI loops.

Q2. The classifier misroutes hard queries down. How do you catch it? Seen at: evals-minded AI loops.

Traffic routed. Now scale the math to the month. Lesson 09: RPS envelope →

Lesson 09 · Scale · The monthly envelope

The RPS Envelope

Wednesday. Thirty cents a turn feels trivial until one million turns a month turn it into a headcount. The envelope math is brutal plus simple: requests per second times seconds per month times unit price. Every pricing conversation that skips this step ends in a surprise.

40 RPS 2.6M sec/mo $0.33/turn $34k/mo modeled envelope, labeled as estimates

Your artifact: a monthly envelope for your endpoint, base plus growth plus peak, all labeled as estimates. Follow the waterfall above to the green total. That total is your budget conversation.

Interview room

Q1. Price 100 RPS of agent traffic for a year, out loud. Seen at: startup plus CTO-round loops.

Q2. Growth triples RPS but the budget doubles. What gives? Seen at: platform plus AI loops.

Envelope sealed. Now add the quality line items. Lesson 10: Guardrail cost →

Lesson 10 · Quality · Safety has a meter

Guardrails Cost Too

Thursday. Evals, judges, plus red-team runs all consume the product they protect. A golden set of five hundred cases re-run on every merge is a second inference bill wearing a quality costume. Budget it openly or quality becomes the first cut when money tightens.

prompt edit golden 500 judge pass eval spend quality is a line item, not a footnote
  • Budget eval spend per merge plus per release. Quality with a meter survives budget season.
  • Verify mechanically where possible. Proofs plus schemas cost less than judges per check.
  • Bridge course: AI Evals That Hold, where the meter becomes a program.

Your artifact: an eval budget, cases times runs times unit price per month. Watch spend accumulate at the red box above, then defend it as insurance with a number. Unpriced quality gets cut first.

Interview room

Q1. The CFO asks to cut eval spend in half. What do you protect? Seen at: startup plus leadership loops.

Q2. When is a code check enough plus when do you need a judge? Seen at: evals-minded loops.

Quality budgeted. Now stop retries from re-billing. Lesson 11: Retry meter →

Lesson 11 · Failure · Failed turns still bill

Retries Multiply the Meter

Friday. Your agents retried failed deploys until retries cost more than features. Every retry is a second billed turn arriving exactly when the system is weakest. At two extra retries per stalled call with no backoff, recovery load lands near six times normal. Failed turns bill like successful ones.

fail $0.33 retry x2 6x load breaker 1x
  • Budget retries like capacity. Run the modeled shape through the retry storm calculator before the incident.
  • Cap attempts plus open breakers while errors rage. Unbounded retries are a self-inflicted invoice.
  • Deep dive: Replit Agent Runtime, where retries cost more than features.

Your artifact: a retry policy with numbers, max attempts plus base delay plus jitter plus breaker threshold, each with a billed cost. Follow the red boxes above to the green one. That green box is your quarter.

Interview room

Q1. Prove your agent retries cannot triple your inference bill. Seen at: Amazon-style plus AI infra loops.

Q2. Design a breaker for billed tool calls with states included. Seen at: backend plus agent loops.

Retries capped. Now catch the billing for standing still. Lesson 12: Idle billing →

Lesson 12 · Serverless · Waiting is billable

Idle Still Bills

Monday. Serverless bills the wait between your function and someone else's API. One provider normalized a million requests per second into trillions of billed events per month, most of them waiting. Dependencies are slow. Your invoice pays for their pace. Keep them fast or pay for standing still.

function slow API: bills fast API: stops timeout cap meter rests
  • Time every dependency per turn. The slowest one sets the billed duration, not your code.
  • Cap timeouts aggressively plus cache slow answers. Waiting is the most boring line item to defend.
  • Deep dive: Vercel Fluid Compute, the bill that waited for a response.

Your artifact: a wait audit, billed wait seconds per turn per dependency, with caps set. Follow the red path above, then starve it. Fast dependencies are margin. Slow ones are rent.

Interview room

Q1. A slow vendor API triples your serverless bill. What changes? Seen at: backend plus AI loops.

Q2. When does provisioned capacity beat serverless for agents? Seen at: platform plus FinOps loops.

Waiting capped. Now watch the meter live. Lesson 13: Metering →

Lesson 13 · Observability · Dashboards before invoices

Meter Every Turn

Tuesday. The invoice arrives. Nobody can say which endpoint, which team, or which experiment spent it. Unmetered turns are anonymous spending with full privileges. Tag every turn with endpoint plus model plus cache status, then alert on burn rate before finance does.

tagged turn chat: calm agent: hot eval: planned burn alert owner paged
  • Emit cost per turn with endpoint plus model plus cache tags. Attribution is the whole game.
  • Alert on burn rate, not on invoices. Invoices are autopsies. Burn alerts are checkups.
  • Deep dive: Cursor Token Exhaustion, the meter nobody watched.

Your artifact: a burn dashboard, spend per endpoint per day with an alert threshold plus an owner. Watch the hot endpoint trip the green alert above, then set your own tripwire. Measured spend is a budget. Unmeasured spend is weather.

Interview room

Q1. Spend doubles overnight with no deploy. How do you find the cause? Seen at: SRE plus AI loops.

Q2. Design cost attribution for ten teams sharing one gateway. Seen at: platform loops.

Meter live. Now rehearse the interview. Lesson 14: Interview bench →

Lesson 14 · Interviews · Draw plus price plus confess

The Interview Bench

Wednesday. The panel asks for a voice feature at a million requests per second. The candidate who draws the meter, prices the month out loud, plus confesses a bill shock passes. The candidate who names three model versions does not. Builders draw. Tourists describe.

draw meter price month shock told offer
  • Draw first, talk second. The meter with a price card opens every strong answer.
  • Price the envelope out loud with the RPS envelope calculator. Panels score math over adjectives.
  • Close with your bill shock. Invoices you survived beat features you shipped.

Rehearse three answers this week: a KV cache design, a routing plan with switch math, plus a burn dashboard with a tripwire story. Bring diagrams. Builders draw. Tourists describe.

Interview room

Q1. Price a voice feature at a million requests per second, out loud. Seen at: startup plus CTO-round loops.

Q2. Tell me about an inference bill you caused plus fixed. Seen at: behavioral plus AI loops.

Answers rehearsed. Now assemble everything into one ship. Lesson 15: Ship endpoint →

Lesson 15 · Capstone · Meter plus batch plus cache

Ship One Endpoint

Sunday. Pick one endpoint from your own work: support answers, voice greetings, or agentic drafts. Price the turn, split the phases, warm the cache, pack the batch, kill the hellos, dodge the cliff, hedge the vendor, route the easy, envelope the month, budget the evals, cap the retries, starve the wait, meter everything. Then ship it to five users, not five thousand.

turn priced cache warm hedges set 5 users, metered five users first, five thousand later
  • Ship the artifacts from lessons 1 through 13 in one folder. Gaps are visible at a glance.
  • Pilot with five users plus a burn alert. Small blast radius, full telemetry, honest notes.
  • Close with the invoice. Turns per month times modeled unit price, labeled as estimates.

Course complete. Ground the answers next: RAG in Production, then browse all courses plus all case studies. New models ship weekly. So will new price lists.

The demo charmed. Production invoices. You just built the study guide. Earn the grade.

Sources plus method

This course teaches from published Buildopsy postmortems linked under each lesson: Perplexity inference bills, Cursor token economics, ElevenLabs voice bills, Muse vendor pricing, Replit agent retries, Vercel wait billing, plus the agent eval program. Incident details follow those accounts. Cost framing is modeled from public list prices.